File management system and file classification coding method thereof

By introducing intelligent dense cabinets, optical character recognition technology and RoBERTa-wwm-ext-based archive classification model in the archive management system, the problems of archive access, electronic review and coding in traditional archive management are solved, and efficient, safe and convenient archive management is achieved.

CN120218543AActive Publication Date: 2025-06-27SHAANXI HUISHENG SPACE-TIME INFORMATION TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510353748.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-27
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

Traditional archive management has many problems in archive access, electronic review and archive coding, resulting in low management efficiency, insufficient information security and inability to meet the modern society's needs for rapid acquisition and precise management of archive information.

Method used

It provides an archive management system, including an archive registration end, an intelligent dense cabinet and an archive review end. It uses optical character recognition technology to generate electronic files, combines the archive classification model based on RoBERTa-wwm-ext for classification, uses NFC tags and structured numbering for encoding, and realizes efficient archive management and secure electronic review through the archive cloud and review and recording module.

Benefits of technology

It realizes rapid positioning and efficient access to archives, improves the automation level and operational convenience of archive management, ensures the security of archives and the rapid acquisition of information, and meets the precise management needs of modern society for archive information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218543A_ABST
    Figure CN120218543A_ABST
Patent Text Reader

Abstract

The invention relates to an archive management system and an archive classification coding method thereof, and belongs to the technical field of file management. The archive registration end is provided with an archive scanning module and an archive coding module; the archive storage library is provided with a plurality of archive storage cabinets, an archive cloud end and an archive destroying module; the archive retrieval end is provided with an electronic archive retrieval module, an entity archive retrieval module and a retrieval recording module; the automation level and the operation convenience of archive management are greatly improved; secondly, the archive scanning module is combined with an archive classification model based on RoBERTa-wwm-ext, the accuracy of archive classification is improved, in addition, the archive coding module adopts a structured numbering mode, archive classification, storage positions and unique identifiers are combined, archive numbers with high readability and uniqueness are generated, and archive retrieval and management are facilitated; the integration of the environment monitoring module effectively prevents archive damage caused by environment factors by monitoring and adjusting the environment parameters of the archive room in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of document management, and specifically relates to an archive management system and its archive classification and coding method. Background Art

[0002] In today's information age, as archives serve as the carriers for various organizations and institutions to record important information, decision-making bases, and historical data, the effectiveness and efficiency of their management play a crucial role in the operation, development, and compliance requirements of the organization. However, when faced with the increasing number of archives and complex management requirements, the traditional archive management mode has exposed many defects and deficiencies.

[0003] Dilemma in archive access and storage: Traditional archive management mainly relies on manual operations for archive access and storage. In large archive repositories or data rooms, archives are usually stored on compact shelves or ordinary shelves. When new archives need to be stored, the administrator needs to manually find a suitable empty space and place the archive box in it. This process depends on the administrator's familiarity with the layout of the archive repository and personal experience, and it is easy to have misplacement or inaccurate recording of the storage location. When retrieving archives, the administrator first needs to determine the approximate location of the archives based on the archive index or retrieval request, and then search through numerous archive shelves one by one. Especially for old or frequently borrowed archives, the entire retrieval process may take a lot of time, seriously affecting work efficiency.

[0004] Lag in electronic retrieval: With the wide popularization of digital technology, electronic documents have shown great advantages in information transmission and sharing. However, traditional archive management has made slow progress in electronic retrieval. Most archives still exist in paper form. Even if some archives have been digitized and scanned, their retrieval methods are relatively cumbersome. Usually, the requester needs to submit a written application to the archive management department, stating information such as the purpose of retrieval and the scope of required archives, and then wait for the administrator to search and extract the corresponding electronic archive files in the local electronic database or storage server. If cross-departmental or cross-regional retrieval requirements are involved, file transfer needs to be carried out through internal network transmission or mailing storage media, etc. This not only easily causes the risk of information leakage but also cannot achieve real-time and multi-person simultaneous online electronic retrieval. For example, when an enterprise conducts an audit, audit teams in different regions need to simultaneously retrieve the financial archives of the company headquarters and each branch. The traditional electronic retrieval method cannot meet this immediacy and collaboration requirement, seriously restricting the efficient development of business.

[0005] Chaos and Inefficiency in File Coding: File coding is an important part of the file management system. Its purpose is to assign a unique identification code to each file for easy classification, storage, retrieval, and management of files. However, there are many problems in file coding in traditional file management. On the one hand, there is a lack of unified coding standards and specifications. The file coding rules vary greatly among different units, different industries, and even different departments within the same unit. For example, some units adopt a coding method of year + department + serial number, while others code according to file type + storage period + number. This makes it extremely difficult for files to be exchanged and shared across departments or institutions. On the other hand, traditional file coding is mostly manually operated, prone to problems such as coding errors, duplicate coding, and untimely coding updates. With the continuous increase in the number of files and the continuous development of business, the classification and management requirements of files are also constantly changing. Manual coding is difficult to adapt to this dynamic management requirement, resulting in low file retrieval efficiency and increased management costs. For example, in the personnel file management of some large enterprises, due to frequent occurrences of employee recruitment, resignation, and job transfer, the coding of personnel files needs to be continuously adjusted and updated. The manual coding method is difficult to ensure the accuracy and timeliness of coding, thus affecting the efficient development of human resource management work.

[0006] In summary, the many problems in traditional file management in aspects such as file access, electronic retrieval, and file coding have severely restricted the efficiency, quality, and security of file management work, and cannot meet the needs of modern society for rapid access to file information, precise management, and long-term preservation. Therefore, developing a file management system with convenient file access, electronic retrieval, and scientific file coding functions has extremely important practical significance and broad application prospects. Summary of the Invention

[0007] This application provides a file management system and its file classification and coding method, which solve many problems existing in the prior art in file access, electronic retrieval, and file coding.

[0008] To achieve the above object, this application adopts the following technical solutions:

[0009] In the first aspect, this application provides a file management system, including: a file registration terminal, a file storage repository, and a file retrieval terminal;

[0010] The file registration terminal is provided with: a file scanning module and a file coding module;

[0011] The file storage repository is an intelligent compact cabinet, provided with: a plurality of file storage cabinets, a file cloud, and a file destruction module;

[0012] The file retrieval terminal is provided with: an electronic file retrieval module, a physical file retrieval module, and a retrieval record module;

[0013] The file scanning module performs optical character recognition on the physical files to be warehoused, recognizes the file content and generates editable electronic files;

[0014] The file repository quickly assigns storage locations for the newly warehoused physical files and prints storage barcodes;

[0015] The file coding module generates file numbers for the scanned electronic files, enters the file numbers into the electronic file text, and at the same time makes NFC tags for the physical files and writes the file codes into the NFC tags;

[0016] The physical files pasted with NFC tags and storage barcodes are received by the file storage cabinet. The file storage cabinet unlocks by identifying the NFC tags or storage barcodes to complete the storage of physical files. The file storage cabinet is equipped with a file access control module, and the administrator operates the file access control module to access files;

[0017] The numbered electronic files are uploaded to the file cloud;

[0018] The file destruction module is equipped with a shredder and a data erasure module for destroying physical files and electronic files that have reached the end of their storage period and have no value for preservation. The destruction process is automatically recorded and a destruction report is generated;

[0019] The electronic file access module provides online access services for electronic files. The physical file access module manages the access process of physical files. The access record module records all file access and update information.

[0020] Furthermore, the file scanning module deploys a file classification model based on RoBERTa-wwm-ext. The model is trained with electronic files obtained by optical character recognition, and the model parameters are updated regularly using new file data. The model is optimized through knowledge distillation and elastic weight consolidation to obtain regularly updated file classifications.

[0021] Furthermore, the file coding module generates file numbers based on file classifications, combined with file classification types, file storage locations, years, and unique identifiers.

[0022] Furthermore, the file access module is tightly integrated with the file cloud. The administrator can quickly retrieve files by keywords, file numbers, classification types, and years; the file access module provides hierarchical permission management. The administrator sets access permissions for different users of the file access module to ensure that sensitive files are only open to authorized personnel.

[0023] Further, the entity file access module manages the access process of entity files. The administrator submits an entity file access application through the file access module. The entity file access module locates the file storage location and notifies the administrator, and drives the file access control module of the file storage cabinet to open the file storage cabinet.

[0024] Further, the access record module records the access and update information of all files, generates an access log, records the file scanning time, accessor, access time, access purpose information, and destruction record, and generates an access data analysis report to help the administrator optimize the file management process.

[0025] Further, the file management system further includes an environmental monitoring module. The environmental monitoring module monitors the environmental parameters of the archive room in real time to ensure the long-term preservation of files; the environmental monitoring module is integrated with: temperature and humidity monitoring and regulating equipment, fire alarm equipment, air conditioning equipment, and lighting regulating equipment;

[0026] The temperature and humidity monitoring and regulating equipment monitors the temperature and humidity of the file storage cabinet in real time. When the temperature and humidity are lower than 14°C or higher than 24°C, and the humidity is lower than 43% or higher than 60%, the temperature and humidity of the monitored file storage cabinet are adjusted.

[0027] The fire alarm equipment monitors the smoke concentration of the file storage cabinet in real time. When smoke is detected, it alarms and extinguishes the fire and removes the smoke.

[0028] The air conditioning equipment monitors the dust concentration and harmful gas concentration of the file storage cabinet in real time. When the dust concentration exceeds 0.15mg / m³ and PM2.5 exceeds 75μg / m³, air purification is carried out. When the concentration of sulfur dioxide and nitrogen oxides exceeds 0.01mg / m³, ventilation is carried out.

[0029] The lighting regulating equipment can adjust the lights of the file storage cabinet and the room, and perform functions such as lighting and ultraviolet irradiation sterilization.

[0030] In a second aspect, the present application provides a file classification and coding method, including: file classification and file numbering;

[0031] In file classification:

[0032] Step A1: Data preparation, collect electronic files extracted by optical character recognition technology, perform data preprocessing, and do data cleaning and word segmentation;

[0033] Specifically, first clean the file text: remove noise characters, stop words, and irrelevant symbols, remove common stop words, and perform standard formatting;

[0034] Secondly, perform word segmentation: Use the Jieba word segmentation technology, load the custom dictionary for word segmentation, add the [CLS] mark at the beginning of the segmented text, and add the [SEP] mark at the end;

[0035] Perform label assignment: Assign a unique label to each type of archival text. For example: {Financial archives: fi, Personnel archives: hr, Administrative archives: ad, Project archives: pr}.

[0036] Step A2: Feature extraction. Convert the text data prepared in Step 1 into numerical values that can be understood by a computer, and extract feature vectors that can reflect their semantic information from different types of archival texts;

[0037] Specifically: Load the RoBERTa-wwm-ext pre-trained model and its corresponding tokenizer;

[0038] Encode archival text: Use the tokenizer of RoBERTa-wwm-ext to encode the cleaned and segmented archival text, convert each word (token) into the corresponding ID in the model vocabulary, and for words out of the vocabulary (Out-of-Vocabulary, OOV), use the special mark [UNK] to represent;

[0039] Generate archival text representation: Input the encoded archival text into the RoBERTa-wwm-ext model to obtain the hidden states of each token. The output of the RoBERTa-wwm-ext model is a hidden state matrix:

[0040]

[0041] Among them, is the hidden state matrix of all tokens, is the number of tokens in the archival text, is the th hidden state of the token, with a dimension of 768.

[0042] Extract global features: Extract the hidden state corresponding to the [CLS] mark from the hidden state matrix , as the global semantic representation of the entire archival text. The hidden state of the [CLS] mark is optimized by the pre-trained model and can capture the overall semantic information of the archival text.

[0043] The feature vector of each archival text sample is:

[0044]

[0045] Among them, is the hidden state marked by [CLS], with a dimension of 768;

[0046] Construct a feature matrix: Concatenate the hidden states of the [CLS] markers of all samples into a feature matrix X, which serves as the input feature for the classification model:

[0047]

[0048] where, is the feature matrix, is the number of archival data samples, is the th sample's feature vector.

[0049] Step A3: Classification model design; The structure of the classification model includes the RoBERTa-wwm-ext model, a pooling layer, and a fully connected layer. The RoBERTa-wwm-ext model is used to generate the context representation of the archival text. The pooling layer performs pooling on the output of RoBERTa-wwm-ext to extract global features. The fully connected layer maps the pooled features to the classification label space and outputs the class probabilities;

[0050] Step A4: Model training. Divide the dataset into a training set, a validation set, and a test set. Use the training set to train the model. Set the hyperparameters of the learning rate, batch size, and number of training epochs. Calculate the difference between the model prediction value and the true label through the cross-entropy loss function. Use the Adam optimizer to update the model parameters by combining the first-order moment estimate and the second-order moment estimate, and evaluate the model performance on the validation set and test its accuracy on the test set. The cross-entropy loss function formula is:

[0051]

[0052] where, is the cross-entropy loss, is the number of archival data samples, is the number of archival categories, is the true label, indicating whether the th sample belongs to the th category. When the sample belongs to the category , is 1, otherwise it is 0. is the probability that the model predicts the sample belongs to the category ;

[0053] Step A5: Model deployment. Deploy the trained model as a RESTful API for the archival scanning module to call, receive the archival text, and return the classification result;

[0054] Step A6: Continuous learning. Set regular tasks to collect the archived data newly stored in the database, update the model parameters using the new data, and prevent forgetting of the existing knowledge.

[0055] Step A61: Data preparation and preprocessing. Regularly collect the archived data newly stored in the archive scanning module, clean and denoise the new data, assign category labels to the new data, and increase the diversity of the data through synonym replacement and random deletion techniques;

[0056] Step A62: Model initialization. Use the previously trained model as the initial model. Freeze some layers of the BERT model in the initial stage and only train the fully connected layer to prevent the model from overfitting on the new data;

[0057] Step A63: Loss function design. Use the initial model as the teacher model and the new model as the student model. Retain the existing knowledge through the distillation loss function. The distillation loss function is:

[0058]

[0059] where, is the distillation loss, is the number of archived data samples, is the number of archive categories, represents the probability that the teacher model believes the sample belongs to the category , represents the probability that the student model believes the sample belongs to the category ;

[0060] Elastic Weight Consolidation (EWC) loss. Elastic Weight Consolidation is a continuous learning technique that prevents the model from forgetting the knowledge of old tasks by penalizing changes in important weights. The Elastic Weight Consolidation loss function is:

[0061]

[0062] where, is the penalty coefficient used to control the intensity of the EWC loss, the larger it is, the stronger the model retains the old knowledge, is the importance of the weight, representing the importance of the -th weight for the old task, is the -th weight of the current model, is the -th weight of the initial model;

[0063] Determined by experimental adjustment, set the initial value according to experience, verify the performance of the evaluation model, and select the value Calculated through the Fisher information matrix is a parameter that is continuously updated during the model training process. In each iteration, it is updated through the Adam optimizer , Train the model on the old task until the model converges, and save the weights;

[0064] Calculate the cross-entropy loss, distillation loss, and EWC loss. Combine the distillation loss, EWC loss, and cross-entropy loss to determine the total loss, and use the total loss function to train the student model:

[0065]

[0066] Among them, is the weight coefficient of the knowledge distillation loss, which is used to control the influence of knowledge distillation on the total loss is the weight coefficient of the elastic weight consolidation loss, which is used to control the influence of EWC on the total loss , , : They are the cross-entropy loss, knowledge distillation loss, and elastic weight consolidation loss respectively;

[0067] The larger it is, the more the model tends to imitate the behavior of the teacher model and retain the knowledge of the old task The larger it is, the more the model tends to retain the key weights of the old task and prevent forgetting and The values are selected through grid search experiments to find the combination that performs well on the new task and forgets the least about the old task to determine.

[0068] Step A64: Use the Adam optimizer to train the model, update the model parameters, record the accuracy and recall metrics of the model, and ensure that the model does not forget the old knowledge;

[0069] Step A65: Model deployment. Deploy the optimized model as a RESTful API, receive the archive text and return the classification result, set a regular task to automatically collect new data and update the model.

[0070] In the archive number:

[0071] Step S1: Formulate the archive number structure: year - classification code - archive storage code - unique identification code;

[0072] Among them, the year is the time when the file is scanned by the file scanning module, represented by 4 digits;

[0073] The classification code is the code generated by the file classification model, represented by 2 - 3 letters;

[0074] The file storage code is the storage location code assigned by the file storage repository to the physical file, represented by 1 letter + 2 digits;

[0075] The unique identification code is the file sequence code in a certain file storage cabinet, represented by 3 digits;

[0076] Step S2: File number generation;

[0077] Step S21: Access the access record module to obtain the file scanning time;

[0078] Step S22: Classify the file through the file classification model to obtain the classification code;

[0079] Step S23: Access the file storage repository to obtain the storage location code assigned to the physical file and the file sequence code in this file storage cabinet;

[0080] Step S24: Combine the year, classification code, file storage code and unique identifier to generate the final file number;

[0081] Step S3: File number entry. Enter the file number into the electronic file text, and at the same time make an NFC tag for the physical file and write the file code into the NFC tag.

[0082] Due to the adoption of the above technical solutions, this application has the following beneficial effects:

[0083] The file management system and its file classification and coding method provided by this application introduce intelligent compact shelves and NFC tag technology, enabling the system to achieve rapid positioning and efficient access of physical files, greatly improving the automation level and operation convenience of file management. Secondly, the file scanning module combines with a file classification model based on RoBERTa-wwm-ext, which not only improves the accuracy of file classification but also ensures the timeliness and adaptability of the classification model through continuous learning and model optimization, effectively coping with the dynamic changes of file data. In addition, the file coding module adopts a structured numbering method, combining file classification, storage location, and unique identifiers to generate highly readable and unique file numbers, facilitating file retrieval and management. The close integration of the file access module with the file cloud provides efficient electronic file access services and ensures the security of sensitive files through hierarchical permission management. The introduction of the access record module not only details the access and update information of files but also generates data analysis reports, providing data support for administrators to optimize the file management process. The integration of the environmental monitoring module further guarantees the long-term preservation of files by real-time monitoring and adjusting the environmental parameters of the file room, effectively preventing file damage caused by environmental factors. Brief Description of the Drawings

[0084] The drawings described herein are used to provide a further understanding of this application, form a part of this application, and do not constitute an improper limitation to this application. In the drawings:

[0085] Figure 1 It is a schematic structural diagram of the file management system provided by this application;

[0086] Figure 2 It is a schematic flowchart of model training in the file classification and coding method provided by this application;

[0087] Figure 3 It is a schematic flowchart of the coding process of the file classification and coding method provided by this application. Detailed Embodiments

[0088] The following will detail this application in combination with the drawings and specific embodiments. Here, the illustrative embodiments and descriptions of this application are used to explain this application but do not serve as a limitation to this application.

[0089] Embodiment 1: Please refer to Figure 1 , this application provides a file management system, including: a file registration terminal, a file storage repository, and a file access terminal;

[0090] The file registration terminal is provided with: a file scanning module and a file coding module;

[0091] The file storage repository is an intelligent compact shelf, provided with: a plurality of file storage cabinets, a file cloud, and a file destruction module;

[0092] The file access terminal is provided with: an electronic file access module, a physical file access module, and an access record module.

[0093] Specifically, the file scanning module performs optical character recognition on the physical files to be warehoused, identifies the file content and generates editable electronic files; the file repository quickly assigns storage locations to the newly warehoused physical files and prints storage barcodes; the file coding module generates file numbers for the scanned electronic files, enters the file numbers into the electronic file text, and the numbered electronic files are automatically uploaded to the file cloud. At the same time, NFC tags are made for the physical files, and the file codes are written into the NFC tags; the physical files pasted with NFC tags and storage barcodes are received by the file storage cabinet, and the file storage cabinet unlocks by identifying the NFC tag or the storage barcode to complete the storage of the physical files. The file storage cabinet is provided with a file access control module, and the administrator operates the file access control module to access the files; the file destruction module is provided with a shredder and a data erasure module for destroying physical files and electronic files that have reached the expiration of storage and have no value for preservation. The destruction process is automatically recorded and a destruction report is generated; the electronic file access module provides online access services for electronic files, the physical file access module manages the access process of physical files, and the access record module records all file access and update information.

[0094] It should be further noted in this embodiment that the file scanning module deploys a file classification model based on RoBERTa-wwm-ext. The model is trained with electronic files obtained by optical character recognition, and the model parameters are updated regularly using new file data. The model is optimized through knowledge distillation and elastic weight consolidation to obtain a regularly updated file classification.

[0095] It should be noted that RoBERTa-wwm-ext is the product of continuous development and optimization in the field of natural language processing. It originated from the improvement of the BERT model. RoBERTa has made robust optimizations in the pre-training method, removed the next sentence prediction task, focused on the masked language model task, and at the same time increased the training batch size and extended the training time; "wwm" means whole word masking, which changes the traditional masking method and masks all the tokens of a complete vocabulary; "ext" means further expansion in data or training methods. This model has many remarkable features; in terms of semantic understanding, the whole word masking technology enables it to accurately capture the semantic information of complete vocabulary, enhancing the understanding accuracy of Chinese text vocabulary; its optimized pre-training strategy allows the model to learn on large-scale text data for a long time and have stronger generalization ability. These features give RoBERTa-wwm-ext outstanding advantages in the file classification task. File texts often contain a large number of professional terms and vocabulary in specific contexts. The model's powerful semantic understanding ability can accurately identify the meanings of these vocabulary and sentences, thus classifying the file content more precisely. At the same time, due to the diverse types and wide sources of file data, the model's generalization ability enables it to have stable and excellent performance on different types of file data, effectively improving the efficiency and accuracy of file classification.

[0096] It should be further noted in this embodiment that, based on file classification, the file encoding module generates a file number by combining the file classification type, file storage location, year, and unique identifier.

[0097] It should be further noted in this embodiment that the file retrieval module is closely integrated with the file cloud. The administrator can quickly retrieve files by keywords, file numbers, classification types, and years; the file retrieval module provides permission hierarchical management. The administrator sets access permissions for different users of the file retrieval module to ensure that sensitive files are only open to authorized personnel.

[0098] It should be further noted in this embodiment that the physical file retrieval module manages the retrieval process of physical files. The administrator submits a physical file retrieval application through the file retrieval module. The physical file retrieval module locates the file storage location and notifies the administrator, and drives the file retrieval control module of the file storage cabinet to open the file storage cabinet.

[0099] It should be further noted in this embodiment that the retrieval record module records all file retrieval and update information, generates a retrieval log, records file scanning time, retrievers, retrieval time, retrieval purpose information, and destruction records, and generates a retrieval data analysis report to help the administrator optimize the file management process.

[0100] Embodiment 2: Please refer to Figure 1, based on the first embodiment of this application, the file management system further includes an environmental monitoring module, which monitors the environmental parameters of the file storage room in real time to ensure the long-term preservation of files.

[0101] The environmental monitoring module integrates: temperature and humidity monitoring and adjustment equipment, fire alarm equipment, air conditioning equipment, and lighting adjustment equipment;

[0102] The temperature and humidity monitoring and adjustment equipment monitors the temperature and humidity of the file storage cabinet in real time. When the temperature is lower than 14°C or higher than 24°C, and the humidity is lower than 43% or higher than 60%, it adjusts the temperature and humidity of the monitored file storage cabinet.

[0103] The fire alarm equipment monitors the smoke concentration in the file storage cabinet in real time. When smoke is detected, it alarms and extinguishes the fire and removes the smoke.

[0104] The air conditioning equipment monitors the dust concentration and harmful gas concentration in the file storage cabinet in real time. When the dust concentration exceeds 0.15mg / m³ and PM2.5 exceeds 75μg / m³, it purifies the air. When the concentration of sulfur dioxide and nitrogen oxides exceeds 0.01mg / m³, it ventilates.

[0105] The lighting adjustment equipment can adjust the lights in the file storage cabinet and the room, and perform the functions of lighting and ultraviolet irradiation sterilization.

[0106] Embodiment 3: A file classification and coding method is provided, including: file classification and file numbering;

[0107] Please refer to Figure 2 , in file classification:

[0108] Step A1: Data preparation, collect electronic files extracted by optical character recognition technology, perform data preprocessing, and do data cleaning and word segmentation;

[0109] Specifically, first clean the file text: remove noise characters, stop words, and irrelevant symbols, such as special characters like "¥", "$", etc., irrelevant content such as table borders and annotations, remove common stop words, such as "de", "shi", etc., and perform standard formatting: unify the format of the amount to "¥100.00", unify the format of the date to "YYYY-MM-DD", and convert numbers to unified Arabic numerals;

[0110] Secondly, perform word segmentation: use the Jieba word segmentation technology, load the custom dictionary for word segmentation, add the [CLS] mark at the beginning of the segmented text, and add the [SEP] mark at the end;

[0111] Perform label assignment: Assign a unique label to each type of archival text. For example: {Financial archives: fi, Personnel archives: hr, Administrative archives: ad, Project archives: pr}.

[0112] Step A2: Feature extraction. Convert the text data prepared in Step 1 into numerical values that can be understood by a computer, and extract feature vectors that can reflect their semantic information from different types of archival text.

[0113] Specifically: Load the RoBERTa-wwm-ext pre-trained model and its corresponding tokenizer. RoBERTa-wwm-ext is a pre-trained language model based on the Transformer architecture, suitable for feature extraction of Chinese text.

[0114] Archival text encoding: Use the tokenizer of RoBERTa-wwm-ext to encode the cleaned and tokenized archival text, convert each word (token) into the corresponding ID in the model vocabulary, and use the special token [UNK] to represent words that are out of the vocabulary (Out-of-Vocabulary, OOV).

[0115] Generate archival text representation:

[0116] Input the encoded archival text into the RoBERTa-wwm-ext model to obtain the hidden states of each token. The output of the RoBERTa-wwm-ext model is a hidden state matrix:

[0117]

[0118] Among them, is the hidden state matrix of all tokens, is the number of tokens in the archival text, is the th hidden state of the token, with a dimension of 768.

[0119] Extract global features:

[0120] Extract the hidden state corresponding to the [CLS] token from the hidden state matrix as the global semantic representation of the entire archival text. The hidden state of the [CLS] token is optimized by the pre-trained model and can capture the overall semantic information of the archival text.

[0121] The feature vector of each archival text sample is:

[0122]

[0123] ​Among them, is the hidden state marked by [CLS], with a dimension of 768;

[0124] Construct a feature matrix:

[0125] Concatenate the hidden states of the [CLS] markers of all samples into a feature matrix X as the input feature of the classification model:

[0126]

[0127] Among them, is the feature matrix, is the number of archival data samples, is the th feature vector of the sample.

[0128] Step A3: Classification model design; the structure of the classification model includes the RoBERTa-wwm-ext model, a pooling layer, and a fully connected layer. The RoBERTa-wwm-ext model is used to generate the context representation of the archival text. Pool the output of RoBERTa-wwm-ext to extract global features. The fully connected layer maps the pooled features to the classification label space and outputs class probabilities;

[0129] Specifically, RoBERTa-wwm-ext outputs:

[0130] After the input archival text passes through the RoBERTa-wwm-ext model, the hidden states of all tokens are obtained:

[0131]

[0132] Among them, is the hidden state matrix of all tokens, is the number of tokens in the archival text, is the th hidden state of the token, with a dimension of 768;

[0133] Pooling layer output: To extract the global features of the entire archival text, use the max pooling method to perform a pooling operation on the hidden state matrix;

[0134] Take the maximum value by dimension:

[0135] For the hidden state matrix take the maximum value for each column (i.e., each dimension of the hidden state) to obtain the global feature vector :

[0136]

[0137] Among them, Denote the th token's th hidden state dimension.

[0138] Output the global feature vector:

[0139] The global feature vector after max pooling is a 768-dimensional vector that can capture the most significant feature information in the archive text.

[0140] Output of the fully connected layer: Input the global feature vector output by the pooling layer into the fully connected layer to further map it to the classification label space and complete the final classification task;

[0141] Linear transformation: Input the global feature vector into the fully connected layer for linear transformation:

[0142]

[0143] Where: is the weight matrix with dimension ( is the number of archive categories); is the bias vector with dimension C; is the classification score vector with dimension .

[0144] Softmax activation: Perform Softmax operation on the classification score vector to obtain the class probability distribution of the archive sample:

[0145]

[0146] Where, denotes the probability that the th sample belongs to the th class, denotes the classification score of the th sample in the th class, denotes the classification score of the th sample in the th class, denotes the number of classes.

[0147] Output the classification result:

[0148] Select the class with the highest probability as the final classification result:

[0149]

[0150] Step A4: Model Training. Divide the dataset into a training set, a validation set, and a test set. Use the training set to train the model. Calculate the difference between the model's predicted values and the true labels through the cross-entropy loss function. Use the Adam optimizer to update the model's parameters, and evaluate the model's performance on the validation set and test its accuracy on the test set.

[0151] Specifically, use the training set to perform model training on the model. Set hyperparameters such as the learning rate, batch size, and number of training epochs. Calculate the difference between the model's predicted values and the true labels through the cross-entropy loss function. Use the Adam optimizer to update the model's parameters through first-order moment estimation and second-order moment estimation, and evaluate the model's performance on the validation set and test its accuracy on the test set. The formula for the cross-entropy loss function is:

[0152]

[0153] where, is the cross-entropy loss, is the number of archive data samples, is the number of archive categories, is the true label, indicating whether the th sample belongs to the th category. When the sample belongs to category , is 1, otherwise it is 0. is the probability that the model predicts that the sample belongs to category ;

[0154] Step A5: Model Deployment. Deploy the trained model as a RESTful API for the archive scanning module to call, receive the archive text, and return the classification result.

[0155] Step A6: Continuous Learning. Set up a regular task to collect newly archived data, use the new data to update the model's parameters, and at the same time prevent forgetting of existing knowledge.

[0156] Step A61: Data Preparation and Preprocessing. Regularly collect newly archived data from the archive scanning module, clean and denoise the new data, assign category labels to the new data, and increase the diversity of the data through techniques such as synonym replacement and random deletion.

[0157] Step A62: Model Initialization. Use the previously trained model as the initial model. Freeze some layers of the model in the initial stage and only train the fully connected layers to prevent the model from overfitting on the new data.

[0158] Step A63: Loss function design. Use the initial model as the teacher model and the new model as the student model. Retain the existing knowledge through the distillation loss function. The distillation loss function is as follows:

[0159]

[0160] where is the distillation loss, is the number of archive data samples, is the number of archive categories, represents the probability that the teacher model believes the sample belongs to the category ; represents the probability that the student model believes the sample belongs to the category ;

[0161] Elastic Weight Consolidation (EWC) loss. Elastic Weight Consolidation is a continual learning technique that prevents the model from forgetting the knowledge of old tasks by penalizing changes in important weights. The EWC loss function is as follows:

[0162]

[0163] where is the penalty coefficient used to control the intensity of the EWC loss. The larger is, the stronger the model retains the old knowledge. is the importance of the weight, representing the importance of the -th weight for the old task, is the -th weight of the current model, is the -th weight of the initial model;

[0164] is determined by experimental adjustment. Set the initial value according to experience, verify and evaluate the performance of the model, and select the value that makes the model perform well on the new task and forgets the old task the least. is calculated through the Fisher information matrix. is a parameter that is continuously updated during the model training process. In each iteration, is updated through the Adam optimizer. Save the weights by training the model on the old task until the model converges.

[0165] Calculate the cross-entropy loss, distillation loss, and EWC loss. Combine the distillation loss, EWC loss, and cross-entropy loss to determine the total loss. Use the total loss function to train the student model:

[0166] ​

[0167] Among them, is the weight coefficient of the knowledge distillation loss, which is used to control the influence of knowledge distillation on the total loss. is the weight coefficient of the elastic weight consolidation loss, which is used to control the influence of EWC on the total loss. , , : They are the cross-entropy loss, the knowledge distillation loss, and the elastic weight consolidation loss respectively;

[0168] The larger , the more the model tends to imitate the behavior of the teacher model and retain the knowledge of the old task. The larger , the more the model tends to retain the key weights of the old task and prevent forgetting. and The values of are selected through grid search experiments to determine the combination that performs well on the new task and forgets the old task the least. , combination.

[0169] Step A64: Use the Adam optimizer to train the model, update the model parameters, record the accuracy and recall metrics of the model, and ensure that the model does not forget the old knowledge;

[0170] Step A65: Model deployment, deploy the optimized model as a RESTful API, receive the archive text and return the classification result, set a regular task to automatically collect new data and update the model.

[0171] Please refer to Figure 3 , in the archive number:

[0172] Step S1: Formulate the archive number structure: year - classification code - archive storage code - unique identification code;

[0173] Among them, the year is the time when the archive is scanned and stored by the archive scanning module, which is represented by 4 digits. For example, for the archive scanned and stored in 2024, the year number is represented as 2024;

[0174] The classification code is the label classified and matched by the archive classification model, which is represented by 2 - 3 letters. For example, the label of the financial archive in the archive classification model is fi. If there are many stored financial archives and refined classification can be carried out, after the archive classification model is optimized regularly, the labels after the refinement of the financial archive can be obtained. For example, the label of the financial voucher is fiv, the label of the financial ledger is fil, and the label of the financial report is fis;

[0175] The file storage code is the storage location code assigned to the physical file by the file storage repository, which is represented by 1 letter + 2 digits. For example, for a file stored in the 10th file storage cabinet in Area A, its file storage code is A10;

[0176] The unique identification code is the file sequence code within a certain file storage cabinet, which is represented by 3 digits. For example, for the 5th file on the 2nd layer in a certain file storage cabinet, its unique identification code is 205;

[0177] Step S2: File number generation;

[0178] Step S21: Access the retrieval record module to obtain the file scanning time;

[0179] Step S22: Classify the file through the file classification model to obtain the classification code;

[0180] Step S23: Access the file storage repository to obtain the storage location code assigned to the physical file and the file sequence code within this file storage cabinet;

[0181] Step S24: Combine the year, classification code, file storage code, and unique identifier to generate the final file number. For example, for a file scanned and stored in 2024, stored in the 5th file on the 2nd layer in the 10th file storage cabinet in Area A, and the file classification type is financial vouchers, then the file number can be 2024fivA10205;

[0182] Step S3: File number entry. Enter the file number into the electronic file text, and at the same time make an NFC tag for the physical file and write the file code into the NFC tag.

[0183] It should be noted that in all the examples shown and described here, any specific value should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.

[0184] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatuses, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0185] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present application.

Claims

1. A file management system, characterized in that: Including, archive registration terminal, archive storage repository and archive retrieval terminal; The file registration terminal is provided with: a file scanning module and a file encoding module; The archive storage library is an intelligent compact cabinet, which is equipped with: multiple archive storage cabinets, an archive cloud and an archive destruction module; The file access terminal is provided with: an electronic file access module, a physical file access module and a access record module; The file scanning module performs optical character recognition on the physical files to be stored, identifies the file contents and generates editable electronic files; The archive storage library quickly allocates storage locations for newly stored physical archives and prints storage barcodes; The file encoding module generates a file number for the scanned electronic file, enters the file number into the electronic file text, uploads the numbered electronic file to the file cloud, and simultaneously makes an NFC tag for the physical file and writes the file code into the NFC tag; The physical files with NFC tags and storage barcodes are received by the file storage cabinet, which is unlocked by identifying the NFC tags or storage barcodes to complete the physical file storage. The file storage cabinet is equipped with a file access control module, and the administrator accesses the files by operating the file access control module. The file destruction module is equipped with a shredder and a data erasure module, which is used to destroy physical files and electronic files that have expired and have no preservation value. The destruction process is automatically recorded and a destruction report is generated; The electronic archive access module provides online access services for electronic archives, the physical archive access module manages the access process of physical archives, and the access record module records the access and update information of all archives.

2. The file management system according to claim 1, characterized in that: The archive scanning module deploys an archive classification model based on RoBERTa-wwm-ext. The model is trained by using electronic archives with optical character recognition, and the model parameters are updated regularly with new archive data. The model is optimized through knowledge distillation and elastic weight solidification to obtain regularly updated archive classification.

3. The file management system according to claim 2, characterized in that: The archive coding module generates an archive number based on the archive classification, combined with the archive classification type, archive storage location, year and unique identifier.

4. The file management system according to claim 1, characterized in that: The archive retrieval module is tightly integrated with the archive cloud, and the administrator can quickly retrieve archives by keyword, archive number, classification type, and year; the archive retrieval module provides hierarchical authority management, and the administrator can set access rights for different users for the archive retrieval module to ensure that sensitive archives are only open to authorized personnel.

5. The file management system according to claim 1, characterized in that: The physical file access module manages the access process of physical files. The administrator submits a physical file access application through the file access module. The physical file access module locates the file storage location to notify the administrator and drives the file access control module of the file storage cabinet to open the file storage cabinet.

6. The file management system according to claim 1, characterized in that: The access record module records the access and update information of all files, generates access logs, records the file scanning time, accessor, access time, access purpose information and destruction records, and generates access data analysis reports to help administrators optimize the file management process.

7. The file management system according to claim 1, characterized in that: The archive management system also includes an environmental monitoring module, which monitors the environmental parameters of the archive room in real time to ensure the long-term preservation of the archives; the environmental monitoring module integrates: temperature and humidity monitoring and adjustment equipment, fire alarm equipment, air conditioning equipment, and light adjustment equipment; The temperature and humidity monitoring and adjustment device monitors the temperature and humidity of the archive storage cabinet in real time. When the temperature and humidity are lower than 14°C or higher than 24°C, and the humidity is lower than 43% and higher than 60%, the temperature and humidity of the monitored archive storage cabinet are adjusted; The fire alarm device monitors the smoke concentration of the file storage cabinet in real time, and when smoke is detected, it issues an alarm and performs fire extinguishing and smoke removal; The air conditioning equipment monitors the dust concentration and harmful gas concentration of the filing cabinet in real time. When the dust concentration exceeds 0.15mg / m³ and PM2.5 exceeds 75μg / m³, air purification is performed. When the concentration of sulfur dioxide and nitrogen oxides exceeds 0.01mg / m³, ventilation is performed. The light adjustment device adjusts the lights in the file storage cabinet and the room, and performs the functions of lighting and ultraviolet irradiation sterilization.

8. A file classification coding method, characterized in that: It includes two steps: file classification and file numbering; In the archives category: Step A1: Data preparation, collecting electronic files extracted by optical character recognition technology, performing data preprocessing, data cleaning and word segmentation; Step A2: Feature extraction. Convert the text data prepared in step 1 into computer-understandable numerical values, and extract feature vectors that can reflect the semantic information of different types of archival texts. Load the RoBERTa-wwm-ext pre-trained model and its corresponding tokenizer. Use the RoBERTa-wwm-ext tokenizer to encode the cleaned and tokenized archival texts, and convert each word (token) into the corresponding ID in the model vocabulary. For words out of vocabulary (Out-of-Vocabulary, OOV), use the special tag [UNK] to represent them. Input the encoded archival text into the RoBERTa-wwm-ext model to obtain the hidden states of each token. The output of the RoBERTa-wwm-ext model is a hidden state matrix: ; in, is the hidden state matrix of all tokens, is the number of tokens in the archive text, It is The hidden state of a token has a dimension of 768; From the hidden state matrix Extract the hidden state corresponding to the [CLS] tag , as the global semantic representation of the entire archival text, the hidden state of the [CLS] tag is optimized by the pre-training model and can capture the overall semantic information of the archival text; the feature vector of each archival text sample is: ; in, is the hidden state of the [CLS] tag, with a dimension of 768; Construct the feature matrix: Concatenate the hidden states of the [CLS] tags of all samples into a feature matrix X as the input features of the classification model: ; in, is the feature matrix, is the number of archival data samples, It is The feature vector of samples; Step A3: Classification model design; the structure of the classification model includes the RoBERTa-wwm-ext model, the pooling layer and the fully connected layer. The RoBERTa-wwm-ext model is used to generate the contextual representation of the archive text. The pooling layer pools the output of RoBERTa-wwm-ext to extract global features. The fully connected layer maps the pooled features to the classification label space and outputs the category probability. Step A4: Model training, divide the data set into training set, validation set and test set, use the training set to train the model, set the hyperparameters of learning rate, batch size and number of training rounds, calculate the difference between the model prediction value and the true label through the cross entropy loss function, use the Adam optimizer to update the model parameters through first-order moment estimation and second-order moment estimation, and evaluate the model performance on the validation set, and test its accuracy on the test set; the formula of the cross entropy loss function is: ; in, is the cross entropy loss, is the number of archival data samples, is the number of archive categories, is the true label, indicating Does the sample belong to categories, when the sample Belongs to category hour, is 1, otherwise it is 0. Predict samples for the model Belongs to category probability; Step A5: Model deployment: deploy the trained model as a RESTful API for the archive scanning module to call, receive archive text and return classification results; Step A6: Continuous learning, setting regular tasks, collecting newly stored archival data, using new data to update model parameters, and preventing the forgetting of existing knowledge; In the file number: Step S1: Formulate the file number structure: year - classification code - file storage code - unique identification code; The year is the time when the archive is scanned by the archive scanning module, expressed in 4 digits; The classification code is a code generated by the archive classification model, represented by 2-3 letters; The archive storage code is the storage location code assigned by the archive repository to the physical archive, represented by 1 letter + 2 digits; The unique identification code is the file sequence code in a certain file storage cabinet, represented by a three-digit number; Step S2: Generate file number; Step S21: access the access record module to obtain the file scanning time; Step S22: classify the archives using the archive classification model to obtain the classification code; Step S23: accessing the archive storage library to obtain the storage location code assigned to the physical archive and the archive sequence code in the archive storage cabinet; Step S24: Generate a final file number by combining the year, classification code file storage code and unique identifier; Step S3: Enter the file number. Enter the file number into the electronic file text. At the same time, make an NFC tag for the physical file and write the file code into the NFC tag.

9. The file classification coding method according to claim 8, characterized in that: In the archive classification, step A3 includes: RoBERTa-wwm-ext output: After the input archive text passes through the RoBERTa-wwm-ext model, the hidden states of all tokens are obtained: ; in, is the hidden state matrix of all tokens, is the number of tokens in the archive text, It is The hidden state of a token has a dimension of 768; Pooling layer output: In order to extract the global features of the entire archive text, the maximum pooling method is used to perform pooling operations on the hidden state matrix; Take the maximum value by dimension: for the hidden state matrix Take the maximum value of each column (that is, the dimension of each hidden state) to obtain the global feature vector : ; in, Indicates The first token hidden state dimensions; Output global feature vector: global feature vector after maximum pooling is a 768-dimensional vector that can capture the most significant feature information in the archive text; Fully connected layer output: The global feature vector output by the pooling layer Input to the fully connected layer and further mapped to the classification label space to complete the final classification task; Linear transformation: transform the global eigenvector Input the fully connected layer and perform linear transformation: ; in: is the weight matrix with dimension ( is the number of archive categories); is the bias vector, dimension is C; is the classification score vector with dimension ; Softmax activation: classification score vector Perform Softmax operation to obtain the category probability distribution of archive samples: ; in, Indicates The samples belong to The probability of the categories, Indicates The sample in The classification scores of the categories, Indicates The sample in The classification scores of the categories, Indicates the number of categories; Output classification results: select the category with the highest probability as the final classification result.

10. The file classification coding method according to claim 8, characterized in that: Step A6 in the archive classification includes the following steps: Step A61: Data preparation and preprocessing: regularly collect newly stored archive data from the archive scanning module, clean and denoise the new data, assign category labels to the new data, and increase data diversity through synonym replacement and random deletion techniques; Step A62: Model initialization, using the previously trained model as the initial model, freezing some layers of the model in the initial stage, and only training the fully connected layers to prevent the model from overfitting on new data; Step A63: Loss function design, using the initial model as the teacher model and the new model as the student model, retaining the existing knowledge through the distillation loss function, the distillation loss function is: ; in, is the distillation loss, is the number of archival data samples, is the number of archive categories, Indicates that the teacher model believes that the sample Belongs to category The probability of Indicates that the student model believes that the sample Belongs to category probability; Elastic weight solidification loss, elastic weight solidification is a continuous learning technique that prevents the model from forgetting the knowledge of old tasks when learning new tasks by penalizing changes in important weights. The elastic weight solidification loss function is: ; in, is the penalty coefficient, which is used to control the intensity of EWC loss. The larger it is, the stronger the model retains old knowledge. is the importance of weight, indicating the The importance of the weights to the old tasks, The current model weights, The initial model weights; Determine through experimental adjustment, set the initial value based on experience, verify the performance of the evaluation model, and select the model that performs well on new tasks and forgets the least about old tasks. value, Calculated by Fisher information matrix, It is a parameter that is continuously updated during model training. In each iteration, it is updated by the Adam optimizer , The weights are saved by training the model on the old task until the model converges; Calculate the cross entropy loss, distillation loss, and EWC loss, combine the distillation loss, EWC loss, and cross entropy loss to determine the total loss, and use the total loss function to train the student model: ; in, is the weight coefficient of knowledge distillation loss, which is used to control the impact of knowledge distillation on the total loss. is the weight coefficient of the elastic weight solidification loss, which is used to control the impact of EWC on the total loss. , , : They are cross entropy loss, knowledge distillation loss and elastic weight solidification loss respectively; The larger it is, the more the model tends to imitate the behavior of the teacher model and retain knowledge of old tasks. The larger it is, the more the model tends to retain the key weights of old tasks to prevent forgetting. and The value is selected through grid search experiments to select the best performing task on the new task and the least forgetting the old task. , Combination to determine; Step A64: Use the Adam optimizer to train the model, update the model parameters, and record the accuracy and recall indicators of the model to ensure that the model does not forget old knowledge; Step A65: Model deployment, deploy the optimized model as a RESTful API, receive archive text and return classification results, set up periodic tasks, automatically collect new data and update the model.

Citation Information

Patent Citations

  • Human body motion system data medical model construction method and system and application thereof

    CN115344702A

  • Intelligent tailoring system based on domain self-adaption

    CN118427300A

  • Archive management system based on cloud archive library

    CN118568327A

  • Electronic archive retrieval method and system based on large language model

    CN118643148A

  • Tobacco industry document automatic classification and storage method using deep learning

    CN118916484A