An intelligent archive management system based on multi-type analysis

By designing multi-type analysis modules and classification verification modules in the archive management system, optimizing document classification methods and verifying classification results, the problems of low classification efficiency and poor accuracy in the existing technology are solved, and efficient and accurate archive classification and management are achieved.

CN119646278BActive Publication Date: 2025-05-13ANHUI YIBAI INTERNET TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510174188.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-13
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The existing archive management system fails to deeply extract the classification characteristics and priorities of various archives during the document classification process, resulting in a long classification time, high system resource utilization, and lack of verification of classification results, which easily leads to classification errors.

Method used

An intelligent archive management system based on multi-type analysis is designed. Through the classification term priority determination module, professional term extraction module, target document classification module, classification verification module and document information change processing module, the document classification method is optimized, and the result verification is carried out after classification.

Benefits of technology

It improves the efficiency and accuracy of archive classification, reduces system resource usage, ensures the high-performance operation of the archive management system, and promptly discovers and corrects classification errors through the classification verification mechanism, improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119646278B_ABST
    Figure CN119646278B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of archive management, and specifically relates to an intelligent archive management system based on multi-type analysis. The classification term priority of each type of archive is determined based on the distribution of professional terms in classified documents in the multi-type archives. In the actual classification operation, the key terms in the document to be classified are extracted and matched with the predefined classification term priority to achieve efficient classification. Not only can the classification speed be improved, but also the occupancy rate of system resources can be effectively reduced to ensure the high-performance operation of the archive management system. At the same time, after the document classification is completed, a verification operation of the classification result is added. The classification accuracy is verified by comparing the similarity of the newly classified document with the classified document in the classified archive type to which it belongs. The error in the initial classification can be discovered and corrected in time, which is conducive to rapid reclassification and ensures that no significant impact is caused on subsequent document retrieval and access.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of archive management, and in particular relates to an intelligent archive management system based on multi-type analysis. Background Art

[0002] As laws and regulations continue to strengthen, companies must ensure that all their business activities strictly comply with national and local laws and regulations, which involves the proper preservation and timely updating of various types of documents such as contract documents, business documents and financial reports. Through effective archive management practices, companies can not only quickly meet the requirements of regulatory agencies, but also improve internal information retrieval efficiency, reduce duplication of work, and thus improve overall work efficiency.

[0003] In the archive management process, document classification is based on established classification standards and rules to systematically classify and label various types of documents. This is the basic link in the archive management system, and all subsequent management and operation activities are based on this. Given the complexity and diversity of multiple types of archives, such archives are particularly prone to classification errors, which affects the accuracy and efficiency of overall archive management.

[0004] Several technical solutions for archive classification management have been proposed in the prior art. For example, the Chinese invention patent with publication number CN113409020A proposes an electronic archive management system and method, including a file collection module, an archive management module, a report design module, an archive statistics module, an archive detection module, and an archive destruction module. The file collection module is responsible for connecting with the external business system and collecting multiple electronic files; the archive management module classifies and archives these files according to preset classification rules.

[0005] Another example is a Chinese invention patent with publication number CN113902421A, which proposes an archive classification management system. The system obtains the basic information of the archive through a scanning module and sends the basic information of the archive to the judgment module. The judgment module matches the basic information of the archive with each archive category, and then classifies the archive into the archive category with the highest text matching degree. If the self-judgment unit cannot judge the archive, the basic information of the archive is sent to the administrator terminal, and the archive is classified through human judgment.

[0006] The above two solutions mainly focus on the attribute information of the document itself during the document classification process, but fail to deeply refine the classification features and priorities of various archives. In particular, when classifying based on content terms, it is necessary to extract key terms from the documents to be classified and fully match them with the classification terms of various archives. Due to the lack of refinement of classification feature priorities, this method may lead to a long classification time and occupy a large amount of system resources, resulting in low classification efficiency and affecting user experience.

[0007] In addition, the above two solutions do not set up an effective classification result verification step after document classification. A single classification is prone to classification errors that cannot be discovered and corrected in time, which may lead to a series of problems such as data inconsistency, retrieval difficulties and inefficiency of subsequent operations. Summary of the invention

[0008] The purpose of the present invention is to improve the deficiencies in the prior art and to provide an intelligent archive management system based on multi-type analysis, which optimizes the document classification method when managing multi-type archives and adds a verification link for the classification results after classification to enhance the efficiency and accuracy of archive classification.

[0009] The purpose of the present invention can be achieved through the following technical solutions: an intelligent archive management system based on multi-type analysis, comprising the following modules: a classification term priority determination module, used to obtain the existing archive types in the enterprise document library and build a professional term library for each type of archive, thereby determining the classification term priority of each type of archive based on the classified documents corresponding to each type of archive.

[0010] The professional term extraction module is used to record the document to be classified as the target document and extract the professional terminology of the target document.

[0011] The target document classification module is used to match the extracted target document professional terms with the priority classification terms of various archives, thereby classifying the target document.

[0012] The classification verification module is used to compare the target document with the classified documents in the classification file type to perform classification verification.

[0013] The document information change processing module is used to configure the meta-information items of various types of archives, and to fill in the meta-information items after document classification, and then to monitor and process changes in the meta-information items of the classified documents within the document retention period, and to identify missing changes after the result of the change processing is that the change is allowed.

[0014] Combining all the above technical solutions, the present invention has the following positive effects: 1. The present invention determines the classification term priority of each type of archive based on the distribution of professional terms in classified documents in multiple types of archives. In the actual classification operation, the key terms in the documents to be classified are extracted and matched with the predefined classification term priority to achieve efficient classification. This classification method can not only greatly improve the classification speed, but also effectively reduce the occupancy rate of system resources, ensuring the high-performance operation of the archive management system.

[0015] 2. After the document classification is completed, the present invention adds a verification operation for the classification results, and verifies the classification accuracy by comparing the similarity between the newly classified documents and the classified documents in the classified archive type. This method reflects the pursuit of rigor and reliability of the initial classification results. Through this classification verification mechanism, errors in the initial classification can be discovered and corrected in a timely manner, which is conducive to rapid reclassification and ensures that there is no significant impact on subsequent document retrieval and access.

[0016] 3. The present invention executes the filling of meta-information items after the document classification is completed, and monitors and processes the changes of these meta-information items during the document preservation period, ensuring that the preservation status of the document always remains dynamic and updated in real time, thereby improving the long-term preservation value of the document. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The present invention is further described using the accompanying drawings, but the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative work.

[0018] Figure 1 It is a schematic diagram of system module connection of the present invention.

[0019] Figure 2 It is a schematic diagram of the target document classification process in the present invention.

[0020] Figure 3 This is a diagram of the target document meta-information item change identification and processing operation in the present invention. DETAILED DESCRIPTION

[0021] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0022] See also Figure 1 As shown, the present invention proposes an intelligent archive management system based on multi-type analysis, including a classification term priority determination module, a professional term extraction module, a target document classification module connected to the classification term priority determination module and the professional term extraction module respectively, a classification verification module connected to the target document classification module, and a document information change processing module connected to the classification verification module.

[0023] The classification term priority determination module is used to obtain the existing file types in the enterprise document library and build a professional term library for each type of file, thereby determining the classification term priority of each type of file based on the classified documents corresponding to each type of file.

[0024] In the example operation of the above module, the existing file types in the enterprise document library include document files, business files, accounting files, etc., among which document files include company policy documents, meeting minutes, employee handbooks, etc., business files include project documents, market research reports, etc., accounting files include financial statements, tax declaration materials, audit reports, etc.

[0025] Since each file type usually corresponds to a specific business field or function, and each field has unique expertise and technical terms, the expressions of different file types show significant differences. By building a professional terminology library for each type of file, we can fully understand and grasp the unique attributes and specific needs of each file type. In the process of building these professional terminology libraries, we can extract and organize professional terms in related fields by referring to existing academic papers, research reports and technical documents. Another example is that we can organize and build according to the standards and specifications widely recognized in the industry to ensure the consistency and universality of the terms.

[0026] As an example applied to the above scenario, the professional terminology library constructed by the document archives includes professional terms such as "agreement", "rules and regulations", "decision-making", etc.; the professional terminology library constructed by the business archives includes professional terms such as "project plan", "progress", "delivery", "risk", etc.; the professional terminology library constructed by the accounting archives includes professional terms such as "assets", "liabilities", "profits", "taxes", etc.

[0027] In the further operation of the above module, the classification term priority of each type of archive is determined according to the classified documents corresponding to each type of archive as follows: professional terms are extracted from the classified documents of each type of archive, and compared with the professional term library of the corresponding archive type, and consistent professional terms are selected as the confirmed professional terms of the document.

[0028] By comparing the professional terms extracted from the classified archives with the professional terminology database, it can be ensured that the extracted professional terms are consistent with the existing standard terms, avoiding confusion or errors caused by inconsistent use of terms in different documents. In addition, the professional terminology database is systematically organized and verified, and contains terms that are widely recognized and used in the industry. By comparing and screening out consistent professional terms, the possibility of misscreening of professional terms is reduced, ensuring the data-driven and objective nature of term selection, and avoiding the impact on the accuracy of subsequent classification.

[0029] The frequency of occurrence of professional terms confirmed in the classified documents in each type of archive is counted, and a priority list of classification terms for each type of archive is generated by arranging them from high to low in order of frequency.

[0030] In the above, the professional terms confirmed in the classified documents are arranged from high to low according to the frequency of occurrence to generate a classification term priority list. This is because frequently occurring terms are usually important concepts in the field, reflecting the focus and needs in actual work. The classification term priority arrangement can provide more effective features for the classification model, making it more efficient and accurate in the classification process.

[0031] As an example of the above implementation, assume that professional terminology extraction and priority determination of business archive types are performed. First, the professional terms extracted and confirmed from the classified business documents are "project plan", "risk", and "delivery", among which project plan appears 10 times, "delivery" appears 7 times, and "risk" appears 3 times. At this time, the generated classification term priority list is first level: project plan, second level: delivery, third level: risk.

[0032] The professional term extraction module is used to record the document to be classified as the target document and extract the professional terminology of the target document.

[0033] It is important to know that when extracting professional terms from classified documents, since it is impossible to determine in advance the specific professional field to which the document belongs, the text content of the document can be decomposed into basic language units (such as words or phrases), and then the frequency of occurrence of each word or phrase in the document is calculated to identify high-frequency words and phrases. Finally, based on the frequency statistics, those expressions that appear repeatedly and have field-specific meanings are selected as potential professional terms. In this way, by using natural language processing technology to automatically complete word segmentation and frequency statistics, the need for manual intervention is reduced. By identifying high-frequency words and phrases, key terms in the document can be quickly located, improving the efficiency of term extraction.

[0034] The target document classification module is used to match the extracted target document professional terms with the priority classification terms of various archives, thereby classifying the target document.

[0035] Applied to the above scheme, see Figure 2 As shown, the target document classification is implemented as follows: the number of professional terms extracted from the target document is counted. If there is only one professional term extracted, the professional term ranked first is first selected from the classification term priority list of each type of archive.

[0036] The professional terms extracted from the target document are matched with the professional terms ranked first in each category of archives. If the professional terms of the target document are successfully matched with the professional terms ranked first in a certain category of archives, the category of archives will be used as the classification result of the target document. If the professional terms of the target document do not match the professional terms ranked first in any category of archives, the professional terms ranked second in each category of archives will be selected in turn, and the professional terms of the target document will continue to be matched with the professional terms ranked second in each category of archives until a match is found or the end of the priority list is traversed.

[0037] Further applied to the above scheme, the target document classification continues to include the following process: if the number of professional terms extracted from the target document is more than one, each professional term is matched according to the classification term priority list of each type of archive to obtain the archive type that matches each professional term in the target document.

[0038] Identify whether the file type matched by each professional term in the target document is the same file type. If it is the same file type, then the file type is used as the classification result of the target document. If it is a different file type, then the file type matched by each professional term in the target document is recorded as a candidate file type, and the number of matched professional terms in the candidate file type is counted, and the matching priority number of each matching professional term in the candidate file type is recorded.

[0039] It should be understood that since the classification term priority list is arranged in descending order of the frequency of occurrence of the professional terms, the smaller the matching priority number of the professional term, the greater the matching degree.

[0040] Substitute the number of matching professional terms in each candidate file type and the matching priority number of each matching professional term into the expression Calculate the classification fitness of the candidate archive type , where Indicates the number of matching professional terms in the candidate file type. Indicates the number of professional terms extracted from the target document, Indicates the first matching candidate file type. The matching priority number corresponding to the professional term, Indicates the matching professional term number in the candidate file type. , Represents the total number of priorities in the taxonomy term priority list of candidate profile types.

[0041] In the above classification suitability formula, the more the number of matching professional terms in the candidate file type is, the smaller the matching priority number of the matching professional terms is, and the greater the classification suitability of the candidate file type is.

[0042] The classification fitness of each candidate file type is compared, and the candidate file type corresponding to the maximum classification fitness is selected as the classification result of the target document.

[0043] The present invention simplifies the classification process by identifying whether the file type matched by each professional term in the target document is the same file type when classifying based on the professional terms extracted from the target document and the classification term priority list of various types of files. If all terms match the same file type, the classification result can be directly determined. When there are multiple different file types, the classification fitness is calculated by counting the number of matched professional terms and the matching priority number, which can comprehensively consider the number and priority of terms to ensure the rationality of the classification result.

[0044] The present invention determines the classification term priority of each type of archive based on the distribution of professional terms in classified documents in multiple types of archives. In the actual classification operation, the key terms in the documents to be classified are extracted and matched with the predefined classification term priority to achieve efficient classification. This classification method can not only greatly improve the classification speed, but also effectively reduce the occupancy rate of system resources, ensuring the high-performance operation of the archive management system. In addition, the method reduces the delay caused by excessive resource occupation, guarantees the user's needs for fast and accurate retrieval and access to archives, thereby significantly improving the user's overall usage experience.

[0045] The classification verification module is used to compare the target document with the classified documents in the classification file type to which it belongs for classification verification. The specific process includes the following: selecting several comparison samples from the existing classified documents in the classification file type to which the target document belongs. The specific implementation process is: comparing the contents of the existing classified documents in the classification file type to which the target document belongs with the professional terminology library corresponding to the classification file type, and counting the proportion of professional terms that have been successfully compared as the professional coverage of the classified document.

[0046] It is important to understand that professional coverage reflects the professionalism and relevance of a document in a specific field, and helps to screen out representative documents as comparison samples.

[0047] In the access time window after the classified document is created, the access interval duration and number of visitors of the classified document are counted from the archive access information, and the use value is calculated using this information. The specific calculation formula is: , where Indicates the access interval duration of classified documents, specifically the average access interval duration. Indicates the length of the access time window after the classified document is created. Indicates the number of visitors to the classified documents. It indicates the total number of visitors to all classified documents in the classification file type to which the target document belongs within the access time window. represents a natural constant, The coefficient of dispersion that represents the access interval duration can be obtained by calculating the standard deviation of the access interval duration of the classified documents and dividing it by the average access interval duration. It reflects the degree of dispersion of multiple access interval durations of the classified documents within the access time window. The greater the degree of dispersion, the more unstable the impact of the access interval duration on the use value, and the smaller the influence, which makes the influence of the number of visitors on the use value greater. From this formula, it can be seen that the shorter the access interval duration of the classified documents and the more visitors there are, the greater the use value of the classified documents.

[0048] It is important to note that the reason for selecting the access interval duration and the number of visitors to analyze the usage value of classified documents is that the access interval duration and the number of visitors directly reflect the frequency of use and popularity of the document in actual work, which helps to screen out truly valuable comparison samples.

[0049] The representativeness index of the classified documents is obtained by taking the weighted average of the professional coverage and usage value of the classified documents.

[0050] The weight factors of professional coverage and use value in the above representative index calculation can be 0.6 and 0.4. This is because the purpose of representative index statistics is to screen comparison samples and target documents for classification verification. At this time, professional coverage is more critical for screening comparison samples, while use value is used as an auxiliary indicator to help improve the quality of sample use.

[0051] The comparison samples are selected based on the representativeness index of the classified documents.

[0052] It should be emphasized that when selecting comparison samples, you can select according to the number of comparison samples required. For example, if 10 comparison samples are needed, they can be arranged from large to small according to the representativeness index, and the top 10 classified documents can be selected as comparison samples. This ensures that the selected comparison samples have high quality and representativeness, and avoids low-quality documents affecting the results of classification verification.

[0053] Through the above steps and methods, representative comparison samples can be selected from the classification file type to which the target document belongs, ensuring the accuracy and comprehensiveness of the classification verification process. This method not only improves the reliability of classification verification, but also ensures the rationality of sample selection through scientific calculation methods.

[0054] Feature extraction is performed from the contents of the target document and the comparison sample respectively.

[0055] It should be supplemented that the features mentioned above can be text features, image features, numerical features, etc. This is because a document may include not only text content but also image content, where text features can be text terms, including but not limited to professional terms. Specifically, TF-IDF, Word2Vec, BERT and other methods can be used to extract text features, and image features can be extracted using color histograms, SIFT feature points and other methods. The detailed extraction method will not be repeated in this invention.

[0056] The similarity between the features extracted from the target document and the features extracted from the comparison samples is calculated to obtain the content similarity between the target document and each comparison sample.

[0057] It should be further added that when calculating content similarity, for text features, the similarity can be calculated by obtaining the cosine value of the angle between vectors; for numerical features, the Euclidean distance between feature vectors can be obtained to calculate the similarity; for image features, the structural similarity between two images can be measured to calculate the content similarity.

[0058] It is also worth pointing out that when calculating content similarity based on extracted features, when there is more than one extracted feature, the similarities calculated from different features need to be normalized and averaged to obtain the overall content similarity between the target document and the comparison sample. This is because different types of features (such as text, numerical, and image features) provide different information dimensions. Through normalization and average calculation, these different dimensions of information can be comprehensively considered to more comprehensively reflect the similarity of the documents. In addition, the similarity values ​​of different features may be in different scale ranges. Normalization can map all similarity values ​​to the same scale range, which is convenient for comparison and weighted averaging.

[0059] The storage formats of each comparison sample are obtained for comparison and classification, and the occurrence ratios of different storage formats are counted.

[0060] The storage format of the target document is matched with the storage format of the comparison sample in the classification file type to which it belongs. If the storage format of the target document is successfully matched, the occurrence ratio of the successfully matched storage format is used as the storage format commonality of the target document. If the match fails, the storage format commonality of the target document is 0.

[0061] The creation department is extracted from the meta-information items of each comparison sample, and compared and classified, and the occurrence ratio of different creation departments is counted.

[0062] Match the creation part of the target document with the creation part of the comparison sample in the classification file type to which it belongs. If the creation part of the target document is matched successfully, the occurrence ratio of the successfully matched creation part is used as the commonality of the creation part of the target document; if the match fails, the commonality of the creation part of the target document is 0.

[0063] Import the content similarity, storage format commonality and creation department commonality of the target document and each comparison sample into the evaluation formula , get the classification conformity of the target document , where Indicates the target document and Compare the content similarity of samples, Indicates the number of the comparison sample, , represents the number of comparison samples, , They represent the commonality of the target document's creation department and storage format, respectively. represents the penalty factor, in , Respectively represent the maximum similarity and minimum similarity between the target document and the comparison sample. represents a natural constant, represents the trade-off factor for storage format commonality, and .

[0064] The present invention conducts a comprehensive similarity evaluation of the target document and the comparison sample in terms of content, storage format and creation department, wherein the content similarity can reflect the semantic similarity of the document and ensure that the subject and core content of the document are consistent, which is crucial for classification verification because similar documents usually have similar content structures and subjects; storage format commonality refers to the similarity of the document in terms of file format (such as PDF, Word, Excel, etc.) and data structure. Documents of the same type usually use the same storage format and data structure, which helps the system manage and process documents; creation department commonality refers to the similarity of the department to which the document creator belongs. Documents created by the same department usually have similar business backgrounds and requirements, which facilitates classification and management. For example, documents generated by the financial department may involve budgets, audits, etc., and are more likely to belong to accounting archives. Content similarity, storage format commonality and creation department commonality are selected as the basis for classification compliance statistics because these parameters can comprehensively measure the similarity and correlation between documents from different dimensions, which helps to improve the comprehensiveness and accuracy of classification verification.

[0065] In addition, the weight factors for content similarity, creation department commonality, and storage format commonality in the above formula are set as , , The reason is that the content of the document is its core information carrier, which determines the subject and purpose of the document. Content similarity directly reflects the semantic consistency and subject relevance between documents, so it has the highest weight. The commonality of the creation department reflects the organizational background and collaborative relationship of the document, which helps to clarify the classification and permission management, so it has the second highest weight. Finally, changes in storage format usually do not significantly change the core content and purpose of the document. For example, the same report can exist in both PDF and Word formats, but this does not change its content and classification. Storage format commonality mainly affects the technical attributes and processing methods of the document, and has little impact on the content and business purpose, so it has the lowest weight.

[0066] Specifically, it is necessary to explain that when using content similarity, storage format commonality, and creation department commonality for classification compliance statistics, since the target document and each comparison sample have a content similarity, if only the average content similarity is used for classification compliance statistics, it is easy to ignore the volatility of content similarity. There may be significant differences in the content similarity between the target document and each comparison sample. In order to more reasonably and accurately reflect the impact of content similarity on classification compliance, the role of average content similarity can be adjusted by introducing a content similarity volatility penalty mechanism. Specifically, when the content similarity fluctuates greatly, its contribution to classification compliance is weakened by increasing the volatility penalty, thereby achieving more accurate classification verification.

[0067] The classification conformity of the target document is compared with a preset classification conformity threshold. Exemplarily, the classification conformity threshold is 0.8. If the classification conformity reaches the classification conformity threshold, the classification verification is successful, otherwise the classification verification fails.

[0068] In the innovative implementation of the above scheme, the following process is also included after the classification verification fails: judging whether there are candidate file types in the target document classification operation. If so, the candidate file types existing in the target document classification operation are arranged in order of classification suitability from large to small, and the second-ranked candidate file type is taken as the re-classification result after the target document classification verification fails.

[0069] After the target document is classified again, the classification verification continues. If the classification verification fails, the candidate file types ranked later in the candidate file type arrangement results are selected in turn for classification verification until the classification verification succeeds or all candidate file types are traversed.

[0070] If the classification verification still cannot be passed after traversing all candidate file types, the manual classification operation will be triggered and the final classification decision will be made by domain experts.

[0071] The document information change processing module is used to configure the meta-information items of various types of archives, and to fill in the meta-information items after document classification. It then monitors and processes changes in the meta-information items of the classified documents within the document retention period, and identifies missed changes when the result of the change processing is that the change is allowed.

[0072] In the above operation examples, meta-information items refer to structured information that describes the document itself and its content. This information helps users and systems better understand and manage documents, including but not limited to the creation department, creation date, author, version number, access rights, etc. Meta-information item information may change over time and as business needs change. Dynamically updating meta-information items can ensure that the data in the document management system is always kept up to date and accurate, thereby supporting more efficient document management and use.

[0073] Preferably, monitoring changes in the meta-information items of the classified documents within the document retention period is implemented as follows: setting an update cycle for the classified documents within the document retention period.

[0074] Specifically, the update cycle (such as monthly, quarterly, etc.) can be set according to the importance of the document.

[0075] See also Figure 3 As shown, each meta-information item of the classified document is taken as a separate block, and the initial hash value of the block corresponding to each meta-information item is calculated after the meta-information item is initially filled in.

[0076] During the update cycle of the classified document, the hash value of each block corresponding to the meta-information item is calculated and matched with the initial hash value of the block corresponding to each meta-information item. If the hash value of a block does not match, it is recognized that the meta-information item of the classified document has changed.

[0077] Further preferably, the change processing is performed as follows: the operator is extracted from the change operation log, and it is determined whether the changer is an authorized person. If the operator is an authorized person, the change of the classified document meta-information item is allowed; if the operator is not an authorized person, a locking operation is triggered to lock the changed meta-information item of the classified document.

[0078] This method not only improves the security of the document management system, but also ensures the transparency and traceability of data through scientific management and maintenance mechanisms. The entire process can be automated, reducing manual intervention and improving work efficiency. At the same time, when unauthorized changes are detected, the locking mechanism can be triggered in time to ensure the accuracy and security of the final information.

[0079] Further preferably, after the result of the change processing is that the change is allowed, the change omission identification includes the following process: extracting and recording the change items from the classified documents of various types of archives that allow changes.

[0080] The change items corresponding to the documents classified as changes allowed by each type of archive are classified and counted, and then the change items with the highest occurrence ratio are taken as the tendency change items.

[0081] The content similarity between the classified documents that are allowed to change and other classified documents that have not changed corresponding to the tendency change items is calculated, thereby screening out the classified documents that have not changed and that meet the similarity threshold.

[0082] The tendency change items of the filtered classified documents that have not changed are actively pushed to the operator to prompt the operator for changes and omissions.

[0083] It is important to understand that when operators make changes to the metadata items of a document, changes may be missed due to human error, system design defects, etc. This is reflected in the fact that it is easy to make mistakes or make mistakes when manually updating metadata items. For example, operators may forget to update all related documents, or become distracted when multitasking, or operators may not follow standardized operating procedures and may ignore certain steps or details, resulting in incomplete updates of metadata items.

[0084] In the specific example understood above, assuming that the name of a project has changed, when the operator updates the "project name" meta-information item of the project report document, he may ignore the corresponding updates of other documents related to the project (such as meeting minutes, budget plans, etc.).

[0085] Through the above steps and methods, the change items in the classified documents that are allowed to change can be systematically extracted in the document management process, and the tendency change items can be classified and counted. The content similarity between the allowed change documents corresponding to these tendency change items and other unchanged documents can be calculated, and the unchanged documents that meet the similarity threshold can be screened out. The tendency change items in these documents can be actively pushed to the operator to prompt potential omissions of changes, and the operator can be reminded to make corrections in time to ensure the accuracy and security of the final information.

[0086] The present invention executes the filling of meta-information items after the document classification is completed, and monitors and processes the changes of these meta-information items during the document preservation period, ensuring that the preservation status of the document always remains dynamic and updated in real time, thereby improving the long-term preservation value of the document. In addition, users can quickly locate the required documents through more precise search conditions, improving the accuracy and speed of retrieval. More importantly, it records in detail the various status changes of the document throughout its life cycle, which is convenient for comprehensive tracking and management of the development process of the document, and significantly improves the overall efficiency and accuracy of document management.

[0087] The above contents are merely examples and explanations of the structure of the present invention. The technicians in this technical field may make various modifications or additions to the specific embodiments described or replace them in a similar manner. As long as they do not deviate from the structure of the invention or exceed the scope defined by the present invention, they should all fall within the protection scope of the present invention.

Claims

1. An intelligent archive management system based on multi-type analysis, characterized in that: Includes the following modules: The classification term priority determination module is used to obtain the existing file types in the enterprise document library and build a professional term library for each type of file, thereby determining the classification term priority of each type of file based on the classified documents corresponding to each type of file; A professional term extraction module is used to record the document to be classified as a target document and extract professional terms from the target document; A target document classification module is used to match the extracted target document professional terms with the priority classification terms of various archives, thereby classifying the target document; A classification verification module is used to compare the target document with the classified documents in the classification file type to which it belongs for classification verification; The document information change processing module is used to configure the meta information items of various archives, and fill in the meta information items after the document classification, and then monitor and process the changes of the meta information items of the classified documents within the document retention period, and identify the omission of changes after the result of the change processing is that the change is allowed; The classification verification refers to the following process: Selecting a number of comparison samples from existing classified documents of the classification file type to which the target document belongs; Extract features from the target document and the comparison sample content to calculate the similarity to obtain the content similarity between the target document and each comparison sample; Obtain and compare the storage formats of the comparison samples, and match the storage format of the target document. If the match is successful, the occurrence ratio of the storage format is used as the storage format commonality; Extract and compare the creation departments of each comparison sample, and match the creation department of the target document. If the match is successful, the occurrence ratio of the creation department is used as the creation department commonality; Through evaluation , get the classification conformity of the target document , where Indicates the target document and Compare the content similarity of samples, Indicates the number of the comparison sample, , represents the number of comparison samples, , They represent the commonality of the target document's creation department and storage format, respectively. represents the penalty factor, in , Respectively represent the maximum similarity and minimum similarity between the target document and the comparison sample. represents a natural constant, represents the trade-off factor for storage format commonality, and ; The classification conformity of the target document is compared with the preset classification conformity threshold. If the classification conformity reaches the classification conformity threshold, the classification verification is successful, otherwise the classification verification fails.

2. The intelligent archive management system based on multi-type analysis as claimed in claim 1, characterized in that: The determination of the classification term priority of each type of archive is as follows: Extract professional terms from the classified documents of each type of archive, compare them with the professional terminology library of the corresponding archive type, and select consistent professional terms as the confirmed professional terms of the document; The frequency of occurrence of professional terms confirmed in the classified documents in each type of archive is counted, and a priority list of classification terms for each type of archive is generated by arranging them from high to low in order of frequency.

3. The intelligent archive management system based on multi-type analysis as claimed in claim 2, characterized in that: The target document classification is implemented as follows: Count the number of professional terms extracted from the target document. If only a single professional term is extracted, match the extracted professional term with the professional term ranked first in the priority list of each type of archive classification term. If the match is successful, determine the classification result. If there is no match, select the professional term ranked second in each type of archive in turn for matching until a match is found or traverse to the end of the priority list.

4. The intelligent archive management system based on multi-type analysis as claimed in claim 3, characterized in that: The target document classification continues to include the following process: When multiple professional terms are extracted from the target document, each extracted professional term is matched hierarchically according to the classification term priority list of each type of archive to obtain the archive type corresponding to each term; If all professional terms match the same file type, the file type is used as the classification result of the target document. If different professional terms match different file types, these matched file types are recorded as candidate file types. Count the number of matching professional terms in each candidate file type, and record the matching priority number of each matching professional term; Using expressions Calculate the classification fitness of the candidate archive type , where Indicates the number of matching professional terms in the candidate file type. Indicates the number of professional terms extracted from the target document, Indicates the first matching candidate file type. The matching priority number corresponding to the professional term, Indicates the matching professional term number in the candidate file type. , The total number of priorities in the priority list of classification terms representing the candidate archive type; The classification fitness of each candidate file type is compared, and the candidate file type corresponding to the maximum classification fitness is selected as the classification result of the target document.

5. The intelligent archive management system based on multi-type analysis as claimed in claim 1, characterized in that: The process of selecting several comparison samples is as follows: Compare the existing classified document contents in the classified file type to which the target document belongs with the professional terminology library corresponding to the classified file type, and count the percentage of professional terms that have been successfully compared as the professional coverage of the classified document; The access interval duration and the number of visitors of the classified documents are counted from the archive access information within the access time window after the classified documents are created to calculate the usage value; The representativeness index of the classified documents is obtained by calculating the weighted average of the professional coverage and use value of the classified documents; The comparison samples are selected based on the representativeness index of the classified documents.

6. The intelligent archive management system based on multi-type analysis as claimed in claim 1, characterized in that: After the classification verification fails, the following process is also included: Determine whether there is a candidate file type in the target document classification operation. If so, sort the candidate file types in descending order according to the classification suitability, and select the next candidate file type for reclassification verification; If the classification verification fails again, the subsequent candidate file types will be selected for classification verification in turn until the classification verification succeeds or all candidate file types are traversed; If the classification verification still fails after traversing all candidate file types, the manual classification operation will be triggered.

7. The intelligent archive management system based on multi-type analysis as claimed in claim 1, characterized in that: The change monitoring of the meta information items of the classified documents within the document retention period is implemented as follows: Set an update cycle for classified documents within the document retention period; Treat each meta-information item of the classified document as a separate block, and calculate the initial hash value of each block corresponding to the meta-information item after the meta-information item is initially filled in; During the update cycle of the classified document, the hash value of the block corresponding to each meta-information item is calculated and matched with the initial hash value. If the hash value of a meta-information item does not match, it is recognized that the meta-information item has changed.

8. The intelligent archive management system based on multi-type analysis as claimed in claim 1, characterized in that: The change process is as follows: After detecting a change in the metadata item, the operator is extracted from the change operation log and the authority is verified. If the operator is an authorized person, the change is allowed. If the operator is not an authorized person, a lock operation is triggered.

9. The intelligent archive management system based on multi-type analysis as claimed in claim 8, characterized in that: The change omission identification is referred to the following process: Extract and record changes from classified documents that are allowed to change in various archives; The change items corresponding to the classified documents that are allowed to change in each type of archive are classified and counted, and then the change items with the highest occurrence ratio are taken as the tendency change items; Calculate the content similarity between the classification documents that are allowed to change and other classification documents that have not changed, thereby screening out the classification documents that have not changed and have reached the similarity threshold; The tendency change items of the filtered classified documents that have not changed are actively pushed to the operator to prompt the operator for changes and omissions.

Citation Information

Patent Citations

  • Electronic archive management system and method

    CN113409020A

  • Archive classification management system

    CN113902421A

  • Problem classification method and apparatus, computer device and storage medium

    CN108509482A

  • Electronic document version management method and device based on block chain

    CN112835612A