A method for recording an archive directory, a program product, a device and a storage medium

By obtaining archive images, recorded data and directory templates, using preset directory templates to fill in recorded data, and attaching basic directory and archive images, the challenge of automated directory recording is solved, and efficient and accurate directory data generation is achieved.

CN119557495BActive Publication Date: 2025-05-06BEIJING STARSHINE DIGITAL SYST CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510088602.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The automated recording of archive catalogs has challenges in the field of information management technology. The existing technology relies on manual operations, is inefficient and prone to errors.

Method used

By obtaining archive images, record data and directory templates, using preset directory templates to fill in recorded data, attaching basic directory and archive images, and filling directory data based on identification information, generating titles, topic words and classification numbers to realize automated recording of archive directory.

Benefits of technology

Automatic recording of archive catalogs has been realized, the dependence on manpower has been reduced, efficiency and accuracy have been improved, and the catalog data has been ensured that the catalog data complies with relevant standards and regulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557495B_ABST
    Figure CN119557495B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a method, program product, device and storage medium for recording an archive directory, the method comprising: obtaining an archive image, recording data and a directory template of a target archive, the directory template comprising a plurality of fields to be filled; filling the recording data into a first field corresponding to the directory template to obtain a basic directory of the target archive; attaching the basic directory to the archive image, and filling the second field of the basic directory based on the identification information of the archive image; generating one or more of the title, subject word and classification number of the target archive, and filling them into a third field corresponding to the basic directory to obtain the directory data of the target archive. By generating the data required for directory recording from multiple aspects, and the obtained directory data conforming to relevant regulations and standards in terms of format and content, the automatic recording of the archive directory is realized, reducing the degree of dependence on manpower.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information management technology, and in particular to a method, program product, device and storage medium for recording an archive directory. Background Art

[0002] The cataloging of archives is an important concept in archives management and information organization, which aims to ensure that archives can be preserved for a long time and can be effectively retrieved and used in the future. Unlike ordinary documents or files, the catalog content of archives is more detailed, usually including more metadata, and pays special attention to the formation background, source and related information of the archives to support research and historical research. At the same time, the cataloging of archives usually follows international or national standards, such as GB / T 20163 "China Archives Machine Readable Catalog Format", "International Standard Archives Cataloging Rules (General Principles)", etc. Therefore, the cataloging of archives is a complex and rigorous process. How to realize the automated cataloging of archives is a technical problem that needs to be solved in this field. Summary of the invention

[0003] The purpose of the embodiments of the present application is to provide a method, program product, device and storage medium for recording an archive directory, so as to achieve the technical effect of automatic recording of the archive directory.

[0004] A first aspect of an embodiment of the present application provides a method for recording an archive directory, the method comprising:

[0005] Acquire an archive image, bibliographic data, and a catalog template of a target archive, wherein the catalog template includes a plurality of fields to be filled;

[0006] Fill the bibliographic data into the first field corresponding to the directory template to obtain the basic directory of the target archive;

[0007] Mounting the basic directory and the archive image, and filling the second field of the basic directory based on the identification information of the archive image;

[0008] Generate one or more of the title, subject word and classification number of the target file, and fill them into the third field corresponding to the basic directory to obtain the directory data of the target file.

[0009] In the above implementation process, by presetting a catalog template including multiple fields to be filled, for each target archive, the catalog data carried by it is filled into the first field, and by attaching the basic catalog and the archive image, the second field can be filled based on the identifier of the archive image, and the title, subject word and classification number of the target archive are generated and filled into the third field, thereby completing the filling of the catalog template. The data required for catalog recording is generated from multiple aspects, and the obtained catalog data conforms to relevant regulations and standards in terms of format and content, realizing the automatic recording of the archive catalog and reducing the degree of dependence on manpower.

[0010] Furthermore, before filling the bibliographic data into the first field corresponding to the catalog template, the method further includes:

[0011] Determine whether the catalog data passes data integrity verification; the data integrity verification includes one or more of total number of pieces consistency detection, total number of bytes consistency detection, metadata item integrity detection, metadata mandatory catalog item detection, process information integrity detection, continuity metadata item detection, content data integrity detection, attachment data integrity detection and information package content data integrity detection.

[0012] In the above implementation process, by verifying the data integrity of the acquired description data, the integrity of the description data can be guaranteed, ensuring that the subsequent directory description is based on the complete and correct description data.

[0013] Further, the filling of the second field of the basic directory based on the identification information of the archive image includes:

[0014] Acquire identification information of the archive image, the identification information including one or more of the start and end numbers of the files in the volume of the target archive, the start image number of the archive image and the end image number;

[0015] A second field of the base directory is populated based on the identification information, and the second field is used to index to the target archive.

[0016] In the above implementation process, the second field is filled with the start and end numbers of the files in the volume of the target archive, and the starting image number and ending image number of the archive image, so that the filled second field can be indexed to the target archive, establishing a one-to-one correspondence between the directory data and the target archive.

[0017] Furthermore, the generating of one or more of the title, subject word and classification number of the target archive includes:

[0018] Generating a title of the target archive according to the subject information identified from the archive image;

[0019] Based on a pre-established subject word database, determining the subject word of the target file from the title of the target file;

[0020] The classification number of the target file is generated based on a pre-established category word database.

[0021] In the above implementation process, the nomination of the target archive is generated by identifying the subject information from the archive image, the subject word of the target archive is determined from the subject word database and the title, and the classification number of the target archive is generated based on the category word database. The target archive nomination, subject word and classification number are automatically recorded, and the degree of automation of catalog recording is improved.

[0022] Furthermore, the method further comprises:

[0023] Performing verification processing on the data to be verified, the data to be verified includes one or more of the basic directory, the filled second field, and the review mark carried by the archive image;

[0024] The archive image is identified and analyzed, and if the identification and analysis result includes the target data corresponding to the field to be filled, the target data is filled into the fourth field corresponding to the basic directory.

[0025] In the above implementation process, by performing recognition and analysis on the archive image, the target data is extracted from the recognition and analysis results to fill in the fourth field of the basic directory, further improving the directory data of the target archive. At the same time, by checking the data to be checked, it is ensured that the filled fields in the basic directory and the review marks carried by the archive image are accurate, thereby improving the accuracy of the recording.

[0026] Furthermore, the method further comprises:

[0027] Obtain the computing performance of multiple positions to be assigned tasks;

[0028] Determine a first camera position and a second camera position from the plurality of camera positions based on the computing performance, wherein the computing performance of the first camera position is higher than the computing performance of the second camera position;

[0029] The task of generating the title, subject words and classification number of the target file is assigned to the first camera, and the task of checking the data to be checked is assigned to the second camera.

[0030] In the above implementation process, on the one hand, the efficiency of catalog recording can be improved, and on the other hand, different computing positions can be jointly allocated to perform tasks with different computing power requirements, thereby improving the rationality of position scheduling.

[0031] Furthermore, the method further comprises:

[0032] Performing multi-level quality checks on the catalog data of the plurality of target archives;

[0033] Modify catalog data that fails quality inspection;

[0034] Submit catalog data that passes quality check.

[0035] In the above implementation process, multi-level quality checks are performed on the directory data of all target archives, and the directory data obtained by automatic recording is strictly checked from multiple aspects and dimensions to ensure the accuracy of the directory data and improve the reliability of the directory data obtained by automatic recording.

[0036] A second aspect of an embodiment of the present application provides a computer program product, wherein the computer program product includes a computer program, and when the computer program is executed by a processor, any method described in the first aspect is implemented.

[0037] According to a third aspect of the present application, an electronic device is provided, the electronic device comprising:

[0038] processor;

[0039] a memory for storing processor-executable instructions;

[0040] Wherein, when the processor calls the executable instruction, it implements the operation of any method described in the first aspect.

[0041] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of any method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0043] Figure 1 A schematic diagram of a flow chart of a method for recording an archive directory provided in an embodiment of the present application;

[0044] Figure 2-Figure 3 A schematic diagram of a flow chart of another method for recording an archive directory provided in an embodiment of the present application;

[0045] Figure 4 A hardware structure diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0046] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0047] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0048] Archives management emphasizes the historical value of archives and their function as evidence, and needs to ensure their authenticity, integrity and reliability. At the same time, in order to ensure interoperability across institutions or systems, the format and content of archives need to be highly consistent, and international or national standards must be followed. Moreover, once the catalog of archives is completed, it will not be easily changed unless errors are found or new information needs to be supplemented. Therefore, archive catalog recording is a complex process that requires strict compliance with relevant standards. Compared with ordinary document files, archive catalogs record more data and more complex processes. Therefore, in related technologies, archive catalog recording still has a high dependence on manpower. How to improve the degree of automation of archive catalog recording has become a technical problem that needs to be solved in this field.

[0049] To this end, this application provides a method for recording an archive directory, which can be applied to an archive management system. Figure 1 As shown, the recording method includes steps 110 to 140.

[0050] Step 110: Acquire the archive image, bibliographic data, and catalog template of the target archive, wherein the catalog template includes a plurality of fields to be filled.

[0051] Exemplarily, the target archive refers to any archive of the directory to be generated. The target archive can be stored in the form of an image, and the archive content is specifically presented in the form of an image. An image that records the archive content is called an archive image. The target archive can carry bibliographic data. For example, before executing the bibliographic method of this embodiment, part of the bibliographic data of the target archive is manually or automatically extracted, but the bibliographic data may be incomplete or the format does not meet the standard.

[0052] In order to generate directory data that complies with regulations or standards, a directory template can be pre-customized. The directory template includes multiple fields to be filled, and each field is used to record different bibliographic information of the target file. The fields may include but are not limited to title fields, responsible person fields, classification number fields, file number fields, subject word fields, document types fields, confidentiality levels fields, storage period fields, and time fields, etc. The fields to be filled in the directory template can be determined according to actual regulations or standards, and this application does not limit this.

[0053] Step 120: Fill the catalog data into the first field corresponding to the directory template to obtain the basic directory of the target archive.

[0054] Among them, different target archives may carry the same type or different types of catalog data. For example, the catalog data carried by target archive A includes responsible person data and confidentiality data, while the catalog data carried by target archive B includes confidentiality data and time data. It can be seen that the catalog data carried by different target archives are uneven. In step 120, the existing catalog data is filled into the first field corresponding to the directory template. The first field refers to the field corresponding to the catalog data in the directory template, or in other words, the first field refers to the field filled with catalog data in the directory template, and the first field does not specifically refer to a certain field. For example, in the above example, for target archive A, the first field includes the responsible person field and the confidentiality field; and for target archive B, the first field includes the confidentiality field and the time field. Filling the previously generated catalog data into the directory template can obtain the basic directory of the target archive. The basic directory is an incomplete directory, including some filled fields and some fields to be filled.

[0055] Step 130: Mount the basic directory and the archive image, and fill the second field of the basic directory based on the identification information of the archive image.

[0056] The so-called linking of the basic directory and the archive image refers to establishing a correspondence between the basic directory and the archive image. If the target archive includes one archive image, the basic directory of the target archive and the archive image are in a one-to-one correspondence; if the target archive includes multiple archive images, the basic directory of the target archive and the archive image are in a one-to-many correspondence.

[0057] In addition, since the accuracy of the attachment of the basic directory and the archival image will affect the accuracy of the subsequent filling of the second field, optionally, after the attachment is completed and before the second field is filled, the method may also include: checking the number of archival images attached to the basic directory. If the number of frames fails to be checked, return to reattach the basic directory and the archival image until the attachment is completely correct. If the number of frames passes, continue to perform subsequent steps. By checking the number of attached archival images, it is possible to avoid wrong attachment of archival images of other archives or missing attachment of archival images of target archives, thereby ensuring that 100% of the archival images and the basic directory are attached correctly.

[0058] After the attachment is completed, the second field of the basic directory can be filled based on the identification information of the archive image. Each archive image has its own identification information. The second field of the basic directory is filled based on the identification information of the archive image, so that the target archive can be indexed using the second field, and the archive image included in the target archive can be read. At the same time, the filled second field can also be used for subsequent verification. The second field refers to the second field corresponding to the identification information of the archive image in the basic directory, or the second field refers to the field in the basic directory filled with the identification information of the archive image.

[0059] Step 140: Generate one or more of the title, subject terms and classification number of the target file, and fill them into the third field corresponding to the basic directory to obtain the directory data of the target file.

[0060] Title, subject word and classification number are important archival catalog information, so in step 140, one or more of the title, subject word and classification number of the target archive is generated, and the generated result is filled into the third field corresponding to the basic directory. The third field may include one or more of the title field, subject word field and classification number field. After completing the filling of the first field to the third field, the directory data of the target archive can be obtained.

[0061] It can be seen that this application presets a catalog template including multiple fields to be filled, and for each target archive, fills the catalog data it carries into the first field, and by connecting the basic catalog and the archive image, fills the second field based on the identifier of the archive image, and generates the title, subject word and classification number of the target archive and fills them into the third field, thereby completing the filling of the catalog template. The data required for catalog recording is generated from multiple aspects, and the obtained catalog data complies with relevant regulations and standards in terms of format and content, realizing the automatic recording of the archive catalog and reducing the degree of dependence on manpower.

[0062] Steps 110 to 140 are described in detail below.

[0063] According to some embodiments of the present application, after obtaining the bibliographic data of the target archive in step 110 and before executing step 120 to fill the bibliographic data into the first field, the method further includes the steps of:

[0064] Determine that the bibliographic data passes data integrity verification.

[0065] The data integrity verification includes one or more of the following: total number of items consistency detection, total number of bytes consistency detection, metadata item integrity detection, metadata mandatory entry item detection, process information integrity detection, continuity metadata item detection, content data integrity detection, attachment data integrity detection, and information package content data integrity detection. The specific process of data integrity verification can be found in the relevant technology, and this application will not be expanded here.

[0066] If the description data fails the data integrity check, the process returns to step 110 to re-download the archive image and description data of the target archive until the description data passes the data integrity check. If the description data passes the data integrity check, the subsequent steps are continued.

[0067] It can be seen that this embodiment can ensure the integrity of the recorded data by performing data integrity verification on the acquired recorded data, and ensure that the subsequent directory recording is based on the complete and correct recorded data.

[0068] According to some embodiments of the present application, filling the second field based on the identification information of the archival image in step 130 may specifically include steps 131 and 132.

[0069] Step 131: Acquire identification information of the archive image, wherein the identification information includes one or more of the start and end numbers of the files in the volume of the target archive, the start image number and the end image number of the archive image.

[0070] Step 132: Fill a second field of the base directory based on the identification information, where the second field is used to index to the target archive.

[0071] Exemplarily, the start and end numbers of the files in the volume refer to the continuous number range assigned to each file after it is arranged in a certain order in a case file (i.e., a collection of related files). The identification information is filled into the second field of the basic directory, so that the target file can be indexed according to the second field, for example, the file image of the target file can be indexed. In addition, the identification information recorded in the second field can also be used to verify whether the target file has missing pages and the file image is discontinuous.

[0072] It can be seen that this embodiment uses the start and end numbers of the files in the volume of the target archive, the starting image number and the ending image number of the archive image to fill the second field, so that the filled second field can be indexed to the target archive, establishing a one-to-one correspondence between the directory data and the target archive, and can be used for subsequent archive verification.

[0073] According to some embodiments of the present application, the process of generating the title of the target archive in step 140 may include step 141 .

[0074] Step 141: Generate a title for the target archive based on the subject information identified from the archive image.

[0075] Exemplarily, the subject information of the target archive is first identified from the archive image. Then, the title of the target archive is determined based on the subject information. For example, the subject information can be identified from the archive image using OCR (Optical Character Recognition) technology. For the specific identification process, please refer to the relevant technology, which will not be elaborated in this application.

[0076] In addition, the process of generating the subject words of the target archive in step 140 may include step 142 .

[0077] Step 142: Based on a pre-established subject word database, determine the subject word of the target file from the title of the target file.

[0078] Exemplarily, a subject word database may be established in advance according to relevant regulations or standards, and the subject word database records a plurality of subject words. When executing step 142, one or more candidate subject words may be determined from the title of the target file, and the candidate subject words recorded in the subject word database may be determined as the subject words of the target file. The subject words of the target file may include one or more.

[0079] In addition, the process of generating the classification number of the target file in step 140 may include step 143.

[0080] Step 143: Generate a classification code for the target file based on a pre-established category word database.

[0081] Exemplarily, a category word database can be established in advance according to relevant regulations or standards, and the category word database includes professional terms that have been strictly screened and defined to describe the content characteristics of the archive. When executing step 143, the target category word used to describe the content characteristics of the target archive can be matched from the category word database first. Then, based on the matched target category word, refer to the corresponding classification system to determine the classification number of the target archive in the corresponding classification system. For example, for the target archive about "ancient Chinese architectural art", the target category words of the target archive that can be matched from the category word database include "ancient architecture", "architectural history" and "architectural style". Subsequently, based on the target category word, refer to the CLC (Chinese Library Classification) classification method, it can be determined that the classification number of the target archive in the CLC system is TU-73 (Architectural Science-Architectural History).

[0082] It can be seen that this embodiment generates a nomination for a target file by identifying subject information from an archive image, determines the subject terms of the target file from a subject term database and a title, and generates a classification number for the target file based on a category term database, thereby completing the automatic recording of target file nominations, subject terms and classification numbers, and improving the degree of automation of catalog recording.

[0083] According to some embodiments of the present application, the recording method may also include: Figure 2 Steps 210 to 220 are shown.

[0084] Step 210: Verify the data to be verified, where the data to be verified includes one or more of the basic directory, the filled second field, and the review mark carried by the archive image.

[0085] Exemplarily, the data to be checked includes a plurality of filled fields in the basic directory, such as a filled first field and a filled second field, so as to correct the fields filled incorrectly.

[0086] In addition, the archive image may carry review marks, which include indexing and / or annotation. The indexing refers to the content in the archive image that is highlighted in the form of modifying the background color, font, font color, highlighting, underlining, etc. The annotation refers to the information added to the archive content for description, explanation, or reminder. If the archive image carries review marks, then the review marks can also be checked in step 210.

[0087] Step 220: Perform recognition and analysis processing on the archive image. If the recognition and analysis result includes the target data corresponding to the field to be filled, fill the target data into the fourth field corresponding to the basic directory.

[0088] Exemplarily, the archival content recorded in the archival image may carry the data required for the catalog entry. Therefore, the archival image can also be identified and analyzed, for example, the archival image can be identified and analyzed using OCR technology to obtain the identification and analysis results. If the identification and analysis results contain the data required for the catalog entry, that is, the target data corresponding to a field to be filled, then the target data is filled into the fourth field corresponding to the basic directory. Among them, the fourth field refers to the field in the basic directory corresponding to the target data obtained after the archival image is identified and analyzed, or in other words, the fourth field refers to the field in the basic directory filled with target data, and the fourth field does not specifically refer to a certain field. For example, if the time information of the target archive is identified and analyzed from the archival image, then the fourth field is the time field. For another example, if the person responsible for the target archive is identified and analyzed from the archival image, then the fourth field is the person responsible field.

[0089] It can be seen that this embodiment further improves the directory data of the target archive by performing recognition and analysis on the archive image, extracting the target data from the recognition and analysis results to fill in the fourth field of the basic directory. At the same time, by checking the data to be checked, it is ensured that the filled fields in the basic directory and the review marks carried by the archive image are accurate, thereby improving the accuracy of the recording.

[0090] Regarding the execution order of step 140 and step 210 - step 220 , as an example, step 140 may be executed first, and then step 210 - step 220 , or step 210 - step 220 may be executed first, and then step 140 .

[0091] As another example, step 140 and step 210-step 220 can be executed in parallel. For example, step 140 and step 210-step 220 can be executed in parallel by different cameras. In this way, the computing performance of multiple cameras to be assigned tasks can be first obtained. The computing performance is determined based on the computing power and load conditions of the camera. Subsequently, based on the computing performance of each camera, a first camera and a second camera are determined from the multiple cameras. Among them, the computing performance of the first camera is higher than the computing performance of the second camera. The first camera may include one or more, and the second camera may also include one or more. In addition, the first camera and the second camera are determined from the cameras equipped in the physical room recorded in the specified directory.

[0092] It is understandable that since the task of generating the title, subject word and classification number of the target file requires high computing power, a neural network model may need to be used. Therefore, the task of generating the title, subject word and classification number of the target file can be assigned to the first camera with relatively high computing performance, and the first camera performs the above step 140. And the task of checking the data to be checked can be assigned to the second camera, and the second camera performs the above steps 210-220. The specific number of the first camera and the second camera can be determined based on the number of their respective tasks. For example, the maximum number of index entries can be set according to the starting file number and the ending file number, the entry data can be retrieved according to the file number, and then the specific first camera and the second camera can be selected. After the tasks are assigned, the total number of items, the allocation ratio, and the allocated and unallocated entries can be counted.

[0093] It can be seen that in this embodiment, the tasks of checking the check items and generating the target data, as well as the tasks of generating the titles, subject terms and classification numbers are divided into two processes and processed in parallel. On the one hand, it can improve the efficiency of catalog recording, and on the other hand, it can jointly allocate different computing stations to perform tasks with different computing power requirements, thereby improving the rationality of station scheduling.

[0094] According to some embodiments of the present application, the recording method may also include: Figure 3 Steps 310 to 320 are shown.

[0095] Step 310: Perform multi-level quality checks on the directory data of multiple target files;

[0096] Step 320: Perform supervised training on the quality inspection model using the error data obtained in the quality inspection.

[0097] Exemplarily, the target archives include multiple ones, and each target archive can generate corresponding directory data through any combination of the above embodiments. After the directory data of multiple target archives are obtained, multi-level quality inspection can be performed on all the directory data.

[0098] Exemplarily, the quality check includes standard character check, standard word check, and time item check according to the set starting file number and ending file number. Among them, the standard character check refers to selecting or setting the field to be checked to check whether the field contains illegal characters. Specifically, you can enter the file number range, select the query field, enter the characters that are allowed or not allowed to appear, and then perform a standard character check on the selected field.

[0099] The standard word check refers to checking whether a certain field contains non-standard organization words. For example, non-standard organization words can be added to the non-standard organization word database in advance. When performing the standard word check, select the check field and check whether there is non-standard organization words in the selected field.

[0100] The time check refers to checking whether the directory data is illogical or the time format is incorrect. For example, you can select the check field and / or query item, and then check the time field in the directory data.

[0101] Exemplarily, the multi-level quality checks can be performed by multiple parties in sequence. For example, after the catalog data is generated, one or more quality checks can be performed by the archive management system. Specifically, after the first camera completes the task of generating the title, subject terms and classification number, the quality check task of the title, subject terms and classification number is performed. And after the second camera completes the verification task, the quality check task of the data to be verified and the filled fourth field is performed. If the quality check fails, the catalog data is modified until the quality check passes. If the quality check passes, the next step can be one or more quality checks by the archives. Similarly, if the quality check fails, the catalog data is returned to the archive management system for modification until the quality check at the archive level passes. If the quality check passes, the next step can be one or more quality checks by a third party.

[0102] The multi-level quality inspection process may include one or more of the following:

[0103] 1. Entry inspection, which is divided into quality inspection tasks for titles, subject terms and classification numbers, and quality inspection tasks for data to be checked and the fourth field that has been filled. Specifically, the embedded rule engine can be used to perform quality inspection tasks for titles, subject terms and classification numbers. Among them, the rule engine is used to define and execute various inspection rules, and the rule engine includes rules corresponding to titles, subject terms and classification numbers. For example, for the quality inspection tasks for titles, subject terms and classification numbers, the rule engine may include the following rules: the title length should be within a specific range, the subject terms should comply with specific vocabulary specifications, and the classification number should follow the established classification system, etc. Of course, the rule engine is not limited to including the above rules. By encoding these rules into the rule engine, the archive management system can automatically scan the catalog data and quickly and accurately detect catalog entries that do not comply with the rules.

[0104] In addition, the entry check may also include using the historical chronology knowledge graph to check the accuracy of the original chronology of the target archive. Specifically, time series analysis and pattern recognition technology can be used in combination with the historical chronology knowledge graph to detect whether the original chronology is extracted correctly. For example, by analyzing the time-related expressions in the text and matching them with the known chronology rules and the timeline of historical events, the accuracy of the original chronology can be determined.

[0105] In addition, entry checking can also include comprehensive detection of catalog data for typos, inappropriate simplified and traditional Chinese, improper placement, extraction errors, incomplete or redundant information, incomplete elements, subject deviation, unclear semantics, inappropriate verbs, missing or incorrect subject terms, missing or incorrect subject free terms, missing or incorrect free terms, missing or incorrect verbs, and other error types.

[0106] The error rate of the above-mentioned various quality inspection standards needs to not exceed the preset error rate threshold, for example, the error rate threshold is one thousandth. The directory data of all target archives needs to be checked. When an error is detected, the system automatically records the first error data, including the error type, error location (such as a specific entry field), error content, etc., and feeds back to the front-end quality inspection personnel in real time for correction and automatic problem recording. At the same time, the first error data is classified and counted and stored in the knowledge base to provide data support for subsequent machine learning model training, so as to continuously optimize the inspection rules and model performance and improve quality inspection efficiency.

[0107] 2. Data integration, including checking the target files for generated catalog data, automatically normalizing the standard fields in the catalog data, including the fields corresponding to the file-level entries and the fields corresponding to the document-level entries, checking for errors in subject terms, missing file-level titles, incorrect classification numbers, etc., and troubleshooting for format errors, blank required entries, etc.

[0108] Exemplarily, a field rule library can be pre-established to cover the specification requirements of all standard fields corresponding to file-level entries and document-level entries. For example, for file-level titles, it is stipulated that they must contain specific key information and the format must comply with certain standards; for classification numbers, their coding rules and hierarchical structures are clarified. In this way, the field rule library can be used to check the file-level entry fields and file-level entry fields of the directory data. Specifically, the fields in the directory data can be compared with the specifications in the field rule library to automatically detect whether there are problems such as no subject word errors, missing file-level titles, and incorrect classification numbers.

[0109] In addition, regular expressions are used to check the format of directory data. Specifically, format parsing technology can be used to interpret the format of directory data, such as date format, number format, character encoding, etc. Then regular expressions are used to check whether the format of directory data meets the requirements. For fields that do not meet the format, the format is automatically converted or the quality inspector is prompted to make corrections. At the same time, it is checked whether the required items are empty. If they are empty, a warning is issued and the items are required to be completed.

[0110] In addition, the consistency and logic between different fields in the directory data can also be checked through the pre-established data association model. The data association model is used to characterize the time association and expression consistency between multiple fields. For example, it can be checked whether the file creation time recorded in the file-level entry field in the directory data is earlier than the file formation time recorded in the file-level entry field, or whether the expression of the responsible person information in different fields is consistent. If inconsistencies or logical errors are found, they will be marked in time and fed back to the quality inspection personnel for verification and correction.

[0111] 3. General quality inspection, including sampling the directory data according to the sampling rate to obtain the second error data. The sampling rate is positively correlated with the importance of the target archive and / or the historical error rate of the directory data. If necessary, the general quality inspection can be carried out in full inspection. Specifically, the sampling rate of each batch can be automatically determined according to the preset risk threshold and sampling rate range. For example, for data with high importance and high historical errors, the sampling rate is increased; for data with relatively stable and few historical errors, the sampling rate is appropriately reduced. In this way, the efficiency and pertinence of sampling inspection can be improved under the premise of ensuring the quality of quality inspection. The quality of quality inspection is ensured by checking the target data and target archive after data integration. And the general inspection results are directly fed back to the project management group and the front-end quality inspection group in a timely manner. Support real-time message push, problem tracking and collaboration functions to facilitate timely communication and problem solving among all parties. The second error data obtained through the general quality inspection can be recorded in the knowledge base in real time, and the efficiency of future quality inspections can be improved through machine self-learning. According to the distribution and type of errors in the general quality inspection results, the inspection rules and model parameters in the item inspection and data integration links are adjusted to improve the accuracy and coverage of automatic detection. Through this closed-loop learning and optimization mechanism, the efficiency and quality of the entire quality inspection process can be gradually improved.

[0112] 4. Random inspection: The directory data of multiple target files can be divided into multiple batches for random inspection. Taking into account the workload and skill level of quality inspectors, the random inspection tasks can be reasonably allocated. For example, the task progress and workload of each quality inspector can be monitored in real time, and random inspection tasks can be allocated to personnel with lighter workloads and corresponding skills (such as being familiar with specific types of files or being good at certain error detection) to ensure fairness and efficiency in task allocation. At the same time, the system instantly displays the random inspection ratio, allowing quality inspectors to clearly understand the scope and requirements of random inspections.

[0113] During the sampling process, quality inspectors can directly annotate the problems found in the sampling on the system interface. The annotation content includes problem description, error type, and suggested correction method. The archive management system can obtain the annotation content of the catalog data input by the quality inspectors, save the annotation content to the database in real time, and automatically associate the annotation content with the relevant archive data. At the same time, the new error types and problem patterns involved in the annotations are automatically extracted and updated to the knowledge base, enriching the content of the knowledge base and providing more comprehensive reference and guidance for subsequent quality inspection work.

[0114] For archive image content that requires special marking, the archive management system also provides a convenient image indexing function. Quality inspectors can use blocks of different colors to mark archive images. The marking results are synchronized with the process and will not damage the original image. In this way, the archive management system can also obtain the annotations added by quality inspectors to the archive images, thereby obtaining third-party error data. In addition, the sampling process can trace back the recording results of each process, which is convenient for quality inspectors to view the processing of archive data in each link, which helps to accurately determine the root cause of the problem and trace the responsibility, and also provides strong support for data analysis and process optimization.

[0115] 5. A general random inspection includes using a preset error classification rule to classify one or more of the first error data, the second error data and the third error data, and calculating and outputting the error rate of the directory data.

[0116] Specifically, the error classification of the sampling results can be automatically performed in combination with the preset error classification rules. The rules can define some clear error types and classification standards, and automatically identify and classify some complex error situations that are difficult to clearly define through rules through a large amount of annotated error data. For example, for some semantically ambiguous but non-compliant error statements, they are classified according to context and semantic similarity.

[0117] The system automatically calculates the error rate of the sampling data and displays it to the chief sampling inspector in the form of intuitive charts (such as bar charts, line charts, etc.). The charts can be subdivided by different dimensions (such as error type, file category, time period, etc.), which makes it convenient for the chief sampling inspector to quickly understand the quality inspection situation and discover problem trends. At the same time, the error rate statistics are associated with the knowledge base to provide data support for subsequent quality analysis and improvement.

[0118] When the general sampling personnel find an error, they can make an electronic mark, which includes the error location, error type, correction suggestions, etc. The marked information is synchronized to the database in real time and associated with the relevant archive data and process flow. If the error found needs to be rolled back to the previous process for modification, the system provides a convenient process rollback function, which automatically transmits the relevant data and marked information to the corresponding process to ensure that the problem can be handled promptly and accurately.

[0119] 6. Final review: This includes obtaining the batch error rate of each sampled batch or the total sampled batch, and weighting the batch error rate based on the error type and the importance of the target archive to obtain the weighted error rate. Specifically, in the final review, a large amount of sampled data can be quickly processed to calculate the error rate of each batch. Consider the weights and importance of various error types to ensure that the calculation results of the error rate can truly reflect the quality level of the batch data. For example, for some key error types that seriously affect the use of archives, a higher weight is given to make them play a greater role in the error rate calculation.

[0120] Based on the calculated error rate and the preset error rate threshold, the archive management system provides intelligent batch rollback decision support. When the weighted error rate exceeds the error rate threshold, the system automatically triggers the batch rollback process, returns the batch of data to the cataloging process for data modification and rework, and sends notifications and reminders to relevant personnel. At the same time, the system can provide some suggestions and guidance based on historical data and the characteristics of the current batch to help cataloging personnel better modify the data. For batch data whose error rate does not exceed the rated value, the system will also return the found error information to the cataloging process for correction to ensure continuous improvement of data quality. When the weighted error rate does not exceed the error rate threshold, the catalog data is submitted.

[0121] For the tasks of reworking the whole batch, the sampling process can be re-inspected. When assigning the tasks of re-inspection, the system adopts an optimized task allocation algorithm, giving priority to the tasks that have not been inspected, to ensure the comprehensiveness and effectiveness of the inspection. At the same time, according to the characteristics of the rework data and the results of the first sampling inspection, the sampling ratio and the focus of the sampling inspection are adjusted to improve the pertinence and efficiency of the second sampling inspection.

[0122] After passing the above multi-level quality inspection, data can be submitted. Specifically, the qualified data will be uploaded to the database intermediate library, which includes the archive title subject word library, archive title free word library, archive title verb library, etc. in addition to the catalog database with completed recording. After uploading, data verification confirmation and verification will be carried out.

[0123] During the data submission process, the data uploaded to the database intermediate library is also verified in multiple dimensions. In addition to checking the integrity and consistency of the data (such as the data association and consistency between the catalog database, the archival title thesaurus, the archival title free thesaurus, and the archival title verb library), it also includes data accuracy (such as comparison with the original archives), standardization (such as compliance with established data standards and format requirements), etc. During the verification process, the system automatically records the verification results and promptly feeds back the problems found to the relevant personnel for processing.

[0124] Establish an electronic verification process. After the data is uploaded and verified, relevant personnel can perform electronic verification operations in the system. The verification process supports identity authentication and permission management to ensure that only personnel with corresponding permissions can perform verification. At the same time, the system tracks and records the verification process in real time, including verification time, verification personnel, verification results and other information, forming a complete process log for easy auditing and tracing. In addition, the system can set an automatic reminder function to remind relevant personnel to perform verification operations in a timely manner to ensure the smooth progress of the data submission process.

[0125] It is understandable that, compared with the directory recording of ordinary document files, the directory content of archives is often more detailed, the directory needs to strictly follow relevant regulations or standards, and the recording process is more complicated. The accuracy of the directory recording will affect the management and use of the subsequent directory, so it is necessary to strictly check the directory data and not allow any errors. In this embodiment, by performing multi-level quality inspection on the directory data of all target archives, the directory data obtained by automatic recording is strictly checked from multiple aspects and dimensions to ensure that the directory data is accurate and improve the reliability of the directory data obtained by automatic recording.

[0126] Based on the method for recording an archive directory described in any of the above embodiments, the present application also provides a computer program product, which includes one or more computer programs or instructions. The computer program or instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. When the computer program is executed by a processor, the method for recording an archive directory described in any of the above embodiments is implemented.

[0127] Based on the method for recording an archive directory described in any of the above embodiments, the present application also provides the following Figure 4 A schematic diagram of the structure of an electronic device is shown in FIG. Figure 4At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement a method for recording an archive directory as described in any of the above embodiments.

[0128] The present application also provides a computer storage medium, which stores a computer program. When the computer program is executed by a processor, it can be used to execute a method for recording an archive directory as described in any of the above embodiments.

[0129] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0130] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0131] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.

[0132] The above description is only an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.

[0133] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0134] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

Claims

1. A method for recording an archive directory, characterized in that: The method comprises: Acquire an archive image, bibliographic data, and a catalog template of a target archive, wherein the catalog template includes a plurality of fields to be filled; Fill the bibliographic data into the first field corresponding to the directory template to obtain the basic directory of the target archive; Mounting the basic directory and the archive image, and filling the second field of the basic directory based on the identification information of the archive image; Generate the title of the target archive according to the subject information recognized from the archive image by using OCR technology, determine candidate subject words from the title, and determine the candidate subject words recorded in the subject word database as the subject words of the target archive, match the target category words used to describe the content characteristics of the target archive from the category word database, and determine the classification number of the target archive from the classification system based on the target category words, and fill the title, the subject words and the classification number into the third field corresponding to the basic directory to obtain the directory data of the target archive; Performing multi-level quality checks on the catalog data of the plurality of target archives, including: Data integration, used to check the consistency and logic between multiple fields in the directory data using a data association model, including whether the file creation time recorded in the file-level entry field in the directory data is earlier than the case file formation time recorded in the case file-level entry field, and whether the representation of the responsible person information in different fields is consistent; Total quality inspection, used to perform random inspection on the catalog data according to a random inspection rate, where the random inspection rate is positively correlated with the historical error rate of the catalog data; Sampling inspection is used to obtain the annotations input by the quality inspectors on the catalog data, as well as the annotations added to the archive images of the target archives using the image indexing function, and the annotations are used to trace back to the recording results of each process; The total sampling is used to classify the error data by error classification rules and calculate the error rate of the directory data.

2. The method according to claim 1, characterized in that Before filling the bibliographic data into the first field corresponding to the catalog template, the method further includes: Determine whether the catalog data passes data integrity verification; the data integrity verification includes one or more of total number of pieces consistency detection, total number of bytes consistency detection, metadata item integrity detection, metadata mandatory catalog item detection, process information integrity detection, continuity metadata item detection, content data integrity detection, attachment data integrity detection and information package content data integrity detection.

3. The method according to claim 1, characterized in that The filling of the second field of the basic directory based on the identification information of the archive image comprises: Acquire identification information of the archive image, wherein the identification information includes one or more of the start and end numbers of the files in the volume of the target archive, and the start image number and the end image number of the archive image; A second field of the base directory is populated based on the identification information, and the second field is used to index to the target archive.

4. The method according to claim 1, characterized in that: The method further comprises: Performing verification processing on the data to be verified, the data to be verified includes one or more of the basic directory, the filled second field, and the review mark carried by the archive image; The archive image is identified and analyzed, and if the identification and analysis result includes the target data corresponding to the field to be filled, the target data is filled into the fourth field corresponding to the basic directory.

5. The method according to claim 4, characterized in that The method further comprises: Obtain the computing performance of multiple positions to be assigned tasks; Determine a first camera position and a second camera position from the plurality of camera positions based on the computing performance, wherein the computing performance of the first camera position is higher than the computing performance of the second camera position; The first camera is assigned the task of generating the title, subject words and classification number of the target file, and the second camera is assigned the task of checking the data to be checked.

6. The method according to any one of claims 1 to 5, characterized in that: Before the data integration, the method further includes: item checking, including using an embedded rule engine to perform a quality check task of the title, subject word and classification number, and using a historical chronology knowledge graph to check the accuracy of the original chronology of the target archive to obtain first error data; wherein the rule engine includes rules corresponding to the title, subject word and classification number; The data integration further includes using a field rule library to check the file-level entry fields and the file-level entry fields of the directory data, and using a regular expression to check the format of the directory data. The error data used in the total random inspection includes the first error data, the second error data obtained by the total quality inspection, and the third error data obtained by the random inspection; After the total sampling inspection, the method further includes: final review, including obtaining a batch error rate of each sampling inspection batch or the total sampling inspection batch, and weighting the batch error rate based on the error type and the importance of the target archive to obtain a weighted error rate, and if the weighted error rate exceeds a preset error rate threshold, returning the catalog data for modification, and if the weighted error rate does not exceed the error rate threshold, submitting the catalog data; The method further comprises: performing supervised training on a quality inspection model using error data obtained in the quality inspection.

7. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

8. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing processor-executable instructions; Wherein, when the processor calls the executable instruction, it implements the operation of any method described in claims 1-6.

9. A computer-readable storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed by a processor, the steps of any method described in claims 1-6 are implemented.

Citation Information

Patent Citations

  • Archive filing method and device, electronic equipment and computer readable storage medium

    CN112052749A

  • Sound image file recording method and device based on image recognition

    CN114117095A

  • Automatic file quality inspection method, device and equipment, medium and product

    CN119204992A