An archive digitization automatic classification and retrieval method based on deep learning
Patent Information
- Application Number
- CN202610828146.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-09-04
AI Technical Summary
[0005]本发明的主要目的在于提供一种基于深度学习的档案数字化自动分类及检索方法,可以有效解决缺乏对扫描图像内部区域差异的细粒度识别能力的问题
1、 本发明提供一种基于深度学习的档案数字化自动分类及检索方法,通过对扫描图像灰度矩阵与字符像素密度分布进行像素级加权求和,并结合版面分区坐标执行区域切片及像素均衡比值计算,使不同区域的清晰程度形成可量化分布结果,弱化扫描质量波动对后续处理带来的干扰,使文本识别基础具备区域差异感知能力。
Smart Images

Figure CN122692292A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep learning technology, and in particular to a method for automatic classification and retrieval of digital archives based on deep learning. Background Technology
[0002] The field of data retrieval and classification technology encompasses data organization structure construction, text content parsing, feature representation, similarity calculation, index building, and query matching. This technology focuses on the standardized processing of unstructured and semi-structured data. By extracting content and semantically annotating archival carriers such as text and scanned images, a computable feature set is formed. Based on inverted indexes, vector space representation, and tree-structured classification systems, hierarchical data management and rapid location are achieved.
[0003] Chinese patent document CN121092759A discloses a method and system for digital governance of archives based on a large model. This method includes collecting archive data, which includes images obtained from scanning paper archives, digitized audio and video archives, and directly imported electronic archives; inputting the archive data into a pre-trained large model for multimodal recognition and information extraction to obtain structured data; based on the structured data, automatically classifying and labeling the archives using the large model to generate structured metadata, semantic tags, and a knowledge graph; and forming embedding vectors based on the structured data, structured metadata, semantic tags, and knowledge graph, storing them in a distributed database, and forming an intelligent retrieval interface based on the large model. This application has advantages such as high efficiency, high accuracy, and generalization.
[0004] The existing technology has the following problems: Existing technologies, in actual operation, focus on uniform feature representation and index construction of text content. The processing starting point mostly relies on the overall text extraction results, lacking the ability to fine-grainedly identify differences in the internal regions of scanned images. When the original document has local blurring, damage, or uneven lighting, the differences in information quality in different regions are ignored, leading to deviations in the basis of subsequent feature representation. For example, blurred areas of marginal annotations in scanned contract documents are easily misinterpreted as main text content, affecting classification judgment from the source. Summary of the Invention
[0005] The main objective of this invention is to provide a deep learning-based method for automatic classification and retrieval of digital archives, which can effectively solve the problem of lacking fine-grained recognition capabilities for differences in internal regions of scanned images.
[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A deep learning-based method for automatic classification and retrieval of digitized archives includes the following steps: S1: Image sharpness coefficient acquisition: Obtain the grayscale matrix, character pixel density distribution value, and page layout coordinate set of the scanned image of the paper document. Call the first two sets of data to perform pixel-level weighted summation. Based on the page layout coordinate set, slice the weighted result into regions and calculate the pixel balance ratio of each region. Compare the result with the preset sharpness discrimination threshold region by region and filter the blocks that meet the discrimination threshold to obtain the image sharpness distribution coefficient. S2: Calculation of structural distribution ratio: Based on the image clarity distribution coefficient, obtain the character stroke width set, line spacing value set and paragraph block size sequence, call the first two sets of data to perform difference operation and generate stroke spacing ratio, the paragraph block size sequence classifies the stroke spacing ratio into regions and calculates the classification frequency ratio to generate text structural distribution ratio. S3: Category discrimination value generation: Based on the text structure distribution ratio, obtain the archive directory keyword sequence, archive number encoding bit value and time field number string, call the first two sets of data to perform matching counting operation and generate keyword matching frequency value, judge the format consistency based on the time field number string and the set time format benchmark value, and generate time matching value, perform weighted summation operation on the two data generation results to form a comprehensive discrimination series, divide the result into intervals with the preset classification boundary value and filter the category number to generate archive category discrimination value; S4: Tag weight allocation: For the file category discrimination value, obtain the category tag index table, field mapping code sequence and tag occurrence frequency statistics, call the first two sets of data for mapping matching and generate tag matching count value, perform normalization processing on the tag matching count value of the paragraph block size sequence and calculate the tag weight ratio to obtain the tag weight allocation rate. S5: Matching and sorting value output: Based on the tag weight allocation rate, obtain the search keyword input sequence, tag weight allocation rate sequence and index position offset value, call the product operation of the first two sets of data and generate matching weights, sort the matching weights according to the index position offset value and filter the preceding position records to obtain the search matching and sorting value.
[0007] Preferably, the image clarity distribution coefficient includes the mean regional clarity, pixel balance ratio, and proportion of effective clear blocks; the text structure distribution ratio includes the stroke spacing ratio distribution, paragraph size classification proportion, and line spacing variation proportion; the file category discrimination value includes keyword matching frequency value, time format matching value, and comprehensive discrimination sequence interval category; the tag weight allocation rate includes tag matching count value, tag normalized weight value, and tag frequency proportion; and the retrieval matching ranking value includes the matching weight sequence, ranking position index, and previous retrieval result set.
[0008] Preferably, in the pixel-level weighted summation operation in the acquisition of the image clarity coefficient in S1, the weighting coefficient is adaptively adjusted based on the variance of the pixel gray value gradient of the gray matrix and the character pixel density distribution value, , wherein is a gray matrix gradient value, is the variance of the character pixel density distribution value, is an adaptive adjustment coefficient, with a value range of 0.3-0.7, and is dynamically corrected according to the resolution of the scanned image.
[0009] Preferably, in the acquisition of the image clarity coefficient in S1, the preset clarity discrimination threshold adopts a double-layer threshold design, including a basic clarity threshold and an optimized clarity threshold. The basic clarity threshold is used for preliminary screening of effective clear blocks, and the optimized clarity threshold is used for secondary verification of the blocks after preliminary screening, wherein the value of the basic clarity threshold is 0.65-0.75, the value of the optimized clarity threshold is 0.75-0.85, and the difference between the two is fixed at 0.1.
[0010] Preferably, when the paragraph block size sequence performs region classification on the stroke spacing ratio in the calculation of the structure distribution ratio in S2, the K-means clustering algorithm is adopted, the number of clusters is adaptively determined according to the number of paragraphs on the archive layout, the mean and standard deviation of the stroke spacing ratio are used as the cluster center initialization parameters in the clustering process, the number of iterations is not less than 10 times, and the clustering precision error is controlled within 5%.
[0011] Preferably, in the generation of the category discrimination value in S3, the format consistency judgment between the time field digit string and the set time format reference value adopts a combination of regular expression matching and time logic verification. The set time format reference values include year-month-day format (YYYY-MM-DD, YYYY年MM月DD日) and year-month format (YYYY-MM, YYYY年MM月). Regular expression matching is used for preliminary judgment of format compliance, and time logic verification is used for verifying the rationality of the date (e.g., the month does not exceed 12, and the date does not exceed the maximum number of days of the corresponding month).
[0012] Preferably, in the weighted summation operation of the comprehensive discrimination sequence in the generation of the category discrimination value in S3, the weight coefficient of the keyword matching frequency value is 0.6-0.7, and the weight coefficient of the time matching value is 0.3-0.4. The weight coefficients can be manually adjusted according to the archive type (document archive, scientific and technological archive, accounting archive), and the sum of the two weight coefficients is always 1.
[0013] Preferably, the normalization processing in the label weight allocation in S4 adopts the min-max normalization algorithm to map the label matching count value to the 0-1 interval, and the specific calculation formula is: , wherein Let be the normalized weight value of the i-th label. The match count value of the i-th tag, The maximum value of the match count for all tags. Find the minimum match count for all tags.
[0014] Preferably, when sorting the matching weights in the S5 matching sorting value output, a quick sorting algorithm is used, with the sorting priority from high to low matching weights. When the matching weights are the same, the records with smaller index position offsets are sorted first. The number of records to be filtered in the preceding position can be set according to user needs, with the default being the first 10 records, and the number of records to be filtered can be manually adjusted by the user.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention provides an automatic classification and retrieval method for digital archives based on deep learning. By performing pixel-level weighted summation of the grayscale matrix and character pixel density distribution of the scanned image, and combining it with the layout partition coordinates to perform region slicing and pixel balance ratio calculation, the clarity of different regions forms a quantifiable distribution result, weakening the interference of scan quality fluctuations on subsequent processing, and enabling the text recognition foundation to have the ability to perceive regional differences.
[0016] 2. This invention provides a deep learning-based method for automatic classification and retrieval of digital archives. It introduces the difference and classification relationship between character stroke width, line spacing and paragraph block size to form the stroke spacing ratio and structure frequency proportion, so that the text layout features are transformed from discrete descriptions to structural expressions with statistical regularities, thereby enhancing the ability to distinguish the differences between different archive formats.
[0017] 3. This invention provides a deep learning-based method for automatic classification and retrieval of digital archives. It combines keyword matching and counting, number encoding digits, and time field format consistency judgment to form a multi-dimensional joint discrimination sequence, enabling semantic clues and format features to work synergistically, reducing the risk of misjudgment based on a single feature, and making the classification process more stable and interpretable.
[0018] 4. This invention provides a method for automatic classification and retrieval of digital archives based on deep learning. It further generates tag matching counts by mapping tag indexes and field codes, and performs normalization processing by combining paragraph size information, so that the tag weight distribution and text structure are coupled, thereby improving the responsiveness of tag expression to content hierarchy. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram of the S1 process of the present invention; Figure 3 This is a schematic diagram of the S2 process of the present invention; Figure 4 This is a schematic diagram of the S3 process of the present invention; Figure 5 This is a schematic diagram of the S4 process of the present invention; Figure 6 This is a schematic diagram of the S5 process of the present invention. Detailed Implementation
[0020] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.
[0021] Example 1: This method is based on a deep learning framework. The specific implementation environment is as follows, ensuring the efficiency and accuracy of data processing in each step: The following preset values are also provided: Example 2: Image Clarity Coefficient Acquisition: 100 paper documents were scanned, and the grayscale matrix, character pixel density distribution values, and page layout coordinate set of each document's scanned image were obtained. Taking one document (scanned image size 210mm×297mm, grayscale matrix 2480×3508) as an example, the weighting coefficient α was corrected to 0.52 based on the image resolution. The image clarity distribution coefficient of this document was obtained through pixel-level weighted summation, region slicing, and equalization ratio calculations. Specific data are shown in the table below. The remaining 99 documents were calculated using the same method. The effective clear block screening pass rate reached 98%, meeting the requirements for subsequent processing. Structural ratio calculation: Based on the aforementioned image clarity distribution coefficients, the character stroke width set, line spacing value set, and paragraph block size sequence are extracted for each document. The K-means clustering algorithm is used to classify the stroke spacing ratios into regions. The number of clusters is adaptively determined based on the number of paragraphs in each document (3-5 clusters for administrative documents, 4-6 for scientific and technical documents, and 2-4 for accounting documents). Taking one scientific and technical document (containing 6 paragraphs) as an example, with 5 clusters and 12 iterations, the clustering accuracy error is 3.8%. The generated text structure distribution ratio is shown in the table below. The remaining documents are calculated using the same logic, achieving a structural classification accuracy of 97.5%.
[0022] Generation of category discrimination value: Extract the directory keyword sequence, file number coding bit value and time field numeric string of each file, generate the file category discrimination value through keyword matching, time format verification and weighted summation, and then determine the category to which the file belongs (the prefix of document file number is WS, that of scientific and technological files is KJ, and that of accounting files is KJ). A combination of regular expression matching and time logic verification is used to verify the format of the time field (supports formats such as YYYY-MM-DD, YYYY年MM月DD日), eliminating interference from invalid time fields. The category discrimination results of 100 samples are shown in the following table. The classification accuracy reaches 99%, and only 1 accounting file has discrimination deviation due to ambiguous keywords.
[0023] Label weight allocation: For files with determined categories, call the classification label index table, field mapping coding sequence and statistical value of label occurrence frequency, obtain label matching count values through mapping matching, map them to the 0-1 interval by using the min-max normalization algorithm, and calculate the proportion of label weights. Taking an accounting file (with a category discrimination value of 3, corresponding to an accounting file) as an example, 3 core labels are extracted, and their weight allocation is shown in the following table; the weight allocation of all samples conforms to this logic, and the rationality of weight allocation reaches 98.5%, which provides reliable support for subsequent retrieval.
[0024] Input retrieval keyword sequences (in this example, 3 groups of retrieval keywords "2025 annual accounting vouchers", "scientific and technological project report", and "document notice" are selected), call the label weight allocation rate sequence and index position offset value, generate matching weights through product operation, sort from high to low according to matching weights by using the quick sort algorithm, and screen the top 10 pre-order position records. Taking "2025 annual accounting vouchers" as the retrieval keyword, the sorting of retrieval results is shown in the following table. The retrieval accuracy reaches 98%, and the sorting response time is ≤ 0.5 seconds, which meets the actual retrieval requirements; the other two groups of retrieval keywords have similar retrieval effects, and all achieve the expected objectives.
[0025] Example 3: In order to verify the effectiveness, stability and practicability of the method of the present invention, three common types of files, namely document files, scientific and technological files and accounting files, are selected, with 1000 files for each type, a total of 3000 files are used as test samples, and tests are carried out in the above preset implementation environment. The test indicators include classification accuracy, retrieval response time and retrieval accuracy, and the test results are shown in the following table: The working principle of this invention is as follows: First, by performing pixel-level weighted summation on the grayscale matrix and character pixel density distribution of the scanned image, and combining it with the layout partition coordinates to perform region slicing and pixel equalization ratio calculation, the clarity of different regions forms a quantifiable distribution result, weakening the interference of scan quality fluctuations on subsequent processing, and enabling the text recognition foundation to have the ability to perceive regional differences. Second, by introducing the difference and classification relationship between character stroke width, line spacing, and paragraph block size, a stroke spacing ratio and structural frequency proportion are formed, transforming the text layout features from discrete descriptions to a structural expression with statistical regularity, enhancing the ability to distinguish differences in different document formats. Then, by combining the directory keyword matching count, number encoding digits, and time field format consistency judgment, a multi-dimensional joint discrimination sequence is formed, enabling semantic clues and format features to work synergistically, reducing the risk of misjudgment of a single feature, and making the classification process more stable and interpretable. Finally, by further generating tag matching counts through tag index and field encoding mapping, and combining them with paragraph size information for normalization processing, the tag weight distribution and text structure are coupled, improving the responsiveness of tag expression to content hierarchy.
[0026] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A method for automatic classification and retrieval of digitized archives based on deep learning, comprising the following steps, characterized in that: S1: Image sharpness coefficient acquisition: Obtain the grayscale matrix, character pixel density distribution value, and page layout coordinate set of the scanned image of the paper document. Call the first two sets of data to perform pixel-level weighted summation. Based on the page layout coordinate set, slice the weighted result into regions and calculate the pixel balance ratio of each region. Compare the result with the preset sharpness discrimination threshold region by region and filter the blocks that meet the discrimination threshold to obtain the image sharpness distribution coefficient. S2: Calculation of structural distribution ratio: Based on the image clarity distribution coefficient, obtain the character stroke width set, line spacing value set and paragraph block size sequence, call the first two sets of data to perform difference operation and generate stroke spacing ratio, the paragraph block size sequence classifies the stroke spacing ratio into regions and calculates the classification frequency ratio to generate text structural distribution ratio. S3: Category discrimination value generation: Based on the text structure distribution ratio, obtain the archive directory keyword sequence, archive number encoding bit value and time field number string, call the first two sets of data to perform matching counting operation and generate keyword matching frequency value, judge the format consistency based on the time field number string and the set time format benchmark value, and generate time matching value, perform weighted summation operation on the two data generation results to form a comprehensive discrimination series, divide the result into intervals with the preset classification boundary value and filter the category number to generate archive category discrimination value; S4: Tag weight allocation: For the file category discrimination value, obtain the category tag index table, field mapping code sequence and tag occurrence frequency statistics, call the first two sets of data for mapping matching and generate tag matching count value, perform normalization processing on the tag matching count value of the paragraph block size sequence and calculate the tag weight ratio to obtain the tag weight allocation rate. S5: Matching and sorting value output: Based on the tag weight allocation rate, obtain the search keyword input sequence, tag weight allocation rate sequence and index position offset value, call the product operation of the first two sets of data and generate matching weights, sort the matching weights according to the index position offset value and filter the preceding position records to obtain the search matching and sorting value.
2. The method for automatic classification and retrieval of digital archives based on deep learning according to claim 1, characterized in that: The image clarity distribution coefficient includes the mean regional clarity, pixel balance ratio, and proportion of effective clear blocks; the text structure distribution ratio includes the stroke spacing ratio distribution, paragraph size classification proportion, and line spacing variation proportion; the archive category discrimination value includes keyword matching frequency value, time format matching value, and comprehensive discrimination sequence interval category; the tag weight allocation rate includes tag matching count value, tag normalized weight value, and tag frequency proportion; the retrieval matching ranking value includes the matching weight sequence, ranking position index, and previous retrieval result set.
3. The method for automatic classification and retrieval of digital archives based on deep learning according to claim 1, characterized in that: The pixel-level weighted summation operation in obtaining the image sharpness coefficient in S1 uses weighting coefficients that are adaptively adjusted based on the variance of the pixel grayscale value gradient and the character pixel density distribution values in the grayscale matrix. in, The gradient value of the grayscale matrix. The variance of the character pixel density distribution values. It is an adaptive adjustment coefficient, with a value range of 0.3-0.7, and is dynamically corrected according to the resolution of the scanned image.
4. The method for automatic classification and retrieval of digital archives based on deep learning according to claim 3, characterized in that: The image sharpness coefficient acquisition in S1 adopts a dual-threshold design for the preset sharpness discrimination threshold, including a basic sharpness threshold and an optimized sharpness threshold. The basic sharpness threshold is used to initially screen effective sharp blocks, and the optimized sharpness threshold is used to perform secondary verification on the initially screened blocks. The basic sharpness threshold is 0.65-0.75, the optimized sharpness threshold is 0.75-0.85, and the difference between the two is fixed at 0.
1.
5. The method for automatic classification and retrieval of digital archives based on deep learning according to claim 1, characterized in that: When calculating the S2 structure distribution ratio, the paragraph block size sequence is used to classify the stroke spacing ratio into regions. The K-means clustering algorithm is used, and the number of clusters is adaptively determined according to the number of paragraphs on the archive page. During the clustering process, the mean and standard deviation of the stroke spacing ratio are used as the initialization parameters for the cluster centers. The number of iterations is not less than 10, and the clustering accuracy error is controlled within 5%.
6. The method for automatic classification and retrieval of digital archives based on deep learning according to claim 1, characterized in that: The consistency judgment of the time field numeric string with the set time format baseline value in the generation of the S3 category discrimination value adopts a combination of regular expression matching and time logic verification. The set time format baseline value includes year-month-day format and year-month format. Regular expression matching is used to initially judge the compliance of the format, and time logic verification is used to verify the rationality of the date.
7. The method for automatic classification and retrieval of digital archives based on deep learning according to claim 6, characterized in that: In the generation of the S3 category discrimination value, the weighted summation operation of the comprehensive discrimination sequence has a weight coefficient of 0.6-0.7 for keyword matching frequency value and a weight coefficient of 0.3-0.4 for time matching value. The weight coefficients can be manually adjusted according to the file type, and the sum of the two weight coefficients is always 1.
8. The method for automatic classification and retrieval of digital archives based on deep learning according to claim 1, characterized in that: The normalization process in the S4 label weight allocation uses the min-max normalization algorithm to map the label matching count value to the 0-1 interval. The specific calculation formula is as follows: in Let be the normalized weight value of the i-th label. The match count value of the i-th tag, The maximum value of the match count for all tags. Find the minimum match count for all tags.
9. The method for automatic classification and retrieval of digital archives based on deep learning according to claim 1, characterized in that: When sorting the matching weights in the S5 matching sorting value output, a quick sorting algorithm is used. The sorting priority is from high to low matching weights. When the matching weights are the same, the records with smaller index position offsets are sorted first. The number of records to be filtered in the preceding position can be set according to user needs. The default is to filter the first 10 records, and users can manually adjust the number of records to be filtered.
Citation Information
Patent Citations
File digital governance method and system based on large model
CN121092759A