Multidimensional Space Storage Management Method, System and Storage Medium for Digital Archives

By digitizing paper archives, extracting feature curves and encrypting storage, the problem of difficulty in preserving paper archives is solved, and efficient management and secure storage of archives are achieved.

CN119759845BActive Publication Date: 2025-07-11JILIN ZHIXIN SCIENCE & TRADE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510258100.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-11
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

现有技术中,纸质档案的保存困难且时效较短,如何提高档案留存时效和降低存储难度是亟需解决的问题。

Method used

By scanning paper archives to generate digital archives, extracting the feature curves of each image, calculating the anomalies and universality of the image, encrypting and storing them based on the feature curves, realizing multi-dimensional spatial management of archives.

Benefits of technology

It improves the storage time of archives, reduces the storage space requirements, and ensures the security and query efficiency of archives through feature curve encryption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119759845B_ABST
    Figure CN119759845B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of digital storage technology, and specifically discloses a multi-dimensional space storage management method, system and storage medium for a digital archive. The method includes receiving an archive to be stored, scanning the archive to obtain a digital archive; the digital archive is a set of scanned images; for any digital archive, reading the images in sequence, identifying the images, locating text boxes, and extracting the feature curves of each image based on the located text boxes; determining the abnormality degree of any image according to the feature curves, counting the abnormality degrees of all images in the same digital archive, and determining the universality degree of the digital archive; and archiving and storing the digital archive according to the universality degree. The present invention scans the archive to obtain a digital archive, extracts features from the digital archive, and then classifies, archives and stores the features, greatly improving the retention time of data and reducing the data storage space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital storage, and specifically to a multi-dimensional space storage management method, system and storage medium for a digital archive. Background Art

[0002] Archives, a Chinese term, refer to various forms of original records with preservation value directly formed in various social activities. They are very common in enterprises. For various handover-related jobs in enterprises, archiving is required. Most of these archives are paper documents, which are very difficult to preserve. How to reduce the difficulty of archive retention and improve the efficiency of archive retention is the technical problem that the technical solution of the present invention aims to solve. Summary of the Invention

[0003] The purpose of the present invention is to provide a multi-dimensional space storage management method, system and storage medium for a digital archive to solve the problems raised in the above background art.

[0004] To achieve the above purpose, the present invention provides the following technical solutions:

[0005] A multi-dimensional space storage management method for a digital archive, the method includes:

[0006] Receiving the archives to be stored, scanning the archives to obtain digital archives; the digital archives are a set of scanned images;

[0007] For any digital archive, reading the images in sequence, identifying the images, locating the text boxes, and extracting the feature curves of each image based on the located text boxes;

[0008] Determining the abnormality degree of any image according to the feature curve, counting the abnormality degrees of all images in the same digital archive, and determining the universality degree of the digital archive; wherein, the abnormality degree is used to represent the degree of difference between an image and other images, and the universality degree represents the similarity degree of the digital archive to other digital archives;

[0009] Archiving and storing the digital archives according to the universality degree, and encrypting the images based on the feature curves during archiving.

[0010] As a further solution of the present invention: the step of, for any digital archive, reading the images in sequence, identifying the images, locating the text boxes, and extracting the feature curves of each image based on the located text boxes includes:

[0011] For any digital archive, obtaining the image size of each image;

[0012] Sorting the images in ascending order according to the image size, and reading the images in sequence in the images sorted in ascending order;

[0013] Input the image into a preset contour recognition model to locate the text box; among them, the contour recognition model contains a performance adjustment port, and the initial performance is the lowest performance. When the positioning process of any image fails, a preset performance step is added to the lowest performance;

[0014] Extract the feature curve of each image based on the located text box.

[0015] As a further solution of the present invention: the step of extracting the feature curve of each image based on the located text box includes:

[0016] Select points on the boundary of the text box based on a preset step; the step is the distance between adjacent points;

[0017] For any point, calculate its distance from the diagonal of the image and synchronously determine the projection point;

[0018] Calculate the distance between adjacent projection points on the diagonal. When the distance is less than a preset threshold, classify all the points corresponding to the projection points into one category;

[0019] Take the distance between the points of the same category as a vector and calculate the combined vector as the vector at the projection point;

[0020] Fit the feature curve according to the endpoints of the vectors at all projection points.

[0021] As a further solution of the present invention: the step of determining the abnormality degree of any image according to the feature curve, counting the abnormality degrees of all images in the same digital file, and determining the universality degree of the digital file includes:

[0022] Record the feature curve of each image in real time to construct a feature curve library;

[0023] For any image, randomly compare the feature curve of this image with the feature curves in the feature curve library and calculate the similarity;

[0024] When the preset comparison jump-out condition is met, stop the comparison; the comparison jump-out conditions include that the similarity reaches the preset similarity threshold and the number of random comparison times reaches the preset number threshold;

[0025] Select the maximum similarity in all comparison results and determine the abnormality degree according to the maximum similarity and the number of random comparison times;

[0026] Determine the universality degree of the digital file according to the abnormality degree of each image.

[0027] As a further solution of the present invention: the calculation process of the similarity is:

[0028] Obtain the fitting function of the feature curve and calculate the similarity according to the fitting function; ; wherein, represents the similarity of two feature curves, represents the right endpoint of the common segment of the diagonals corresponding to two feature curves, represents the left endpoint of the common segment of the diagonals corresponding to two feature curves, is generally set to 0; represents the length of the longer diagonal among the diagonals corresponding to two feature curves; and respectively represent two feature curves; and respectively represent the fitting functions of two feature curves;

[0029] The calculation process of the abnormality degree is as follows:

[0030] ; wherein, represents the abnormality degree, represents the maximum similarity of a certain image in the random comparison process, represents the number of random comparison times, and are preset correction coefficients;

[0031] The calculation process of the universality is as follows:

[0032] ; wherein, represents the universality, represents the th abnormality degree of the image, represents the total number of images in the digital archive.

[0033] As a further solution of the present invention: The step of archiving and storing the digital archive according to the universality and encrypting the image based on the feature curve during archiving includes:

[0034] Archive and store the digital archive according to a preset universality range;

[0035] During storage, perform text recognition on the image, locate the text, use the center of the text as a point position, and update the feature curve;

[0036] Input the updated feature curve into the trained transcoding model, output the encryption key, and encrypt the image.

[0037] The technical solution of the present invention also provides a multi-dimensional space storage management system for a digital archive, and the system includes:

[0038] A digital processing module, configured to receive the file to be stored, scan the file to obtain a digital file; the digital file is a set of scanned images;

[0039] A feature extraction module, configured to, for any digital file, sequentially read the images, identify the images, locate the text boxes, and extract the feature curves of each image based on the located text boxes;

[0040] An image analysis module, configured to determine the abnormality degree of any image according to the feature curves, count the abnormality degrees of all images in the same digital file, and determine the universality degree of the digital file; wherein, the abnormality degree is used to characterize the degree of difference between the image and other images, and the universality degree represents the similarity degree between the digital file and other digital files;

[0041] An image encryption module, configured to archive and store the digital file according to the universality degree, and encrypt the image based on the feature curves during archiving.

[0042] As a further solution of the present invention: the feature extraction module includes:

[0043] An image size acquisition unit, configured to, for any digital file, acquire the image size of each image;

[0044] An image sorting unit, configured to sort the images in ascending order according to the image size, and sequentially read the images in the images sorted in ascending order;

[0045] A text box positioning unit, configured to input the image into a preset contour recognition model to locate the text box; wherein, the contour recognition model has a performance adjustment port, and the initial performance is the lowest performance. When the positioning process of any image fails, a preset performance step is added to the lowest performance;

[0046] An extraction execution unit, configured to extract the feature curves of each image based on the located text boxes.

[0047] As a further solution of the present invention: the extraction execution unit includes:

[0048] A point selection sub-unit, configured to select points on the boundary of the text box based on a preset step; the step is the distance between adjacent points;

[0049] A projection sub-unit, configured to, for any point, calculate the distance between it and the diagonal of the image, and synchronously determine the projection point;

[0050] A merging sub-unit, configured to calculate the distance between adjacent projection points on the diagonal. When the distance is less than a preset threshold, all the points corresponding to the projection points are classified into one category;

[0051] A calculation subunit, configured to use the distances between same-type points as vectors and calculate a resultant vector as the vector at the projection point;

[0052] A fitting subunit, configured to fit a feature curve based on the endpoints of the vectors at all projection points.

[0053] The technical solution of the present invention also provides a storage medium, in which at least one program code is stored, and when the program code is loaded and executed by a processor, the multi-dimensional space storage management method of the digital archive is implemented.

[0054] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention scans the archives to obtain digital archives, extracts features from the digital archives, and then classifies and archives them, greatly improving the retention time of data and reducing the data storage space. Description of the Drawings

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention.

[0056] Figure 1 It is a flowchart of the multi-dimensional space storage management method for a digital archive.

[0057] Figure 2 It is a block diagram of the composition structure of the multi-dimensional space storage management system for a digital archive. Detailed Embodiments

[0058] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0059] Figure 1 It is a flowchart of the multi-dimensional space storage management method for a digital archive. In an embodiment of the present invention, a multi-dimensional space storage management method for a digital archive, the method includes:

[0060] Step S100: Receive the archives to be stored, scan the archives to obtain digital archives; the digital archives are scanned image sets;

[0061] Regardless of the type of file, basically there is a paper form. The preservation period of paper-based files is relatively short, so digital processing is carried out and then digital storage is performed. In the technical solution of the present invention, the scenario faced is a scenario with paper-based files. The files to be stored are received and converted into digital files by means of a scanner, obtaining multiple images, which are called a scanned image set.

[0062] Step S200: For any digital file, read the images in sequence, identify the images, locate the text boxes, and extract the feature curve of each image based on the located text boxes.

[0063] For any digital file, read the images in sequence, identify the images, and locate the text boxes. Since the pixels in the file are limited, such as black text on a white background and occasionally some red borders, the process of locating the text boxes is very simple. After locating the text boxes, record the positions and sizes of all the text boxes, and determine the features of the image according to the positions and sizes of each text box. In the technical solution of the present invention, it is represented by a curve, which is called a feature curve. Generally speaking, extract a curve in the image as the feature of the image to obtain the feature curve.

[0064] Step S300: Determine the abnormality degree of any image according to the feature curve, count the abnormality degrees of all the images in the same digital file, and determine the universality of the digital file. Among them, the abnormality degree is used to characterize the degree of difference between the image and other images, and the universality represents the similarity degree between the digital file and other digital files.

[0065] By analyzing the feature curve of the image, the abnormality degree of the image can be calculated. The calculation process of the abnormality degree is to compare the difference between the feature curve of the image and the feature curves of other images to determine a parameter, which is called the abnormality degree. On this basis, a digital file is a set of images. According to the abnormality degrees of the images in the digital file, the universality of the digital file can be determined. The universality represents the similarity degree between the digital file and other digital files, and its practical significance is to represent the common degree.

[0066] Step S400: Archive and store the digital file according to the universality. When archiving, encrypt the image based on the feature curve.

[0067] Classify similar digital files according to the universality. When querying is needed, first use the universality as an index to locate the corresponding category of digital files, and then perform subsequent traversal and retrieval to optimize the retrieval architecture. In addition, when specifically storing, since the feature extraction process is performed on each image, the image can be encrypted according to the feature extraction result to ensure the security of the digital file.

[0068] Regarding step S200, the steps of successively reading images of any digital file, identifying the images, locating text boxes, and extracting the feature curve of each image based on the located text boxes include:

[0069] For any digital file, obtain the image size of each image;

[0070] Sort the images in ascending order according to the image size, and successively read the images in the images sorted in ascending order;

[0071] Input the image into a preset contour recognition model to locate the text box; wherein, the contour recognition model has a performance adjustment port, and the initial performance is the lowest performance. When the location process of any image fails, a preset performance step is added to the lowest performance;

[0072] Extract the feature curve of each image based on the located text box.

[0073] The above content describes the image recognition process and the feature extraction process. Any digital file is a set of images. Obtain the image size of each image, sort the images in ascending order according to the image size, and successively read the images in the images sorted in ascending order. Its practical significance is to read small images first and then large images. The advantage of doing this is that at the beginning, the contour recognition model can operate in a low-performance mode. If the recognition fails, the performance is increased. This process is continuously looped until all images are recognized. In this process, since the image size is increasing, the performance of the contour recognition model is also increasing. Compared with the existing contour recognition model with fixed performance, the resource utilization rate is higher.

[0074] Applying the contour recognition model to recognize the image can locate the text box, count the position and size of each text box, and a curve can be fitted, which is called the feature curve.

[0075] Further, the steps of extracting the feature curve of each image based on the located text box include:

[0076] Select points on the boundary of the text box based on a preset step; the step is the distance between adjacent points;

[0077] For any point, calculate its distance from the diagonal of the image and synchronously determine the projection point;

[0078] Calculate the distance between adjacent projection points on the diagonal. When the distance is less than a preset threshold, all the points corresponding to the projection point are classified into one category;

[0079] Take the distance of the same-category points as vectors, and calculate the resultant vector as the vector at the projection point;

[0080] The feature curve is fitted based on the endpoints of the vectors at all projection points.

[0081] In an example of the technical solution of the present invention, the generation process of the feature curve is specifically described. First, a text box is actually some continuous boundary lines. During analysis, continuous line segments are not easy to process. Therefore, it is necessary to discretize them. The discretization method is to select points on the boundary of the text box based on a preset step size. For the selected points, calculate the distance between them and the diagonal of the image. In fact, there are two diagonals for the image. In the technical solution of the present invention, a prior regulation needs to be made. It is best to use the diagonal in the same direction for all images. For example, from the upper left corner to the lower right corner. Then, calculate the projection points of each selected point on the diagonal. Connect the selected point and the projection point. Taking the projection point as the starting point and the selected point as the ending point, a vector can be obtained. If the projection points of two vectors are close enough, the corresponding vectors are regarded as the same type of vectors and added together to finally obtain a combined vector. Finally, fit the endpoints of all combined vectors, and the obtained curve is the feature curve.

[0082] Regarding step S300, the steps of determining the abnormality degree of any image according to the feature curve, counting the abnormality degrees of all images in the same digital file, and determining the universality degree of the digital file include:

[0083] Record the feature curve of each image in real time to construct a feature curve library;

[0084] For any image, randomly compare the feature curve of this image with the feature curves in the feature curve library and calculate the similarity;

[0085] When the preset comparison jump-out condition is met, stop the comparison; the comparison jump-out conditions include that the similarity reaches a preset similarity threshold and the number of random comparisons reaches a preset number threshold;

[0086] Select the maximum similarity from all comparison results, and determine the abnormality degree according to the maximum similarity and the number of random comparisons;

[0087] Determine the universality degree of the digital file according to the abnormality degree of each image.

[0088] The above content makes limitations on the generation process of the abnormality degree and the universality degree. The characteristic curves of each image are recorded in real time, and a characteristic curve library is constructed. It is not necessary to record in the characteristic curve library which characteristic curve corresponds to which image, which saves storage space and there is no storage pressure for storing the corresponding relationship. For any image, a characteristic curve is randomly read in the characteristic curve library, and the randomly read characteristic curve is compared with the characteristic curve of the current image, and the similarity can be calculated. When the similarity obtained by a certain random selection reaches the threshold, it means that it is similar enough to some curves in the characteristic curve library. At this time, the random selection process (which is itself a loop process) is exited; when the number of random comparison times reaches the preset number threshold, it means that after multiple random selection processes, no sufficiently similar characteristic curves are found. At this time, the random selection process is also exited.

[0089] Finally, the abnormality degree is jointly determined according to the maximum similarity and the number of random comparison times. The greater the maximum similarity, the smaller the abnormality degree, and the greater the number of random comparison times, the greater the abnormality degree; each digital file is an image set, and the universality degree of the digital file can be jointly calculated according to the abnormality degrees of each image in the image set.

[0090] Furthermore, the calculation process of the similarity is as follows:

[0091] Obtain the fitting function of the characteristic curve and calculate the similarity according to the fitting function; ; In the formula, represents the similarity of two characteristic curves, represents the right endpoint of the common segment of the diagonals corresponding to two characteristic curves, represents the left endpoint of the common segment of the diagonals corresponding to two characteristic curves, is generally set to 0; represents the length of the longer diagonal among the diagonals corresponding to two characteristic curves; and respectively represent two characteristic curves; and respectively represent the fitting functions of two characteristic curves;

[0092] The calculation process of the abnormality degree is as follows:

[0093] ; In the formula, represents the abnormality degree, represents the maximum similarity of a certain image in the random comparison process, represents the number of random comparison times, and are preset correction coefficients;

[0094] The calculation process of the universality degree is as follows:

[0095] ; wherein, represents the universality, represents the abnormality degree of the th image, and

[0096] Regarding step S400, the step of archiving and storing the digital archive according to the universality and encrypting the image based on the characteristic curve during archiving includes:

[0097] Archiving and storing the digital archive according to a preset universality range;

[0098] During storage, perform text recognition on the image, locate the text, use the center of the text as a point position, and update the characteristic curve;

[0099] Input the updated characteristic curve into a trained transcoding model to output a cipher and encrypt the image.

[0100] The above content defines the specific storage stage. The management party pre-sets some ranges, called universality ranges. For any digital archive, storing it within the corresponding universality range is the process of archiving and storing. Further, during storage, the technical solution of the present invention can also perform text recognition on the image. After locating the text, use the center of the text as a point position and add it to the process of "selecting point positions on the boundary of the text box based on a preset step size". That is, on the basis of determining the point positions based on the text box, new point positions are determined again based on the text recognition result. After the point positions are updated, the fitted characteristic curve is updated. Since there are more point positions, the characteristic curve becomes more refined, and the difference in the characteristic curve becomes more obvious. The characteristic curves of different images will almost all have slight differences. In other words, the probability of having the same characteristic curve is smaller. Previously, as long as the text boxes were the same, the characteristic curves were the same. Now, the text also needs to be exactly the same for the characteristic curves to be the same. The purpose of refining the characteristic curve is to input the updated characteristic curve into a trained transcoding model to output a cipher and encrypt the image. Thus, each image is encrypted independently.

[0101] Figure 2 It is a block diagram of the composition structure of a multi-dimensional space storage management system for a digital archive. In an embodiment of the present invention, a multi-dimensional space storage management system for a digital archive, the system 10 includes:

[0102] A digital processing module 11, configured to receive the archive to be stored, scan the archive to obtain a digital archive; the digital archive is a scanned image set;

[0103] The feature extraction module 12 is used to sequentially read images for any digital file, identify the images, locate the text boxes, and extract the feature curves of each image based on the located text boxes.

[0104] The image analysis module 13 is used to determine the abnormality degree of any image according to the feature curve, count the abnormality degrees of all images in the same digital file, and determine the universality of the digital file; wherein, the abnormality degree is used to characterize the degree of difference between an image and other images, and the universality represents the similarity degree between the digital file and other digital files.

[0105] The image encryption module 14 is used to archive and store the digital file according to the universality, and encrypt the image based on the feature curve during archiving.

[0106] Furthermore, the feature extraction module 12 includes:

[0107] The image size acquisition unit is used to acquire the image size of each image for any digital file.

[0108] The image sorting unit is used to sort the images in ascending order according to the image size, and sequentially read the images in the images sorted in ascending order.

[0109] The text box positioning unit is used to input the image into a preset contour recognition model to locate the text box; wherein, the contour recognition model contains a performance adjustment port, and the initial performance is the lowest performance. When the positioning process of any image fails, a preset performance step is added to the lowest performance.

[0110] The extraction execution unit is used to extract the feature curve of each image based on the located text box.

[0111] Wherein, the extraction execution unit includes:

[0112] The point selection sub-unit is used to select points on the boundary of the text box based on a preset step; the step is the distance between adjacent points.

[0113] The projection sub-unit is used to calculate the distance between any point and the diagonal of the image, and synchronously determine the projection point.

[0114] The merging sub-unit is used to calculate the distance between adjacent projection points on the diagonal. When the distance is less than a preset threshold, all the points corresponding to the projection points are classified into one category.

[0115] The calculation sub-unit is used to use the distance between the points in the same category as a vector, and calculate the combined vector as the vector at the projection point.

[0116] The fitting sub-unit is used to fit the feature curve according to the endpoints of the vectors at all projection points.

[0117] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi-dimensional space storage management method for a digital archive, characterized in that The method includes: Receiving an archive to be stored, scanning the archive to obtain a digital archive; the digital archive is a set of scanned images; For any digital archive, sequentially reading images, identifying the images, locating text boxes, and extracting the feature curve of each image based on the located text boxes; Determining the abnormality degree of any image according to the feature curve, counting the abnormality degrees of all images in the same digital archive, and determining the universality degree of the digital archive; wherein, the abnormality degree is used to characterize the degree of difference between an image and other images, and the universality degree represents the similarity degree between the digital archive and other digital archives; Archiving and storing the digital archive according to the universality degree, and encrypting the image based on the feature curve during archiving; The step of, for any digital archive, sequentially reading images, identifying the images, locating text boxes, and extracting the feature curve of each image based on the located text boxes includes: For any digital archive, obtaining the image size of each image; Sorting the images in ascending order according to the image size, and sequentially reading the images in the images sorted in ascending order; Inputting the image into a preset contour recognition model to locate the text box; wherein, the contour recognition model has a performance adjustment port, the initial performance is the lowest performance, and when the positioning process of any image fails, a preset performance step is added to the lowest performance; Extracting the feature curve of each image based on the located text box; The step of extracting the feature curve of each image based on the located text box includes: Selecting points on the boundary of the text box based on a preset step length; the step length is the distance between adjacent points; For any point, calculating the distance between it and the diagonal of the image, synchronously determining the projection point, connecting the selected point and the projection point, taking the projection point as the starting point and the selected point as the ending point to obtain the vector corresponding to the point; Calculating the distance between adjacent projection points on the diagonal, and when the distance is less than a preset threshold, classifying all points corresponding to the projection point into one category; Calculating the resultant vector of the vectors corresponding to the points in the same category, and fitting the endpoints of all resultant vectors to obtain the feature curve.

2. The multi-dimensional space storage management method of the digital archive according to claim 1, characterized in that The step of determining the abnormality degree of any image according to the feature curve, counting the abnormality degrees of all images in the same digital archive, and determining the universality degree of the digital archive includes: Recording the feature curve of each image in real time to construct a feature curve library; For any image, randomly comparing the feature curve of the image with the feature curves in the feature curve library and calculating the similarity; When a preset comparison jump-out condition is met, stop the comparison; the comparison jump-out conditions include that the similarity reaches a preset similarity threshold and the number of random comparison times reaches a preset number threshold; Selecting the maximum similarity from all comparison results, and determining the abnormality degree according to the maximum similarity and the number of random comparison times; Determining the universality degree of the digital archive according to the abnormality degree of each image.

3. The multi-dimensional space storage management method of the digital archive according to claim 2, characterized in that, The calculation process of the similarity is: Obtain the fitting function of the characteristic curve and calculate the similarity according to the fitting function; ; where, represents the similarity of two characteristic curves, represents the right endpoint of the common segment of the diagonals corresponding to two characteristic curves, represents the left endpoint of the common segment of the diagonals corresponding to two characteristic curves, is generally set to 0; represents the length of the longer diagonal among the diagonals corresponding to two characteristic curves; and respectively represent two characteristic curves; and respectively represent the fitting functions of two characteristic curves; The calculation process of the abnormality degree is: ; where, represents the abnormality degree, represents the maximum similarity of a certain image during the random comparison process, represents the number of random comparison times, and are preset correction factors; The calculation process of the universality degree is: ; wherein, represents the universality, represents the abnormality of the th image, represents the total number of images in the digital file.

4. The multi-dimensional space storage management method of the digital archive according to claim 1, characterized in that The step of archiving and storing the digital archive according to the universality degree, and encrypting the image based on the feature curve during archiving includes: Archive and store digital archives according to a preset popularity range; During storage, perform text recognition on the image, locate the text, use the center of the text as a point position, and update the feature curve; Input the updated feature curve into the trained transcoding model, output the encryption password, and encrypt the image.

5. A multi-dimensional space storage management system for a digital archive, characterized in that, The system includes: A digital processing module for receiving the archives to be stored, scanning the archives to obtain digital archives; the digital archives are a set of scanned images; A feature extraction module for, for any digital archive, sequentially reading the images, performing recognition on the images, locating the text boxes, and extracting the feature curve of each image based on the located text boxes; An image analysis module for determining the abnormality degree of any image according to the feature curve, counting the abnormality degrees of all images in the same digital archive, and determining the popularity of the digital archive; wherein, the abnormality degree is used to characterize the degree of difference between the image and other images, and the popularity represents the similarity degree between the digital archive and other digital archives; An image encryption module for archiving and storing digital archives according to the popularity, and encrypting the images based on the feature curve during archiving; The feature extraction module includes: An image size acquisition unit for, for any digital archive, acquiring the image size of each image; An image sorting unit for sorting the images in ascending order according to the image size, and sequentially reading the images in the images sorted in ascending order; A text box positioning unit for inputting the image into a preset contour recognition model to locate the text box; wherein, the contour recognition model has a performance adjustment port, the initial performance is the lowest performance, and when the positioning process of any image fails, a preset performance step is added to the lowest performance; An extraction execution unit for extracting the feature curve of each image based on the located text box; The extraction execution unit includes: A point position selection sub-unit for selecting point positions on the boundary of the text box based on a preset step length; the step length is the distance between adjacent point positions; A projection sub-unit for, for any point position, calculating the distance between it and the diagonal of the image, and synchronously determining the projection point; connecting the selected point position and the projection point, with the projection point as the starting point and the selected point position as the ending point, to obtain the vector corresponding to the point position; A merging sub-unit for calculating the distance between adjacent projection points on the diagonal, and when the distance is less than a preset threshold, classifying all point positions corresponding to the projection point into one category; A fitting sub-unit for calculating the combined vector of the vectors corresponding to the point positions in the same category, and fitting the ending points of all combined vectors to obtain the feature curve.

6. A storage medium, characterized in that, At least one program code is stored in the medium, and when the program code is loaded and executed by a processor, it implements the multi-dimensional space storage management method of the digital archive as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Character offset detection method and system

    CN111680692A

  • Digital management method and system based on enterprise archives and storage medium

    CN116663549A