Blockchain-based Archival Management Method
Through the blockchain-based archive management method, the problems of high cost and low efficiency of paper archive management are solved, efficient storage and query of archives are achieved, management costs are reduced and utilization efficiency is improved.
Patent Information
- Application Number
- CN202210212848.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-06
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-03-06
AI Technical Summary
Existing paper archives are costly and have low efficiency, difficult to quickly view and will hardly be reused after storage.
The blockchain-based archive management method is adopted, and the distributed storage and query of archives are realized by receiving and scanning paper archives, generating archive numbers and retrieving key values, extracting hash values and uploading blockchain storage, establishing archive indexes, and realizing distributed storage and querying of archives.
It reduces the cost of archive management and improves the efficiency of archive utilization. Through blockchain evidence storage, scanned copies have legal effect, easy to view, and improves security and storage efficiency through distributed storage.
Smart Images

Figure CN114610778B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and particularly to a blockchain-based file management method. Background Art
[0002] The content of file management includes activities such as file collection, sorting, storage, appraisal, statistics, and utilization provision, such as file collection, file sorting, file value appraisal, file storage, file cataloging and retrieval, file statistics, file editing and research, and file utilization provision. File management institutions store a large number of paper files, which are characterized by being scattered, messy, of poor quality, large in quantity, and single copies. Paper files need to be stored in a special place, not only to prevent the oxidation and deterioration of the paper, but also to ensure fire safety, which consumes a large amount of administrative funds. A large number of paper files stored at great expense are difficult to quickly access due to the vast number of volumes and are rarely reused after being stored. However, some of the voucher materials involved must be stored for a certain number of years according to regulations. Although the electronicization of paper files helps to improve the search efficiency, the electronic files are easily forged and do not have legal effect, and the paper files still need to be stored. To solve the financial and management burdens brought by paper files, new file management solutions need to be studied. Summary of the Invention
[0003] The technical problem to be solved by the present invention is the technical problem of high cost and low efficiency in current paper file management. A blockchain-based file management method is proposed, which can reduce the file management cost and improve the file utilization efficiency.
[0004] To solve the above technical problem, the technical solution adopted by the present invention is: A blockchain-based file management method, including: receiving batch paper files and batch file information, where the batch file information includes the file source department, file type, file creation time, and retrieval fields; sequentially scanning the paper files to obtain scanned copies of the paper files, and assigning an archive number to the scanned copies; filling in the retrieval key values of the scanned copies, where the retrieval key values are the values of the retrieval fields in the current scanned copies, and storing the scanned copies; extracting the hash values of the scanned copies, uploading the hash values to the blockchain for storage, and obtaining the corresponding block heights; establishing a file index, where the file index records the archive number, file source department, file type, file creation time, retrieval key values, hash values, and block heights; receiving the identity authentication information and file query requests sent by the requester, where the file query requests include several key values; after verifying the identity authentication information of the requester, retrieving the file index, and if there are scanned copies that match the file query requests, sending the associated hash values and block heights of the matching scanned copies to the requester.
[0005] Preferably, a number of storage nodes are established, the scanned document is sliced into a number of slice regions, the slice regions are numbered, and the number of slice regions are respectively sent to a number of storage nodes for storage. When reading the scanned document, the slice regions are respectively read from the number of storage nodes.
[0006] Preferably, the hash value of each slice region is extracted and denoted as the slice hash value; the hash value of the scanned document is denoted as the file hash value, and the hash values of the file hash value and all slice hash values are extracted together and denoted as the evidence storage hash value; the evidence storage hash value is uploaded to the blockchain for storage to obtain the corresponding block height; the storage node stores each slice region associated with the number, the file hash value, and all slice hash values; when reading the scanned document, the slice regions are respectively read from a number of storage nodes; after verifying the slice hash value of each slice region, the slice regions are arranged according to the number to restore the scanned document.
[0007] Preferably, the scanned documents of the same type of paper files are sliced at the same slicing position and numbered for the slice regions in the same order. The slice regions with the same number are sent to the same storage node for storage; the storage node extracts the average pixel value of each pixel position of a number of slice regions of the same type and the same number; the common slice is constructed using the average pixel value; the pixel difference between each slice region and the common slice at each pixel position is calculated; the pixel difference is represented using a shorter byte length, so that the slice region is represented using a shorter byte length; the exception set is used to store the exception pixels that exceed the representation range of the shortened byte length, and the exception set records the pixel coordinates and pixel values of the exception pixels; when reading the scanned document, after the storage node superimposes the saved pixel difference and the common slice, the pixel value recorded in the exception set is used to replace the pixel value at the corresponding pixel coordinate to restore the slice region.
[0008] Preferably, the scanned copies of the same type of paper archives are sliced at the same slicing position so that each sliced area contains at most one field area. The field area is the area where the fields and filling information are recorded on the paper archives. The field names of the filling fields included in the sliced area are stored. The sliced areas are numbered in the same order, and the sliced areas with the same number are sent to the same storage node for storage. The storage node extracts the average pixel value at each pixel position of several sliced areas of the same type and the same number. The common slice is constructed using the average pixel value. The pixel difference between each sliced area and the common slice at each pixel position is calculated. The pixel difference is represented using a shorter byte length, so that the sliced area can be represented using a shorter byte length. The exception set is used to store the exception pixels that exceed the representation range of the shortened byte length. The exception set records the pixel coordinates and pixel values of the exception pixels. The archive query request includes the requested field name. When reading the scanned copy, the storage node storing the sliced area corresponding to the requested field name superimposes the saved pixel difference and the common slice. The pixel value recorded in the exception set is used to replace the pixel value at the corresponding pixel coordinate to restore the sliced area, and the restored sliced area is fed back. The storage node that does not store the sliced area corresponding to the requested field name feeds back the common slice. The restored sliced areas and the common slice are spliced into the scanned copy and sent to the requester.
[0009] Preferably, a permission table is generated. The permission table records the permission level of the requester and the permission level allowed to view the field name. After verifying the identity authentication information of the requester, it is verified whether the permission level of the requester meets the permission level allowed to view the requested field name. If it does not meet the permission level allowed to view the requested field name, the corresponding storage node feeds back the common slice.
[0010] Preferably, the method for the requester to send the identity authentication information and the archive query request includes: allocating a pair of public and private keys to the requester, and allocating an identity identifier to the requester. The identity identifier is the public key. The requester generates an archive query request, generates a transaction for transferring several tokens to a preset address, uses the archive query request as the transaction attachment information, and uploads the transaction to the blockchain. Poll the blockchain to obtain the archive query request. The signature of the transaction is the identity authentication information, and verifying the signature of the transaction can verify the identity identifier of the requester.
[0011] The substantial effect of the present invention is that by using the blockchain to store the scanned copy, the scanned copy has legal effect, so there is no longer a need to save the paper archives, which not only facilitates the access to the archives, but also reduces the cost of archive management; through distributed storage, the storage security is improved; using sliced areas for storage reduces the storage space occupied by the scanned copy. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a schematic diagram of the archive management method for Embodiment 1.
[0013] Figure 2 Schematic diagram of the sliced storage method for the scanned copy of Example 1.
[0014] Figure 3 Schematic diagram of the method for establishing shared slices in Example 1.
[0015] Figure 4 Schematic diagram of the method for establishing a sliced area in Example 1.
[0016] Figure 5 Schematic diagram of the method for establishing a sliced area in Example 2.
[0017] Figure 6 Schematic diagram of the method for sending an archive query request in Example 2. Specific implementation manners
[0018] The following will further specifically describe the specific implementation manners of the present invention through specific examples and in conjunction with the accompanying drawings.
[0019] Example 1:
[0020] For the archive management method based on blockchain, please refer to the attached Figure 1, including: Step A01) Receive batch paper archives and batch archive information, where the batch archive information includes the archive source department, archive type, archive creation time, and retrieval fields; Step A02) Scan the paper archives in sequence to obtain scanned copies of the paper archives, and assign an archive number to the scanned copies; Step A03) Fill in the retrieval key values of the scanned copies, where the retrieval key values are the values of the retrieval fields in the current scanned copies, and store the scanned copies; Step A04) Extract the hash values of the scanned copies, upload the hash values to the blockchain for storage, and obtain the corresponding block height; Step A05) Establish an archive index, where the archive index records the archive number, archive source department, archive type, archive creation time, retrieval key values, hash values, and block height; Step A06) Receive the authentication information and archive query request sent by the requester, where the archive query request includes several key values; Step A07) After verifying the requester's authentication information, retrieve the archive index. If there are scanned copies that match the archive query request, send the associated hash values and block height of the matching scanned copies to the requester. Archive types such as academic record archives, ID card archives, marriage archives, or real estate archives, etc., record the livelihood data information of the masses. The archive source department is the corresponding government service department. Currently, the materials and archives of people's livelihood institutions are uniformly handed over to the archives for storage and management. The archives need to consume a large amount of funds every year for digitizing paper archives and storing paper archives. Through this technical solution, the scanned copies are authenticated using the blockchain, making the scanned copies legally valid, so there is no longer a need to store paper archives, saving funds. At the same time, an archive index is established, facilitating the search and use of archives. The requesters who send archive query requests are preferably relevant government people's livelihood service departments or service windows, as well as the masses themselves. After the people's livelihood service window queries the archives, it can retrieve the relevant archives of the people handling affairs, eliminating the need for the masses to provide relevant archives, which is convenient for the people handling affairs.
[0021] Establish several storage nodes, divide the scanned copy into several slice areas, number the slice areas, and send the several slice areas to several storage nodes for storage respectively. When reading the scanned copy, read the slice areas from several storage nodes respectively. Provide a distributed storage method, divide the scanned copy into several parts and store them separately, which can improve the security of the scanned copy and avoid leakage during storage.
[0022] Please refer to the appendix Figure 2, the method for storing scanned document slices includes: Step B01) Extract the hash value of each slice area, denoted as the slice hash value; Step B02) Denote the hash value of the scanned document as the archive hash value, and extract the hash value of the archive hash value and all slice hash values together, denoted as the deposit hash value; Step B03) Upload the deposit hash value to the blockchain for storage to obtain the corresponding block height; Step B04) The storage node stores the associated number, archive hash value, and all slice hash values of each slice area; Step B05) When reading the scanned document, read the slice area from several storage nodes respectively; Step B06) After verifying the slice hash value of each slice area, arrange the slice areas according to the number to restore the scanned document. Generate the corresponding slice hash value for each slice area so that each slice area can obtain a deposit. When restoring from different storage nodes, it can be verified whether the slice area provided by each storage node has been tampered with. Even if some slice areas are damaged, the remaining slice areas can also be proven to be true and reliable and form validity.
[0023] This embodiment provides a solution for compressing the storage space occupied by the scanned document to further reduce the archive storage cost. Please refer to the appendix Figure 3, including: Step C01) Splitting the scanned copies of the same type of paper files at the same splitting positions, numbering the sliced areas in the same order, and sending the sliced areas with the same number to the same storage node for storage; Step C02) The storage node extracts the average pixel values of several sliced areas of the same type and with the same number at each pixel position; Step C03) Using the average pixel values to construct a shared slice; Step C04) Calculating the pixel differences between each sliced area and the shared slice at each pixel position; Step C05) Representing the pixel differences with a shorter byte length, so as to represent the sliced areas with a shorter byte length; Step C06) Using an exception set to store the exception pixels that exceed the range represented by the shortened byte length, and the exception set records the pixel coordinates and pixel values of the exception pixels; Step C07) When reading the scanned copy, the storage node adds the saved pixel differences and the shared slice, and replaces the pixels at the corresponding pixel coordinates with the pixel values recorded in the exception set to restore the sliced area. A large number of paper application forms or qualification certificate documents received by a given window have formatted terms, and only basic information needs to be filled in and signed at a small number of specified positions. Therefore, most of the pixel values in such scanned copies are the same, without considering the differences in scanning devices and lighting. For example, various informed consent forms have the same recorded content, and the public needs to sign at the end. Such informed consent forms are important documents for liability determination in post-dispute cases and need to be stored for a preset duration before being destroyed. For example, the scanned copy is stored in the form of a picture, and the pixels of the picture are represented using the RGB color system. Each pixel uses 3 channels, and the channel value of each channel is represented by 1 byte, that is, the value range of each channel is [0, 255], and the set of the three channel values constitutes the pixel value. Each pixel value occupies 3 bytes. Since the scanned copies of a large number of the same type of files are highly similar in the same area. Therefore, the difference between the pixel values of each scanned copy and the shared slice at the same pixel position will be small. Half a byte is sufficient to represent it. That is, the value range of the difference of each channel is [0, 15], plus 1 bit to represent the positive or negative sign of the difference. The formation of the difference is mainly due to the hardware differences during each scan of the scanning device. The differences caused by the placement positions of the files are not discussed in this implementation. Conventional techniques should be used for cropping and alignment, or a scanning device with an alignment function should be used. If conventional techniques are used to perform white balance on the scanned copy, better technical effects can be achieved. For the pixel points where the difference exceeds [0, 15], the exception set method is used to separately record such pixel points. Since the shared slice only needs to be stored once, most pixels can be stored using pixel differences, and the space occupied by each pixel is reduced by half. Although a small number of pixels need to occupy the space of the exception set, for the same area of the same type of files, the number of exception pixels should be small, so that less storage space is occupied overall. The solution provided by this embodiment reduces the storage space occupied by the sliced areas by means of the shared slice, pixel differences, and exception set.
[0024] The beneficial technical effects of this embodiment are as follows: By using the blockchain to deposit the scanned documents, the scanned documents are given legal effect, so that there is no longer a need to save paper archives, which not only facilitates the access to archives but also reduces the cost of archive management; Through distributed storage, the security of storage is improved; Using sliced areas for storage reduces the storage space occupied by the scanned documents.
[0025] Embodiment 2:
[0026] Based on the blockchain-based archive management method, this embodiment provides a new solution for compressing the storage space occupied by scanned documents on the basis of Embodiment 1. Please refer to the appendix Figure 4, including: Step D01) Split the scanned copies of paper archives of the same type at the same splitting position so that the sliced area contains at most one field area. The field area is the area where the fields and filling information are recorded in the paper archive, and store the field names of the filling fields contained in the sliced area; Step D02) Number the sliced areas in the same order, and send the sliced areas with the same number to the same storage node for storage; Step D03) The storage node extracts the average pixel value at each pixel position of several sliced areas of the same type and the same number; Step D04) Use the average pixel value to construct a common slice; Step D05) Calculate the pixel difference between each sliced area and the common slice at each pixel position; Step D06) Use a shorter byte length to represent the pixel difference, so as to use a shorter byte length to represent the sliced area; Step D07) Use an exception set to store the exception pixels that exceed the representation range of the shortened byte length. The exception set records the pixel coordinates and pixel values of the exception pixels; Step D08) The archive query request includes the requested field name. When reading the scanned copy, the storage node that stores the sliced area corresponding to the requested field name will superimpose the saved pixel difference and the common slice; Step D09) Replace the pixel value at the corresponding pixel coordinate with the pixel value recorded in the exception set, restore the sliced area, and feedback the restored sliced area; Step D10) The storage node that does not store the sliced area corresponding to the requested field name feedbacks the common slice; Step D11) Stitch the restored sliced area and the common slice into a scanned copy and send it to the requester. Each sliced area contains at most one field area, that is, there are some sliced areas whose covered range is the area of the standard terms. Such sliced areas do not contain field areas, that is, the difference between the pixel values and the common slice will be very small, and in the case of good scanning, the exception set should be empty. Therefore, a large amount of storage space will be saved. Other sliced areas will only contain one field area, and there are typed or handwritten words in the field area. Both typed and handwritten words record the information corresponding to the archive person, so there are great differences between different archives. At this time, the exception set will have more content. However, since the content and handwriting only occupy a limited number of pixels, storage space can still be saved. For some people's livelihood archives, if photos are pasted or a large amount of text is filled in a certain column, this type of archive should be manually marked and saved in a conventional manner.
[0027] This embodiment splits the multiple field areas recorded in the scanned copy into different sliced areas. On this basis, a technical solution for providing different visible areas of the scanned copy for the requester according to permissions is provided. Please refer to the appendix Figure 5, including: Step E01) Generate a permission table, which records the permission level of the requester and the permission level for viewing the field name; Step E02) After verifying the authentication information of the requester, verify whether the permission level of the requester meets the permission level for viewing the requested field name; Step E03) If it does not meet the permission level for viewing the requested field name, the corresponding storage node feeds back the shared slice. For example, an ID card file records name, gender, ethnicity, birthday, ID card number, and address. A requester with a higher permission level, such as a government civil affairs service window, can obtain the complete ID card file. If the requester has a lower permission level, such as the anti-addiction system of an online game, when verifying whether a game login user is 18 years old and calling the ID card file, only the name, ID card number, and birthday need to be returned for comparison and verification with the name and ID card number provided by the login user. Keep the gender, ethnicity, and address confidential, which is difficult to achieve with current conventional technical solutions. Moreover, the provided name, ID card number, and birthday can all be verified through the corresponding slice hash values. Specifically, when implementing, the public security organ can set up a separate server dedicated to the verification of the online game anti-addiction system, receive the name and ID card number sent by the online game server, and then send a corresponding file query request to the server running this file management method. Set a lower permission level for this server in advance, and it can only view the name, ID card number, and birthday. Therefore, the server running this file management method will return the slice area corresponding to the name, ID card number, and birthday, all slice hash values, file hash values, and deposit hash values. This server dedicated to the verification of the online game anti-addiction system extracts the hash value of the slice area containing the name, ID card number, and birthday and compares it with the slice hash value. Then compare all slice hash values and file hash values with the deposit hash value, and finally compare the deposit hash value with the one stored on the blockchain to confirm the authenticity of the obtained slice area containing the name, ID card number, and birthday, and avoid the situation of being tampered with or damaged during network transmission. Feed back the verification result to the online game server to complete the anti-addiction verification process. During the verification process, the gender, ethnicity, and address of the ID card file do not need to be transmitted over the network, so it is impossible to be eavesdropped and leaked during this process, improving the security of file use.
[0028] Please refer to the appendix Figure 6, the method for the requester to send authentication information and an archive query request includes: Step F01) allocating a pair of public and private keys to the requester and allocating an identity identifier to the requester, where the identity identifier is the public key; Step F02) the requester generates an archive query request, generates a transaction for transferring a number of tokens to a preset address, uses the archive query request as attached information to the transaction, and uploads the transaction to the blockchain; Step F03) polling the blockchain to obtain the archive query request; Step F04) the signature of the transaction is the authentication information, and verifying the signature of the transaction can verify the identity identifier of the requester. Leaving a trace on the blockchain can facilitate and quickly view the situation of the archive history being requested for viewing.
[0029] Compared with Embodiment 1, when providing a scanned copy in this embodiment, whether to provide a complete scanned copy or only provide the field area that conforms to the permission level is determined according to the permission level, which can better protect the privacy of the archive.
[0030] The above-described embodiments are only a preferred solution of the present invention and do not impose any form of limitation on the present invention. There are other variations and modifications without exceeding the technical solutions recorded in the claims.
Claims
1. A blockchain-based file management method, characterized in that, it includes: Receiving batch paper files and batch file information, where the batch file information includes the file source department, file type, file creation time, and retrieval fields; Sequentially scanning the paper files to obtain scanned copies of the paper files, and assigning an archive number to the scanned copies; Filling in the retrieval key values of the scanned copies, where the retrieval key values are the values of the retrieval fields in the current scanned copies, and storing the scanned copies; Extracting the hash value of the scanned copy, uploading the hash value to the blockchain for storage, and obtaining the corresponding block height; Establishing a file index, where the file index records the archive number, file source department, file type, file creation time, retrieval key value, hash value, and block height; Receiving the authentication information and file query request sent by the requester, where the file query request includes several key values; After verifying the requester's authentication information, retrieving the file index. If there is a scanned copy that matches the file query request, then sending the associated hash value and block height of the matching scanned copy to the requester; Establishing several storage nodes, splitting the scanned copy into several slice regions, numbering the slice regions, and sending the several slice regions to several storage nodes for storage respectively. When reading the scanned copy, reading the slice regions from several storage nodes respectively; Splitting the scanned copies of the same type of paper files at the same splitting positions, numbering the slice regions in the same order, and sending the slice regions with the same number to the same storage node for storage; The storage node extracts the pixel value mean at each pixel position of several slice regions of the same type and the same number; Using the pixel value mean to construct a common slice; Calculating the pixel difference between each slice region and the common slice at each pixel position; Using a shorter byte length to represent the pixel difference, so as to use a shorter byte length to represent the slice region; Using an exception set to store the exception pixels that exceed the representation range of the shortened byte length, where the exception set records the pixel coordinates and pixel values of the exception pixels; When reading the scanned copy, the storage node superimposes the saved pixel difference and the common slice, and replaces the pixel value at the corresponding pixel coordinate with the pixel value recorded in the exception set to restore the slice region.
2. The blockchain-based file management method according to claim 1, characterized in that, Extracting the hash value of each slice region, denoted as the slice hash value; The hash value of the scanned copy is denoted as the file hash value, and the hash value of the file hash value and all slice hash values is extracted together, denoted as the deposit hash value; Uploading the deposit hash value to the blockchain for storage, and obtaining the corresponding block height; The storage node stores each slice region associated with the number, file hash value, and all slice hash values; When reading the scanned copy, reading the slice regions from several storage nodes respectively; After verifying the slice hash value of each slice region, arranging the slice regions according to the number to restore the scanned copy.
3. The blockchain-based file management method according to claim 1 or 2, characterized in that, Slice the scanned copies of paper archives of the same type at the same slicing positions so that each sliced area contains at most one field area, where the field area is the area for filling in fields and information in the paper archives, and store the field names of the filled fields contained in the sliced areas; Number the sliced areas in the same order, and send the sliced areas with the same number to the same storage node for storage; The storage node extracts the average pixel value at each pixel position of several sliced areas of the same type and the same number; Use the average pixel value to construct a common slice; Calculate the pixel difference between each sliced area and the common slice at each pixel position; Use a shorter byte length to represent the pixel difference, so as to use a shorter byte length to represent the sliced area; Use an exception set to store the exception pixels that exceed the representation range of the shortened byte length. The exception set records the pixel coordinates and pixel values of the exception pixels; The archive query request includes the requested field name. When reading the scanned copy, the storage node storing the sliced area corresponding to the requested field name will superimpose the saved pixel difference and the common slice; Use the pixel values recorded in the exception set to replace the pixel values at the corresponding pixel coordinates, restore the sliced area, and feedback the restored sliced area; The storage node that does not store the sliced area corresponding to the requested field name feedbacks the common slice; Stitch the restored sliced area and the common slice into a scanned copy and send it to the requester.
4. The blockchain-based archive management method according to claim 3, wherein, Generate a permission table, which records the permission level of the requester and the permission level for viewing the field name; After verifying the identity authentication information of the requester, verify whether the permission level of the requester meets the permission level for viewing the requested field name; If it does not meet the permission level for viewing the requested field name, the corresponding storage node feedbacks the common slice.
5. The blockchain-based archive management method according to claim 1 or 2, wherein, The method for the requester to send identity authentication information and an archive query request includes: Allocate a pair of public and private keys to the requester and assign an identity identifier to the requester, where the identity identifier is the public key; The requester generates an archive query request, generates a transaction for transferring several tokens to a preset address, uses the archive query request as the attached information of the transaction, and uploads the transaction to the blockchain; Poll the blockchain to obtain the archive query request; The signature of the transaction is the identity authentication information, and verifying the signature of the transaction can verify the identity identifier of the requester.
Citation Information
Patent Citations
Archive management method based on block chain
CN111611460A
Notarization file information processing method, system and platform, equipment and storage medium
CN112163241A
Block chain evidence storage method and system based on isomorphic multi-chain architecture
CN113326317A