Method and system for storage and access of data files
The data file structure encrypts documents with unique key pairs and scatters blocks across a distributed ledger to secure and efficiently manage sensitive documents, addressing unauthorized access and manual review inefficiencies.
Patent Information
- Application Number
- GB2024001076
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-26
- Publication Date
- 2025-08-06
AI Technical Summary
Existing document transfer methods allow unauthorized access and inefficient handling of sensitive documents, particularly in legal transactions, due to insufficient permissions and manual review processes, leading to potential exposure of confidential information.
A data file structure that encrypts documents using unique key pairs, scatters blocks across a distributed centralised ledger, and uses a file identifier and expiry indicator to ensure secure, selective access by third parties.
Enhances security and efficiency by preventing unauthorized access and automating document verification, ensuring confidentiality and reducing manual intervention.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
The present invention relates to a data file structure, particularly to a data file structure for storing and sending confidential files. Background In recent times, the circulation of documents has shifted from paper to electronic. Documents that require circulation may contain confidential or sensitive information, particularly in the case where the documents relate to legal transactions, such as the buying and selling of property. For example, a potential buyer or seller will need to have multiple documents containing sensitive or confidential information verified by a third party. Currently, documents are typically transferred by being uploaded to a server through a corresponding network protocol, the document is then given open permissions which determine who can then download the file. If the permission of a user is not sufficient, the corresponding file cannot be downloaded. This document transfer method allows most users to see the uploaded files even if they cannot download them. In addition, there are many work arounds for the permissions and the lack of approval steps can lead to unauthorised access and exposure of confidential information. The documents in this case will also likely be handled by an intermediary, such as an estate agent, who will disseminate the documents as and when needed. The estate agents can be inefficient and costly depending on the property. In addition, the documents are reviewed manually each and every time they are needed to determine the type and verification status of each document. Multiple sensitive documents could potentially be viewed altogether by a number of parties, even when those parties only required access to one or some, but not all, of the documents. It is an aim of the invention to provide a data file structure that mitigates one or more of the above-mentioned problems. Statements of invention According to a first aspect of the invention there is provided a data file store for storing multiple collections of files, each collection of files being provided by a user and containing information specific to that user, the data file store comprising; a plurality of data blocks, each data block comprising an encrypted file from the collection of files for a user and a corresponding file identifier, each encrypted file being encrypted using a separate key pair comprising a private key and a public key, the file identifier for each encrypted file comprising an element from the encrypted file’s public key, wherein the data blocks are scattered in different locations in the data file store and individual files of the collection of files are retrievable by using the element of the public key to determine to which of the collection of files the public key corresponds. Each collection of files may comprise between two and thirty documents, e.g. between two and twenty documents or between two and ten documents. The invention is particularly useful for relatively small numbers of documents in each collection. The element of the public key may be an alphanumeric string, such as a number. The element of the public key may be a prime number. The element may be unique within the files on the data store. The file identifier may further comprise a second element. The second element may comprise an alphanumeric string, such as a number. The second element may be a prime number. The second element may be unique within the files on the file store. Each block may further comprise an encrypted file expiry indicator. The encrypted file expiry indicator may comprise an indicator identifying how long the file can be accessed for. The file expiry indicator may comprise an expiry date for the file. The file may no longer be decrypted or accessed beyond the expiry date, e.g. rendering it unusable. The file identifier and / or file expiry indicator in each block may be encrypted. The file identifier and file expiry indicator in each block may be encrypted together. The file identifier, file expiry indicator and collection ID may be encrypted together. The blocks may be arranged in any order. The blocks do not have to be arranged in sequence and typically would not be arranged in sequence. Therefore the storage location of the files in the data file store is not indicative of any relationship between the files. The scattered nature of the files for a collection means that the collection of files is not readily apparent without knowledge of the collection. The blocks for multiple collections may be intermingled. The blocks of one collection of files are typically scattered with blocks of multiple other collections of files. As such any collection is obfuscated in the date file store, including the individual files making up the collection. The collection of files may be identifiable by a collection ID, e.g. a case ID. The collection ID may or may not be contained in a block. The data file store may comprise a data file server. A collection of documents may relate to a common transaction, e.g. documentation in support of a transaction. The data file structure may comprise a distributed centralised ledger. The data file structure may comprise a scattered centralised ledger. According to a second aspect of the invention there is provided a method of storing a collection of files containing sensitive information for a user to permit subsequent, selective access to individual files of the collection by a third party, the method comprising: uploading a collection of files to an encryption server, encrypting each file of the collection using a separate key pair comprising a private key and a public key, creating a file identifier for each encrypted file using an element from the encrypted file’s public key, each encrypted file and a corresponding file identifier for the collection of files being stored as a block, storing each block in a data file store comprising multiple further collections of files whereby the blocks of the collection of files are scattered and are retrievable by using the element of the public key to determine to which of the collection of files the public key corresponds. A collection ID may be assigned to the collection of files, e.g. when they are uploaded. The files of the collection are typically documents, e.g. sensitive documents. The collection may be uploaded via a gateway. Metadata associated with the collection of files may be logged, e.g. in a database which may be separate from the data file store. According to a third aspect there may be provided a system for implementing the method of the second aspect whereby the encryption server comprises the data file store and a key store. According to a further aspect, there is provided a method or system for managing access to individual documents in a collection of documents uploaded by a user, comprising the method of the second aspect or system of the third aspect, wherein a content access management system manages the dissemination of individual files in response to received requests. The received request may comprise the public key for an encrypted file in the data file store. A collection of encrypted files may be identified from / by the request. The location of the individual files of the collection of files may be identified. The public key may be used to identify the individual file of the collection of files to which public key applies, e.g. by matching of the key element of the public key with the key element of the corresponding file identifier. The blocks may be scattered in the data file storage or database structure of the data file server. Detailed description Practicable embodiments of the invention are described in further detail below with reference to the accompanying drawings, of which: Figure 1 shows an overview of a system 100 for storage and access of files that contain sensitive information in accordance with an example of the invention. Figure 2 shows an overview of a method 200 for storage and access of files that contain sensitive information in accordance with an example of the invention. Figure 3 shows the file storage server structure in accordance with an example of the invention. Overview of the system Figure 1 shows an overview of a system 100 for storage and access of files that contain sensitive information. In the examples of the invention described herein, the documents / files relate to a property transaction and may contain sensitive information regarding buyers and / or sellers of property. The documents / files for this type of transaction may require access by numerous parties throughout the transaction for the purpose of verification. The system 100 comprises a client device 101 for uploading documents / files 102. A client would use the client device 101 for uploading document / files 102. The client is the first of three types of users of the system. The client, or category one user, is a person who provides the documents / files for storage and access. In the present example, the client would be the person who is buying or selling one or a plurality of properties. The number of documents / files 102 for uploading can vary based on the requirements of the transaction. The system is primarily intended for use with no more than ten documents / files 102. In certain examples, the number of documents / files 102 may be less than or equal to five, or ten files. The system according to the present invention is primarily intended for collections of files that are collated manually by a user, i.e. rather than larger numbers of files amassed by a computer programme. The system 100 may be used for a single document / file 102. The documents / files 102 uploaded by a user are specific to the user. The documents / files 102 uploaded by a user are specific to the transaction. The documents / files 102 may comprise any type of documentation which contains sensitive or confidential information about the buyer, seller, or property. The documents / files 102 may comprise proof of sale. The documents / files 102 may comprise personal information about the client. The documents / files 102 may comprise any or any combination of the proof of sale, council approval, client’s passport, driving license, proof of address (such as bills or council tax information), bank account / statement, birth certificate and / or marriage certificate. The files may comprise document and / or image files, e.g. comprising images of any of the relevant documents. The files could comprise other file formats containing data representative of the information contained in any of the aforementioned documents. The client device 101 may comprise any suitable device for uploading the documents / files 102. The client device 101 may be a conventional computing device such as a desktop / laptop PC, a server, a smartphone, tablet or other conventional computing device. The files are uploaded to a gateway 103. The gateway 103 may be a public but secure gateway. The gateway may be a Public API Gateway. The gateway 103 acts as a single point of entry for documents / files 102 into the system. The documents / files 102 may be stored on the device 101, e.g. in a nonvolatile memory device thereof. Alternatively, the documents / files 102 may be accessible to the client device 101 for the purpose of transmitting them to, or allowing access to them. For example, the documents / filed 102 could be stored in a cloud storage facility and the device 101 could send the files form the cloud storage facility or allow the access to the storage facility to retrieve or modify / process the files. The system further comprises a database 104 for storing metadata 105 each of the documents / files 102. The metadata / metainformation provides information about the data within each document / file 102 without disclosing the data itself. The metadata may include descriptive information about the data, characteristics as what the document is e.g., who created it and when it was created, or a combination of descriptive and characteristic metadata. The metadate further comprises a case ID. Each transaction is given a case ID. The case ID is specific to each transaction. The case ID is specific to the documents / files uploaded by the user for each transaction. Once the metadata 105 is stored in the database 104, the gateway 103 sends the documents / file 102 to a first server 106. The first server 106 is an encryption server. The encryption server 106 creates encrypted documents 107 by generating a key pair 108 for each document / file 102. The key pair 108 comprises a private key 109 and a public key 110. The public key 110 may also be referred to as an access key 109. The encryption server 106 sends the encrypted documents 107 to a second server. The second server comprises a file storage server 111. The file storage server 111 may comprise a distributed ledger technology (DLT). The file storage may comprise a scattered centralised ledger. The private key 109 is sent to a third server. The third server comprises a key storage server 112. The system further comprises an access context manager 113 and an access administrator 114. The access context manager 113 allows enterprises to configure access levels which map to a policy defined on request attributes. The access administrator 114 is the second of the three types of users of the system. The third type of user is the verification agents 115, these are the people responsible for verifying the legal documents. For example, the verification agents may be lawyers or estate agents who are required to verify the documents / files 102 during the transaction. The access administrator is the person that provides access to one or a plurality of the documents / files 102 to the verification agents 115. The access administrators 114 do not have access to documents / files 102 themselves, they can only enable access to others. The access administrators 114 cannot obtain private keys 109. Method of encryption and searching Figure 2 shows an overview of a method 200 for storage and access of files that contain sensitive information. Files / documents 102 are uploaded to a gateway 103. Metadata 105 for each of the documents / files 102 is then stored on a database 104. The metadata 105 further comprises the case ID corresponding to the documents / files 102 uploaded for each transaction. The documents / files 102 are sent to the first server, the encryption server 106. Within the encryption server 106 each of the documents / files 102 undergoes and encryption process. Each document / file is encrypted separately. The encryption server 106 generates a separate key pair 108 for each document / file 102. The key pair 108 comprises a pair of mathematically-related keys. A document / file 102 that is encrypted with the private key 109 must be decrypted with the public key 110, and a document / file 102 that is encrypted with the public key 110 must be decrypted with the private key 109. Once encrypted each document / file 102 is given an identifier. The identifier is saved with the encrypted file and used to locate the file; the process is as follows: 1. A random prime number is generated and multiplied with the prime number from the public key 110 to create a document / file identifier. Examples of document / file identifiers or ID are shown below. File ID Proof of Sale 391 Council Approval 899 2. The document / file is provided with an expiry date at which the document / file can no longer be accessed. The expiry date, the case ID, and the document / file identifier are encrypted together with the public key 110. 3. The encrypted file expiry date, case ID, and file identifier are stored with the document / file 102. The encrypted document / file is stored in the file storage 111 with the encrypted file expiry date (EXP) and the file identifier (ID) in blocks. An example of a case with three blocks are shown below: Block 1: Block 2: HASH (EXP+CASE ID+ID) CONTENT HASH (EXP+CASE ID+ID) CONTENT HASH of Block 2 HASH of Block 1 HASH of Block 3 HASH of Block 3 Block 3: HASH (EXP+CASE ID+ID) CONTENT HASH of Block 1 HASH of Block 2 Each block contains its own content and the hashes of all of the other blocks within the collection, e.g. all other blocks sharing the same case ID. Each block contains its own content and the hashes of all of the other blocks sharing the same case ID keeps data integrity. If the data in any block is updated, its hashes in all the other blocks have to updated as well, which can not be done without proper access to all the nodes. The blocks are stored in different locations on the file storage server 111. The blocks are not stored in sequence but are arranged in a scattered ledger within the file storage server. Each transaction has its own scattered ledger. The scattered ledger is identified by the case ID specific to the transaction. The scattered ledger for each transaction can be identified by the case ID. The private key 109 of the key pair 108 is sent to the key storage server 112. The public key 110 is used for encryption. The public key 110 for each document is shared with the verification agents 115. Private keys 109 are used for decryption. To access the encrypted documents / files the private key 109 must be obtained. The private key 109 is obtained by way of the public key 110. When the verification agents 115 require access to a document / file 102 they request authorisation from the access administrator 114. An access request from a verification agent 115 comprises sending the public key 110 to the access administrator 115. The access administrator 115 notifies the file storage server and the private keys server with the public key and the case ID. The case ID is used to determine what scattered ledger the encrypted documents / files are stored on and how many documents / files are being stored on the scattered ledger. The case ID identifies all documents / files relating to the transaction in order for them to be searched with the public key. As the public key 110 is used to create the file identifier, the public key 110 can be used to identify the document / file 102 within the collection which access is requested to. The access administrator 114 obtains the file identifier using the public key 110 provided by the verification agent 115. The access administrator 115 notifies the file storage 111 and key storage 112 of the access request with the public key 110 and file identifier. Once the second and third servers are notified, the private key 109 and encrypted document / file 102 are sent to the access context manager 113. Access administrator. These users don’t have access to documents themselves they only enable access, and they cannot get the private keys themselves. The access context manager decrypts the encrypted document / file 102. The process for identifying and accessing the encrypted documents / files 102 is as follows: 1. The file expiry date and file identifier are decrypted using the private key corresponding to the public key provided by the verification agent. 2. The file expiry date is checked. If the file expiry date has passed, then the file is expired and cannot be accessed. If the file expiry date has not passed, the file can be access. 3. The file identifier is checked to see if it is divisible by the prime number from the public key. If the file identifier is not divisible by the prime number from the public key, then the file is not authentic and cannot be accessed. If the file identifier is divisible by the prime number from the public key, then the file authentic and can be accessed. Once the document / file has been found authentic and has been decrypted by the access context manager, it is sent to the verification agent. File storage server (second server) 111 Figure 3 shows the file storage server structure 300 in accordance with an example of the invention. The file storage server comprises a centralised scattered ledger. Each transaction has its own scattered ledger. The scattered ledger is identified by the case ID specific to the transaction. Each encrypted document is stored on the file storage server 111 in a scattered ledger in the form of a block 301. Each block 301 with the encrypted document / file, the file expiry date and the file identifier. All of the blocks 301 are saved within the file storage server 111. The file storage server comprises a plurality of servers. The blocks are stored across multiple servers. The multiple servers of the file storage server 111 run on a single machine or a plurality of machines. The file storage server 111 is represented in Figure 3 as 302. The blocks are distributed around multiple nodes / servers, these nodes / servers can be running on a single machine or multiple machines. The use of multiple machines enhances security. When the system creates blocks of information they are stores across multiple nodes / servers. Each block in the scattered ledger contains its own content, and hashes of all the other blocks within the same scattered ledger. As each scattered ledger corresponds to each collection, each block in the scattered ledger contains its own content, and hashes of all the other blocks with the same case ID. The purpose of this is to keep data integrity. If the data in any block is updated, its hashes in all the other blocks of the same scattered ledger are also updated. Updated to the blocks can only be done with proper access to all the nodes of the scattered ledger. The metadata stored further comprises information on how many blocks will be stored on the scattered ledger for each transaction. The number of blocks is determined by the number of files / documents uploaded by the user for each transaction. Each of the blocks 301 are stored separately on the file storage server, the blocks are not stored in sequence. The blocks are not stored in any particular order. As such, each of the documents / files are stored separately on the file storage server. The documents / files are not stored in the same place. The documents / files are not stored in a sequence. This makes it more secure and resistant to tampering than traditional centralised databases. The blocks being saved in different locations within the file storage server makes this a distributed yet centralised ledger. A distributed yet centralised system is the category of system wherein the system is distributed, from the location's perspective, yet the system is controlled by a central authority or central entity. In 5 this case, the blocks are distributed but within the file storage server. The blocks are retrieved by the case ID. The blocks are then searched to identify which of the blocks corresponds to the public key using the identifying and accessing method described above. 10
Claims
1. A data file store for storing multiple collections of files, each collection of files being provided by a user and containing information specific to that user, the data file store comprising:a plurality of data blocks, each data block comprising an encrypted file from the collection of files for a user and a corresponding file identifier;each encrypted file being encrypted using a separate key pair comprising a private key and a public key;the file identifier for each encrypted file comprising an element from the encrypted file’s public key;wherein the data blocks are scattered in different locations in the data file store and individual files of the collection of files are retrievable by using the element of the public key to determine to which of the collection of files the public key corresponds.
2. The data file store of claim 1, wherein the element of the public key comprises an alphanumeric string.
3. The data file store of claims 1 or 2, wherein the element of the public key comprises a prime number.
4. The data file store of claim 3, wherein the file identifier further comprises a second element comprising a prime number.
5. The data file store of claim 1, wherein each block further comprises an encrypted file expiry indicator.
6. The data file store of claim 5, wherein the encrypted file expiry indicator comprises an indicator identifying how long the file can be accessed for.
7. The data file store of claims 5 to 6, wherein the file identifier and file expiry in each block are encrypted.
8. The data file store of claim 7, wherein the file identifier and file expiry in each block are encrypted together.
9. The data file store of any of the preceding claims, wherein the blocks are scattered in any order.
10. The data file store of any of the preceding claims wherein the collection of files comprises a case ID.
11. The data file store of claim 10, wherein the case ID is stored within each block of the collection.
12. The data file store of claim 11, wherein the case ID is encrypted.
13. The data file store of claim 12, wherein the case ID, file identifier, and fileexpiry in each block are encrypted together.
14. The data file store of any of the preceding claims wherein each block contains a hash of all other blocks within the collection.
15. A method of storing a collection of files containing sensitive information for a user to permit subsequent, selective access to individual files of the collection by a third party, the method comprising: uploading a collection of files to an encryption server;encrypting each file of the collection using a separate key pair comprising a private key and a public key;creating a file identifier for each encrypted file using an element from the encrypted file’s public key;each encrypted file and a corresponding file identifier for the collection of files being stored as a block;storing each block in a data file store comprising multiple further collections of files whereby the blocks of the collection of files are scattered and areretrievable by using the element of the public key to determine to which of the collection of files the public key corresponds.
16. The method of claim 15, wherein a collection ID is assigned to the collection 5 of files when they are uploaded.
Citation Information
Patent Citations
System and method to protect sensitive information via distributed trust
US20210067320A1
Securing secrets and their operation
WO2022174122A1