Geology big data secrecy sharing method and device supporting program uploading access and medium

By standardizing the geoscience metadata file and AES-256-GCM encryption, unique data ID is generated, users write code packages and upload verification and analysis, the problem of confidential sharing of geoscience big data is solved, and secure sharing and scientific research efficiency are improved.

CN120337261APending Publication Date: 2025-07-18CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510461119.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Geology big data cannot be kept confidential and shared due to large amounts of sensitive geographical location information, resulting in serious lack of public special data sets and hindering scientific research progress.

Method used

By obtaining geoscience metadata files for standardization, AES-256-GCM algorithm is used for block encryption, and a unique data ID is generated. The user writes a code package and uploads it to the data storage for verification and analysis, and uses a strict security mechanism to ensure data sharing.

Benefits of technology

It realizes the secure sharing of geology big data under the premise of confidentiality, ensures data security, improves scientific research efficiency, and promotes the development of geology big data research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337261A_ABST
    Figure CN120337261A_ABST
Patent Text Reader

Abstract

The invention relates to the field of big data storage and management, and discloses a geoscience big data secrecy sharing method and device supporting program uploading access and a medium, and the method comprises the steps: obtaining a geoscience metadata file, and carrying out the standardization processing of the geoscience metadata file, and obtaining a processed file; carrying out block encryption on the processed file by adopting an encryption algorithm to obtain a complete encrypted file, and storing the encrypted file in a data storage end; the user generates a unique data ID for accessing the complete encrypted file according to a naming rule; a user calls the complete encrypted file for verification and writes a code package; the code package is uploaded to a data storage end, and the data storage end verifies and analyzes the code package. The method realizes sharing of geoscience big data on the premise of confidentiality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data storage and management, and particularly to a method for geoscience big data confidential sharing that supports program upload and access. Background Art

[0002] Computer-related scientific research relies on data sets to conduct experiments to verify research results. In the process of scientific research development, some institutions or organizations spontaneously provide data sets for scientific researchers to use widely. Over time, these frequently used data sets gradually form public thematic data sets. In the field of computer-related research, if one hopes to publish a paper on the same topic in a high-level journal or conference, it is usually necessary to conduct experiments based on the same public thematic data set to prove the advancement of one's own method.

[0003] Geoscience big data, as an important category of big data, is closely related to earth science and contains a vast amount of complex data. Its research involves the cross-integration of geoscience and computer science. Like other computer-related fields, the geoscience big data field also needs to verify scientific research results through experiments. However, compared with other types of big data, geoscience big data has a significant characteristic, that is, it is rich in a large amount of sensitive geographical location information. This characteristic leads to a very low degree of sharing of data that has been labeled and preprocessed for specific geoscience big data problems among different institutions, and thus causes a serious lack of public thematic data sets.

[0004] The lack of public thematic data sets has greatly hindered the advancement of geoscience big data research, which is specifically reflected in the following key aspects:

[0005] 1. Method practice is blocked: Even if scientific researchers design innovative methods, it is difficult to carry out practical operations through actual data if they cannot obtain experimental data, making it impossible to verify the implementation of scientific research results.

[0006] 2. Result verification is difficult: After the review and acceptance of a paper, since reviewers and readers cannot obtain the experimental data used in the paper, it is difficult to effectively verify the authenticity and quality of the published research results, which is not conducive to the rigor and reliability of academic research.

[0007] 3. Comparative experiment dilemma: It is difficult to conduct comparative experiments between the scientific research results of different researchers, and it is impossible to intuitively judge the advantages and disadvantages of various methods, which restricts academic exchanges and the rapid development of the discipline.

[0008] 4. Follow-up research is rare: Different from other computer-related research fields, in the field of geoscience big data, it is extremely rare to conduct follow-up research based on the methods in papers, which is not conducive to the accumulation and expansion of knowledge and delays the overall research process of the discipline. Summary of the Invention

[0009] The object of the present invention is to propose a method, device and medium for secure sharing of geoscience big data that supports program upload and access, so as to solve the technical problem that current geoscience big data cannot balance confidentiality and sharing due to the involvement of a large amount of sensitive geographical location information.

[0010] Specifically, a method for secure sharing of geoscience big data that supports program upload and access provided by the present invention includes the following steps:

[0011] S1. Obtain a geoscience metadata file and perform standardization processing on it to obtain a processed file;

[0012] S2. Use an encryption algorithm to perform block encryption on the processed file to obtain a complete encrypted file and store it in a data storage end;

[0013] S3. The user generates a unique data ID according to a naming rule for accessing the complete encrypted file;

[0014] S4. The user calls the complete encrypted file for verification and writes a code package;

[0015] S5. Upload the code package to the data storage end, and the data storage end verifies and parses the code package.

[0016] A storage medium stores instructions and data for implementing a method for secure sharing of geoscience big data that supports program upload and access.

[0017] A device for secure sharing of geoscience big data that supports program upload and access includes: a processor and the storage medium; the processor loads and executes the instructions and data in the storage medium for implementing a method for secure sharing of geoscience big data that supports program upload and access.

[0018] The beneficial effects provided by the present invention are:

[0019] First of all, the proposed method for secure sharing of geoscience big data realizes the sharing of geoscience big data under the premise of confidentiality through data standardization processing, encrypted storage management, strict security mechanisms and standardized user operation processes.

[0020] Secondly, the present invention ensures the security of data and prevents the leakage of sensitive information and the tampering of data.

[0021] The present invention also improves the scientific research efficiency, facilitates scientific research personnel to carry out research work based on shared data, and is of great significance for promoting the development of geoscience big data research. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a simple process schematic diagram of the method of the present invention;

[0023] Figure 2It is a schematic diagram of the operation of the hardware device according to an embodiment of the present invention. Detailed implementation manners

[0024] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0025] Before formally elaborating on the present invention, a general description of the solution of the present invention will be given first for easy understanding.

[0026] Please refer to Figure 1 , a method for secure sharing of geoscience big data supporting program upload and access provided by the present invention includes:

[0027] S1. Obtain a geoscience metadata file and perform standardization processing on it to obtain a processed file;

[0028] It should be noted that in the present invention, information such as the source, format, calling method, field meaning, sensitive level label, etc. of geoscience data are detailedly recorded in the metadata file (metadata.json);

[0029] In addition, in the present invention, a data description document (README.md) is also equipped, and the data collection method, usage restrictions and authorization agreements are described in the data description document.

[0030] The standardization processing described in step S1 specifically refers to: assigning a unique data ID to the data set of the geoscience metadata file, and the rule of the unique data ID is: the first six digits of the value generated after SHA256 hashing of the DOI number of the paper using this data + the data source party ID.

[0031] As an embodiment, in the present invention, if the data source party ID is S and the paper DOI number is D, the generation formula of the unique data set id I is:

[0032] I = s + SUBSTRING(SHA256(D), 0, 6) (1)

[0033] Wherein, SUBSTRING(SHA256(D), 0, 6) represents taking the first six digits of the value generated after SHA256 hashing of the DOI number D.

[0034] S2. Perform block encryption on the processed file using an encryption algorithm to obtain a complete encrypted file and store it in the data storage end;

[0035] It should be noted that in the present invention, the AES-256-GCM algorithm is used to perform block encryption on the data file, and each block is attached with an independent initialization vector (IV) and an authentication tag (MAC) to prevent tampering.

[0036] Step S2 is specifically as follows:

[0037] S21. Generate the master key Km through the key management system; dynamically create the data key Kd through the random number generator;

[0038] In the present invention, the 256-bit (32-byte) master key Km is generated through the key management system (KMS). The master key is only used to encrypt the data encryption key (DEK). The data key K_d is dynamically created through the random number generator, and each file or data block uses an independent data key. The data key itself needs to be encrypted by the master key and then stored, that is, the encrypted data key e(Km, Kd) is stored, where E is the encryption function, and the encryption process can be expressed as:

[0039] E(Km, Kd) = AES-256-GCM-ENCRYPT(Km, Kd) (2)

[0040] Ensure the security of key storage.

[0041] S22. Cut the geoscience metadata file into multiple data blocks according to the fixed size B;

[0042] Cut the original file into multiple data blocks according to the fixed size B. Suppose the size of the original file is N, and the calculation formula for the number of data blocks n is:

[0043]

[0044] where ┌x┐ represents rounding up x, and the last block is allowed to be smaller than the block size (no padding is required). Assign a unique identifier IDi (i = 1, 2,..., n) to each block for subsequent decryption and recombination in order.

[0045] S23. Generate a unique random initialization vector IVi for each data block, and encrypt the plaintext block Pi based on the master key and the data key to generate the ciphertext Ci and the authentication tag MACi, forming an encrypted file;

[0046] Generate a unique 12-byte random IV for each data block. Suppose the IV of the i-th data block is IVi. Use the AES-256-GCM algorithm, with the data key Kd and the current block's IVi as inputs, to encrypt the plaintext block Pi. During the encryption process, the ciphertext Ci and the 16-byte authentication tag (MAC) MACi are generated simultaneously. The mathematical expression of the encryption process is:

[0047] Ci, MACi = AES-256-GCM(Kd, IVi, Pi) (4)

[0048] Among them, MACi is used to verify data integrity. Combine IVi, MACi, and Ci in sequence as the final encryption result of this block. Each encrypted block is guaranteed to be stored and transported independently during the transportation process to avoid overall data damage caused by damage to a single block.

[0049] S24. After adding meta-information to the encrypted file header, write all encrypted blocks into the file in the original order to obtain a complete encrypted file.

[0050] Add meta-information to the encrypted file header, including the encryption algorithm identifier A, block size B, and key version V. Write all encrypted blocks into the file in the original order to form a complete encrypted file.

[0051] In the present invention, sensitive information is further processed. Specifically, a test data set is automatically generated according to the data sensitive label. Let the sensitive label be T. For geographical location information, use the random offset function R(x) to process the real geographical location x, that is, the replaced geographical location is R(x); for numerical information, let the original value be y and the scaling ratio be k, and the scaled value is k×y. The test data and the original data share the same unique ID for the user to verify the code logic.

[0052] Through the above settings, the MAC is generated by the encryption algorithm during encryption and is strongly associated with the ciphertext and IV. Any modification to the ciphertext or IV will cause the MAC verification to fail.

[0053] Force MAC verification during decryption to ensure that the data has not been tampered with during storage or transmission.

[0054] Each file uses an independent data key to reduce the risk of key leakage. At the same time, set the main key update period to Tm, update the main key regularly, and re-encrypt the historical data.

[0055] The IV of each block is generated by a cryptographically secure random number generator to ensure global uniqueness.

[0056] Under the same key, reusing the same IV will undermine the security of the GCM mode, and it is necessary to strictly ensure the uniqueness of the IV.

[0057] S3. The user generates a unique data ID according to the naming rule to access the complete encrypted file;

[0058] Generate a unique data ID according to the naming rule. The naming rule has been described above.

[0059] Retrieve the associated test data set (sharing the same ID as the original data) according to the data ID.

[0060] Return the desensitized test data (such as geographical coordinate offset, numerical value scaled proportionally), ensuring that users cannot reverse-infer the original sensitive information from the test data.

[0061] Record the user's query behavior, generate a temporary access token (Token), and only allow the same Token to be used when submitting subsequent code.

[0062] S4. The user calls the complete encrypted file for verification and writes a code package.

[0063] It should be noted that when the user writes code based on the test data, they need to follow the interface specifications:

[0064] Data input: Load data through platform.load_dataset(DATA_ID).

[0065] Result output: Save the result using platform.save_result(output).

[0066] Specifically, load data through platform.load_dataset(DATA_ID). Among them, DATA_ID needs to conform to the established naming rules, and the input data format should meet [specific format description, such as JSON, CSV, etc.] to ensure that the data can be correctly loaded.

[0067] Save the result using platform.save_result(output). output should be organized according to [specific format requirements, such as including specific fields, data structures, etc.] to ensure that the result can be correctly saved and used subsequently.

[0068] After the code is uploaded, it is saved in the folder where the dataset is located and starts running.

[0069] The user runs the code in the local sandbox simulator using the test data to verify the correctness of the function. The present invention provides a test environment SDK to simulate the data loading and result saving interfaces.

[0070] S5. Upload the code package to the data storage end, and the data storage end verifies and parses the code package.

[0071] It should be noted that step S5 is specifically as follows:

[0072] S51. The data storage end performs static scanning and digital signature on the code package uploaded by the user.

[0073] As an example, static scanning refers to detecting malicious code (such as file operations, network requests, source data access, etc.).

[0074] A digital signature means using the user's private key SK to sign the hash value H(C) of the code. Let the signature function be Sign, then the calculation of the signature result Sig is as follows:

[0075] Sig = Sign(SK, H(C)) (5)

[0076] Ensure code integrity.

[0077] S52. When a complete encrypted file is called in the code package, locate the encrypted file block according to the unique data ID in step S3, and perform reverse parsing on it to obtain the corresponding Kd, IVi, Ci, MACi, and verify whether it is the same as the value during encryption. If so, the verification passes, and it is restored to the complete original metadata file;

[0078] When the code calls platform.load_dataset(DATA_ID), the system locates the encrypted file block according to the data ID, parses the encrypted file header, and obtains information such as the block size B, encryption algorithm A, key version V, etc., to ensure that the decryption parameters are consistent with those during encryption.

[0079] Read data from the encrypted file block by block according to the block size B, with a fixed length read each time. Separate IVi, MACi, and Ci from each block. Initialize the decryptor with the data key Kd and the IVi of the current block, input the ciphertext Ci and verify MACi: If

[0080] Verify(Kd, IVi, Ci, MACi) = True (6)

[0081] (Verify is the verification function), output the plaintext data block Pi; If

[0082] Verify(Kd, IVi, Ci, MACi) = False (7)

[0083] Determine that the data has been tampered with, terminate the decryption and trigger a warning.

[0084] Concatenate all the decrypted plaintext blocks in the original order to restore them to a complete file. Compare the hash value (such as SHA - 256) of the decrypted file with the original file to verify integrity.

[0085] S53. Start a read - only Docker container, mount the encrypted data volume, and prohibit external network access and file writing permissions.

[0086] To ensure the security of sharing, read-only Docker containers are launched in the present invention, and encrypted data volumes (the storage paths where the original data is located) are mounted. External network access is disabled, and the file system write permissions are restricted (only the / tmp directory is allowed). The platform decrypts the data key through the KMS and temporarily injects it into the container memory for AES-GCM decryption. The decryption key is only valid during the execution of the current task and is destroyed immediately after the task ends.

[0087] Finally, some calculation methods regarding the above process in the present invention are described as follows:

[0088] 1. Formula for calculating the hash value of geological data:

[0089] Assume that the content of geological data is D, and the hash algorithm used is Hash (such as SHA-256). Then the calculation formula for the hash value HD of geological data is:

[0090] HD = Hash(D) (8)

[0091] It should be noted that the hash value HD of geological data is used to uniquely identify the data content. In the blockchain network, it serves as the digital fingerprint of the data. When the data is transmitted or stored between nodes, the integrity of the data can be quickly detected by comparing HD.

[0092] 2. Digital signature formula for data application package: Assume that the basic index information of geological data is I, the private key of the data owner is SKowner, and the signature function is Sign. Then the calculation formula for the signature result Sigowner is:

[0093] Sigowner = Sign(SKowner, I) (9)

[0094] It should be noted that digital signature is the key to ensuring the authenticity and integrity of the data application package. In the blockchain environment, the data owner uses the private key to sign the basic index information to prevent others from forging the application package. Other nodes can confirm the reliability of the application source by verifying the signature, ensuring that only the legitimate owner can apply for data operations.

[0095] 3. Public key verification signature formula: Assume that the public key of the data owner extracted is PKowner, the received signature result is Sigrecv, the received basic index information is Irecv, and the verification function is VerifySign. Then the verification process can be expressed as:

[0096] VerifySign(PKowner, Sigrecv, Irecv) = True / False (10)

[0097] It should be noted that after the blockchain node receives the data application packet, it uses the public key PKowner of the data owner to verify the signature Sigrecv. If the verification result is True, it indicates that the signature is valid, that is, the data application packet has not been tampered with during transmission and comes from a legitimate owner; if it is False, the node rejects the application packet to prevent illegal data from entering the blockchain system and maintain the security and trust of the system.

[0098] 4. Block generation formula: Let the ciphertext index information be Cindex, the hash value of the previous block be Hprev, the timestamp be T, and the nonce be R (used for the random allocation mechanism). If the block generation function is GenerateBlock, then the calculation formula for the generated block Block is:

[0099] Block = GenerateBlock(Cindex, Hprev, T, R) (11)

[0100] It should be noted that the hash value Hprev of the previous block connects the new block to the existing chain structure of the blockchain to construct an immutable chain ledger. The timestamp T records the block generation time and is used to sort and trace the chronological order of data operations. The nonce R in consensus mechanisms such as proof of work continuously changes to find a hash value that meets specific conditions, ensuring the security and decentralization of block generation and enabling the new block to be legally added to the blockchain.

[0101] After the operation ends, the result file is encrypted using the user's public key PK (RSA-2048). Let the encryption function be Enc, then the calculation of the encrypted result file Eresult is as follows:

[0102] Eresult = Enc(PK, Result) (12)

[0103] Only the user's private key can decrypt it.

[0104] Generate a one-time HTTPS link (valid for 5 minutes), and the server copy will be automatically deleted after timeout.

[0105] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the operation of the hardware device in an embodiment of the present invention. The hardware device specifically includes: a geoscience big data confidentiality sharing device 401 that supports program upload and access, a processor 402, and a storage medium 403.

[0106] A geoscience big data confidentiality sharing device 401 that supports program upload and access: The geoscience big data confidentiality sharing device 401 that supports program upload and access implements the geoscience big data confidentiality sharing method that supports program upload and access.

[0107] Processor 402: The processor 402 loads and executes the instructions and data in the storage medium 403 to implement the method for secure sharing of geoscience big data supporting program upload and access.

[0108] Storage medium 403: The storage medium 403 stores instructions and data; the storage medium 403 is used to implement the method for secure sharing of geoscience big data supporting program upload and access.

[0109] The beneficial effects of the present invention are as follows:

[0110] First of all, the proposed method for secure sharing of geoscience big data realizes the sharing of geoscience big data under the premise of confidentiality through data standardization processing, encrypted storage management, strict security mechanisms, and standardized user operation processes.

[0111] Secondly, the present invention guarantees the security of data and prevents the leakage of sensitive information and the tampering of data.

[0112] The present invention also improves the scientific research efficiency, facilitates scientific researchers to carry out research work based on shared data, and is of great significance to promoting the development of geoscience big data research.

[0113] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for secure sharing of geoscience big data supporting program upload and access, characterized in that: It includes the following steps: S1. Obtain a geoscience metadata file and perform standardization processing on it to obtain a processed file; S2. Use an encryption algorithm to perform block encryption on the processed file to obtain a complete encrypted file and store it at the data storage end; S3. The user generates a unique data ID according to the naming rule to access the complete encrypted file; S4. The user calls the complete encrypted file for verification and writes a code package; S5. Upload the code package to the data storage end, and the data storage end verifies and parses the code package.

2. The geoscience big data confidential sharing method supporting program upload and access according to claim 1, characterized in that: The standardization processing described in step S1 specifically refers to: assigning a unique data ID to the data set of the geoscience metadata file, and the rule of the unique data ID is: the first six digits of the value generated after SHA256 hashing of the data source party ID + the DOI number of the paper using this data.

3. The geoscience big data confidentiality sharing method supporting program upload and access as claimed in claim 1, wherein: Step S2 is specifically as follows: S21. Generate a master key Km through a key management system; dynamically create a data key Kd through a random number generator; S22. Cut the geoscience metadata file into multiple data blocks according to a fixed size B; S23. Generate a unique random initialization vector IVi for each data block, and encrypt the plaintext block Pi based on the master key and the data key to generate a ciphertext Ci and an authentication tag MACi, forming an encrypted file; S24. After adding meta-information to the head of the encrypted file, write all the encrypted blocks into the file in the original order to obtain a complete encrypted file.

4. The geoscience big data confidential sharing method supporting program upload and access according to claim 3, characterized in that: In step S24, the meta-information includes: encryption algorithm identifier A, block size B, and key version V.

5. The geoscience big data confidentiality sharing method supporting program upload and access as described in claim 3, characterized in that: In step S24, test data is also generated for the sensitive information in the complete encrypted file, specifically as follows: For the encrypted file containing sensitive information, for the sensitive geographical location information, use a random offset function R(x) to process the real geographical location x to obtain a processed geographical location; For the sensitive numerical value y, perform scaling processing to obtain a scaled numerical value; the processed geographical location and the scaled numerical value constitute test data, and the test data shares the same unique data ID as the original data.

6. The geoscience big data confidential sharing method supporting program upload and access according to claim 5, characterized in that: The process of the user calling the complete encrypted file in step S4 is as follows: The user calls the complete encrypted file based on the test data according to the specification interface and runs the verification result in the local sandbox simulator.

7. The geoscience big data confidential sharing method supporting program upload and access according to claim 1, characterized in that: Step S5 is specifically as follows: S51. The data storage end performs static scanning and digital signature on the code package uploaded by the user; S52. When the complete encrypted file is called in the code package, locate the encrypted file block according to the unique data ID in step S3, and perform inverse parsing on it to obtain the corresponding Kd, IVi, Ci, MACi, and verify whether it is the same as the value during encryption. If so, the verification passes and it is restored to the complete original metadata file; S53. Start a read-only Docker container, mount the encrypted data volume, and prohibit external network access and file writing permissions.

8. A storage medium, characterized in that: The storage medium stores instructions and data for implementing a geoscience big data confidentiality sharing method according to any one of claims 1 to 7.

9. A geoscience big data confidential sharing device supporting program upload and access, characterized in that: It includes: Processor and storage medium; the processor loads and executes instructions and data in the storage medium to implement a method for secure sharing of geoscience big data supporting program upload and access according to any one of claims 1 to 7.