Data security protection method

Through dynamic key encryption and blockchain storage result template methods, the security and format unified problem of data transmission among medical systems is solved, and efficient and secure data processing and aggregation is achieved.

CN120389911AInactive Publication Date: 2025-07-29NANJING AIKEMAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510875131.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing AES encryption algorithms have the risk of key leakage in data transmission between medical systems, resulting in security risks of data leakage, and the data formats of different medical systems lead to complex data processing.

Method used

The calculation results are encrypted using dynamic key generation method, and the result template is stored through blockchain to ensure data security, and the large language model is used for semantic matching and data aggregation, and a unified result template is generated to reduce data format differences.

Benefits of technology

It improves the security of data transmission between medical systems, reduces the risk of data leakage, improves the efficiency and accuracy of data processing, and simplifies the process of unified data formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120389911A_ABST
    Figure CN120389911A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric digital data processing, in particular to a data security protection method, which comprises the following steps that: a medical system obtains a result template, and generates a calculation result containing all required fields according to the required fields in the result template; generating a dynamic key according to a pre-shared key, random Salt and experimental data hash, encrypting a calculation result through the dynamic key, generating a ciphertext and a transmission packet with the ciphertext, and uploading the ciphertext and the transmission packet to a summarizing system; and after the summarizing system obtains the transmission packet, decrypting the ciphertext in the transmission packet, extracting an original field from the ciphertext, matching the original field with a standard field, and carrying out subsequent analysis. According to the invention, the experimental result is encrypted by using the experimental data, the experimental result can be stored in the medical system locally through the unified result template, and part of the data is sent to the summarizing system, so that the safety of the medical data is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electrical digital data processing, and in particular, to a data security protection method. Background Art

[0002] With the development of medical informatization, medical information is no longer limited to a single medical system, but involves information transmission and sharing among multiple medical systems. For example: carbon-13 urea breath test, multi-center cardiovascular risk assessment and intervention strategy optimization research based on CACS, multi-center drug safety assessment of adverse drug reactions, etc.

[0003] Since a single medical system is difficult to complete experiments or research, after a medical system completes data collection and calculation, it sends the data to other medical systems. However, in these experiments and research, some sensitive information such as personal information of patients is often involved. Once leaked, it will cause serious adverse effects. Therefore, it is necessary to encrypt the data, and the AES encryption algorithm is usually used. However, the AES encryption algorithm has only a fixed key. Therefore, once the key is leaked, it is easily attacked and the data is leaked. Summary of the Invention

[0004] In order to improve the security of data transmission between medical systems, the present invention provides a data security protection method.

[0005] The present invention provides a data security protection method, adopting the following technical solutions: A data security protection method includes the following steps: A medical system obtains a result template and generates a calculation result including all required fields according to the required fields in the result template; Generate a dynamic key according to a pre-shared key, a random Salt, and the hash of experimental data, and dynamically encrypt the calculation result with the dynamic key to generate a ciphertext and a transmission packet with the ciphertext, and upload it to the summary system; After the summary system obtains the transmission packet, decrypt the ciphertext in the transmission packet, extract the original fields from the ciphertext and match them with the standard fields for subsequent analysis; The summary system converges the obtained calculation results.

[0006] In a specific feasible implementation, before the medical system obtains the result template, the summary system writes the metadata of the result template into the blockchain. The key metadata includes the template ID, version number, hash value, and storage pointer.

[0007] In a specific feasible implementation, the medical system downloads the key metadata from the blockchain, extracts the hash value and the storage pointer, downloads the result template through the storage pointer, and verifies the result template through the hash value.

[0008] In a specific implementation scheme, the result template includes mandatory fields and optional fields. The mandatory fields include project ID, timestamp, and indicator set; the optional fields include symptom description.

[0009] In a specific implementation scheme, the formula for dynamically encrypting the calculation result using a dynamic key is: K=HKDF(PSK||Salt||H(Data)) Where K represents a dynamic key, and HKDF is a key derivation function used to generate a high entropy key from PSK, Salt, and data hash. PSK is a pre-shared key, and its calculation formula is: new =HKDF(PSK old ||Timestamp||random number), Timestamp is the current timestamp; random number is a randomly generated value; PSK new Indicates the updated pre-shared key; PSK old Represents the original pre-shared key; Salt is a random value, and H(Data) is the result of hashing the experimental data.

[0010] In a specific implementation scheme, the structure of the ciphertext is: C = (K, nonce, plaintext data), where C is the ciphertext and nonce is a random number used only once; The structure of the transmission packet P is: P = (C, nonce, metadata), where metadata is metadata of the calculation result.

[0011] In a specific implementation scheme, metadata is signed using HMAC, and the signature formula is: S = HMAC-SHA256 (PSK, metadata), where S is the signature of the metadata; After the aggregation system obtains the transmission package P, it verifies whether the data in the transmission package P has been tampered with through the signature S.

[0012] In a specific implementation scheme, after the aggregation system obtains the transmission packet, it decrypts the ciphertext in the transmission packet, extracts the original field from the ciphertext, and matches it with the standard field, including the following steps: The aggregation system generates a derived ciphertext K' using the pre-shared PSK and Salt, decrypts the ciphertext C using the derived key K', and extracts the original field from the ciphertext C; A large language model is used to perform semantic matching between the original field and the standard field. If the matching degree is greater than or equal to 90%, it indicates that the original field matches the standard field and is unified into the standard field. Otherwise, a query is generated for manual judgment. Calculate the value corresponding to the standard field according to the historical mean corresponding to the standard field. If the value is greater than or equal to the preset threshold, it indicates that the data value corresponding to the standard field is an outlier, mark the outlier and generate a warning.

[0013] In a specific feasible implementation, after the summary system decrypts the ciphertext in the transmission packet, analyze the historical data, calculate the field appearance rate of the optional fields. If the field appearance rate is greater than or equal to 80%, then expand the field into the required fields, thereby generating a new result template.

[0014] In a specific feasible implementation, the summary system converges the calculation results by using real-time stream processing and batch calculation. The time window of real-time stream processing is 1 minute, and the calculation formula is:

[0015] In the above formula, is the mean value per minute, is the total number of data points within the current time window; is the time window; is the th data point value; The formula for batch calculation of the global variance is:

[0016] In the above formula, is the global variance; is the global mean.

[0017] In summary, the present invention includes the following beneficial effects: 1. The medical system generates calculation results by downloading the result template, encrypts the calculation results and transmits them to the summary system, so that the detailed experimental data is saved locally in the medical system, and only the calculation results are transmitted, thereby reducing the security risks in the process of experimental data transmission and improving the security of experimental data.

[0018] 2. The medical systems adopt the same result template, so that after the summary system obtains the calculation results of multiple medical systems, the complex standardization process is omitted, and the data processing efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flowchart of the data security protection method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The following further describes the present invention in detail with reference to the attached Figure 1 drawings.

[0021] For the convenience of understanding, taking the example where a medical system transmits the results of processing items to a summary system for aggregation through the summary system, the data security protection method includes the following steps: S100, the medical system generates a calculation result according to the result template.

[0022] The result template is stored in a specified cloud or system, such as IPFS, etc. The summary system writes the key metadata of the result template to the blockchain. The key metadata includes the template ID, version number, hash value, storage pointer, etc. After the key metadata of the result template is uploaded to the blockchain, it cannot be tampered with and the update history can be traced.

[0023] The medical system downloads the key metadata of the result template from the blockchain, extracts the hash value and the storage pointer, and downloads the result template from the storage location specified by the storage pointer through the storage pointer. Calculate the hash value of the downloaded result template and compare it with the hash value in the key metadata on the blockchain. If the hash values are equal, it indicates that the verification is passed and the result template can be used; otherwise, it indicates that the verification fails, the result template has been tampered with, and it needs to be downloaded again. While avoiding the impact on blockchain performance if the result template is large, the security of the result template is ensured by calculating the hash.

[0024] Taking IPFS as the storage location as an example, the storage pointer is the content identifier (CID) of IPFS. Then the medical system can download the result template from IPFS through the storage pointer and check whether the result template has been tampered with by calculating the hash.

[0025] The result template has mandatory fields. The medical system generates a calculation result that includes at least all the mandatory fields according to the mandatory fields in the result template. That is, the calculation result output by the medical system must include all the mandatory fields in the result template. In addition, it can also include some optional fields required in the corresponding items, or may not include the optional fields.

[0026] The mandatory fields include the project ID, timestamp, indicator set, etc. Exemplarily, the result template of a medical detection project may include the project ID (such as "FPG-GODPOD"), timestamp (2025.04.01.09:18:36AM), indicator set (HbA1c, GLU, etc.) as mandatory fields, and the optional fields may include the symptom description of the patient, etc.

[0027] Multiple medical systems adopt a unified result template, and the generated calculation results unify the data format through the result template, which can save the complex data standardization process caused by different data in subsequent data processing due to different medical systems. At the same time, since the medical system sends the calculation result and the detailed experimental data, etc. are all stored in the medical system, the security risks in the data transmission process of the medical system are reduced.

[0028] S200, encrypt and transmit the calculation result.

[0029] Generate a 256-bit dynamic key K based on the pre-shared key (PSK), random Salt, and the hash of the experimental data H(Data), and perform dynamic encryption on the calculation result. The formula is: K = HKDF(PSK || Salt || H(Data)) where HKDF is a key derivation function used to generate a high-entropy key from PSK, Salt, and the data hash; PSK is the pre-shared key used to ensure that only authorized systems can generate and verify the key. The calculation formula is: PSK new = HKDF(PSK old || Timestamp || random number), where Timestamp is the current timestamp used to ensure the uniqueness of each update; the random number is a randomly generated value to increase the unpredictability of the key; PSK new represents the updated pre-shared key; PSK old represents the original pre-shared key; Salt is a random value, and H(Data) is the result of hashing the experimental data. It is easy to understand that PSK is updated every 24 hours to reduce the risk of long-term key leakage.

[0030] Encrypt the calculation result with the dynamic key K to generate the ciphertext C and the transmission packet P. The structure of the ciphertext C is: C = (K, nonce, plaintext data), where nonce is a random number used only once to prevent replay attacks. The structure of the transmission packet P is: P = (C, nonce, metadata), and the metadata is the metadata of the calculation result, including data such as the source and timestamp of the calculation result.

[0031] To ensure that the data has not been tampered with during transmission, sign the metadata through HMAC. The signature formula is: S = HMAC-SHA256(PSK, metadata), where S is the signature of the metadata. After the subsequent aggregation system obtains the transmission packet P, it can verify whether the data in the transmission packet P has been tampered with through the signature S to ensure the integrity and accuracy of the data.

[0032] S300, the aggregation system decrypts the transmission packet for processing and analysis of the calculation result.

[0033] The aggregation system generates a derived key K' through the pre-shared PSK and Salt, decrypts the ciphertext C with the derived key K', and extracts the original fields from the ciphertext C. The extracted original fields can be restored to the original calculation result.

[0034] Because the same indicator may have different names in different medical systems, such as blood sugar value, which may be called blood sugar level in some other medical systems, a large language model is used to perform semantic matching using algorithms such as cosine similarity. The original fields in the ciphertext are semantically matched with the standard fields. If the match is greater than or equal to 90%, it indicates that the original field matches the standard field and is unified as the standard field. Otherwise, a query is generated and manually judged.

[0035] According to the historical mean value corresponding to the standard field, calculate the corresponding The value is calculated as follows:

[0036] in, is the current data value, is the historical average, is the historical standard deviation. If the value is greater than or equal to the preset threshold, it indicates that the data value corresponding to the standard field is an outlier. The outlier is marked and a warning is generated to facilitate subsequent verification of the outlier.

[0037] After completing the field unification and outlier marking of the calculation results, perform subsequent analysis of the calculation results.

[0038] S400: Optimize the result template according to the calculation result.

[0039] The summary system uses the obtained calculation results and analyzes historical data. Based on the field occurrence rate, it converts optional fields into required fields and further expands the required fields. The calculation formula for field occurrence rate is:

[0040] In the above formula, is the field occurrence rate; is the total number of items; is the frequency of occurrence, that is, the number of times the field appears in all items. Repeated occurrences in the same item are counted as one. , then the field is expanded to the required field, thereby generating a new result template that better meets actual needs. The updated template is re-uploaded to IPFS and obtains a new content identifier, thereby generating new key metadata. The aggregation system writes the new key metadata to a new block on the blockchain.

[0041] S500: The aggregation system aggregates the calculation results.

[0042] The calculation results are aggregated by means of real-time stream processing and batch computing. The real-time stream is used to respond to the latest data with low latency, and the high throughput of batch computing is used to process all historical data, so as to improve the aggregation efficiency of the calculation results. In this embodiment, the time window of real-time stream processing is 1 minute, and the calculation formula is:

[0043] In the above formula, is the average value per minute, is the total number of data points within the current time window; is the time window; [[ID=IS=13]]is the value of the

[0044] The formula for calculating the global variance by means of batch computing is:

[0045] In the above formula, is the global variance; is the global average value.

[0046] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention shall be covered within the protection scope of the present invention.

Claims

1. A data security protection method, characterized in that: It includes the following steps: The medical system obtains the result template and generates a calculation result containing all required fields according to the required fields in the result template; Generate a dynamic key based on the pre-shared key, random Salt, and the hash of the experimental data. Dynamically encrypt the calculation result with the dynamic key to generate a ciphertext and a transmission packet with the ciphertext, and upload it to the aggregation system; After the aggregation system obtains the transmission packet, decrypt the ciphertext in the transmission packet, extract the original fields from the ciphertext, match them with the standard fields, and perform subsequent analysis; The aggregation system aggregates the obtained calculation results.

2. The data security protection method according to claim 1, wherein: Before the medical system obtains the result template, the aggregation system writes the metadata of the result template to the blockchain. The key metadata includes the template ID, version number, hash value, and storage pointer.

3. The data security protection method according to claim 2, wherein: The medical system downloads the key metadata from the blockchain, extracts the hash value and storage pointer, downloads the result template through the storage pointer, and verifies the result template through the hash value.

4. The data security protection method according to claim 1, characterized in that: The result template includes required fields and optional fields. The required fields include the project ID, timestamp, and metric set; the optional fields include symptom descriptions.

5. The data security protection method according to claim 1, wherein: The formula for dynamically encrypting the calculation result with the dynamic key is: K = HKDF(PSK||Salt||H(Data)) Among them, K represents the dynamic key, and HKDF is a key derivation function used to generate a high-entropy key from PSK, Salt, and the data hash; The PSK is a pre-shared key, and its calculation formula is: PSK new =HKDF(PSK old ||Timestamp||Nonce), where Timestamp is the current timestamp; the nonce is a randomly generated value; PSK new represents the updated pre-shared key; PSK old represents the original pre-shared key; Salt is a random value, and H(Data) is the result of hashing the experimental data.

6. The data security protection method according to claim 1, wherein: The structure of the ciphertext is: C = (K, nonce, plaintext data), where C is the ciphertext and nonce is a random number used only once; The structure of the transmission packet P is: P = (C, nonce, metadata), where the metadata is the metadata of the calculation result.

7. The data security protection method according to claim 6, wherein: Sign the metadata through HMAC, and the signature formula is: S = HMAC-SHA256(PSK, metadata), where S is the signature of the metadata; After the aggregation system obtains the transmission packet P, verify whether the data in the transmission packet P has been tampered with through the signature S.

8. The data security protection method according to claim 1, characterized in that: After the aggregation system obtains the transmission packet, decrypt the ciphertext in the transmission packet, extract the original fields from the ciphertext, and match them with the standard fields, including the following steps: The aggregation system generates a derived ciphertext K' through the pre-shared PSK and Salt, decrypts the ciphertext C with the derived key K', and extracts the original fields from the ciphertext C; Perform semantic matching between the original fields and the standard fields through a large language model. If the matching degree is greater than or equal to 90%, it indicates that the original field matches the standard field, and it is unified as the standard field; otherwise, an inquiry is generated and judged manually; Calculate the value corresponding to the standard field according to the historical mean corresponding to the standard field. If the value is greater than or equal to the preset threshold, it indicates that the data value corresponding to the standard field is an outlier, mark the outlier and generate a warning.

9. The data security protection method according to claim 4, characterized in that: After the aggregation system decrypts the ciphertext in the transmission packet, analyze the historical data, calculate the field appearance rate of the optional fields. If the field appearance rate is greater than or equal to 80%, expand the field to the required fields to generate a new result template.

10. The data security protection method according to claim 1, characterized in that: The aggregation system uses real-time stream processing and batch calculation methods to aggregate the calculation results. The time window for real-time stream processing is 1 minute, and the calculation formula is: In the above formula, is the average value per minute, is the total number of data points within the current time window; is the time window; is the value of the data point; The formula for batch calculation of the global variance is: In the above formula, is the global variance; is the global mean.