Medical data analysis system and method based on homomorphic encryption technology

By dividing the data flow and evaluating the complexity of the analysis task in the medical data analysis system, the tasks are divided into simple and complex categories, and using parallel computing to process complex tasks, the problem that homomorphic encryption is difficult to deal with complex calculations is solved, and efficient medical data analysis is achieved.

CN120197198AActive Publication Date: 2025-06-24NANJING YILIAN SUNSHINE INFORMATION TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510629553.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-24
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

In the process of medical data processing, homomorphic encryption technology is difficult to directly execute complex calculations and queries, resulting in the challenge of handling complex calculations and queries in practical applications.

Method used

By dividing the medical data stream into independent segments and evaluating coefficients based on the complexity of the analytical task, the analytical tasks are divided into two categories: simple and complex. For simple analysis tasks, encrypted storage is performed directly; for complex analysis tasks, it is decomposed into sub-tasks, and parallel encryption is used to perform parallel encryption processing, and the summary results are gradually decrypted.

Benefits of technology

It realizes dynamic division of tasks according to the complexity of tasks, improves task processing efficiency, and can effectively analyze while protecting data privacy, overcomes the limitations of homomorphic encryption when processing complex calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197198A_ABST
    Figure CN120197198A_ABST
Patent Text Reader

Abstract

The invention discloses a medicine data analysis system and method based on a homomorphic encryption technology, and relates to the technical field of data analysis. The system comprises a data processing and segmenting module, a task complexity evaluation module, a task division and storage module and a task decomposition and parallel computing module. The data processing and segmenting module segments the medicine data flow into independent segments according to a time window, generates subsets according to business dimensions, and matches analysis tasks with the subsets; the task complexity evaluation module extracts analysis task information, evaluates the complexity of tasks in combination with medical data, and calculates a complexity evaluation coefficient; the task division and storage module divides the task into a simple task and a complex task according to the complexity evaluation coefficient, and encrypts and stores the data of the simple task; the task decomposition and parallel computing module decomposes a complex task into sub-tasks, analyzes a dependency relationship and decrypts a summarized result step by step through parallel computing encryption processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and specifically to a medical data analysis system and method based on homomorphic encryption technology. Background Art

[0002] With the rapid development of information technology, especially in the fields of medical care, medical insurance, drug and consumable management, big data and data analysis technologies are playing an increasingly important role in ensuring drug supply chain management, cost supervision, inventory management, sales analysis, etc. However, these fields involve a large amount of sensitive data, such as drug sales records, drug procurement and inventory data, etc. Therefore, how to ensure data privacy and security while performing efficient data analysis has become an urgent technical problem to be solved.

[0003] To solve this problem, homomorphic encryption technology, as an emerging encryption technology, can perform data operations in the encrypted state, and the calculation result is the same as the operation result on the original data, thus avoiding the risk of plaintext data leakage. This technology has the advantage of protecting data privacy and reduces the need for decryption operations to a certain extent. However, although homomorphic encryption technology has significant advantages in protecting data privacy, it still faces challenges in practical applications. For example, medical data includes drug sales records, procurement and inventory data, etc. Due to the different analysis tasks corresponding to different data, the corresponding data processing processes are also different. Although homomorphic encryption can perform basic operations such as addition and multiplication, its calculation range is relatively limited, and some complex operations may not be directly executable in the encrypted state. Therefore, in practical applications, especially in the process of medical data processing, a major challenge faced by homomorphic encryption technology is how to handle complex calculations and queries. Summary of the Invention

[0004] The purpose of the present invention is to provide a medical data analysis system and method based on homomorphic encryption technology to solve the problems raised in the above background art.

[0005] To solve the above technical problems, the present invention provides the following technical solutions: A medical data analysis method based on homomorphic encryption technology includes the following steps: Step S100. Obtain a medical data stream from a database, divide the medical data stream into independent segments according to a preset time window, further divide each independent segment according to the business dimension to generate a number of subsets; obtain a list of analysis tasks for the medical data, extract the requirement information for each analysis task from the list of analysis tasks, and correspond each analysis task to the corresponding subset; Step S200. Based on the correspondence between the analysis tasks and the subsets, and in combination with the requirement information of the analysis tasks and the medical data of the corresponding subsets, evaluate the complexity of the analysis tasks, so as to obtain a complexity evaluation coefficient; Step S300. According to the complexity evaluation coefficient of the analysis tasks, divide the analysis tasks in the analysis task list of the medical data into simple analysis tasks and complex analysis tasks; for the simple analysis tasks, directly encrypt and store the medical data of the corresponding subsets; Step S400. For complex analysis tasks, according to the corresponding requirement information, decompose the complex analysis tasks into several subtasks, analyze the dependency relationships between the subtasks, so as to obtain a subtask execution order list; for the decomposed subtasks, use a parallel computing architecture to perform parallel encryption processing on each subtask. After the subtasks are completed, according to the order list and execution results of the subtasks, summarize the results of the subtasks by means of step-by-step decryption.

[0006] Further, step S100 includes: S101. Obtain the medical data stream from the database, divide the medical data stream into independent segments according to a preset time window, and label them as {T1, T2,..., Tn}, where T1 represents the medical data segment corresponding to the first preset time window, T2 represents the medical data segment corresponding to the second preset time window, and so on, Tn represents the medical data segment corresponding to the nth preset time window, and n represents the number of independent segments obtained by dividing the acquired medical data stream according to the preset time window; for each independent segment Ti, where i ranges from 1 to n; divide according to the business dimension of the medical data corresponding to the independent segment Ti, where the business dimensions include dimensions such as procurement, sales, and inventory; classify the medical data belonging to the same business dimension into one category, so as to generate several subsets, labeled as {S1, S2,..., Sm}, where S1 represents the set composed of the medical data of the first business dimension, S2 represents the set composed of the medical data of the second business dimension, and so on, Sm represents the set composed of the medical data of the mth business dimension, and m represents the number of business dimensions; S102. Obtain the analysis task list of the medical data for each independent segment, extract the requirement information of each analysis task from the analysis task list, and obtain the business dimension of the medical data corresponding to each analysis task according to the requirement information; sequentially match the business dimension of the medical data corresponding to the analysis task with the business dimension corresponding to the subset Sj, where j ranges from 1 to m; according to the matching result, correspond each analysis task to the corresponding subset.

[0007] Further, step S200 includes: S201. Based on the correspondence between the analysis tasks and the subsets, for each analysis task, obtain the data storage amount of the medical data of the corresponding subset, and calculate the data scale index DSI of the analysis task according to the data storage amount of the medical data of the subset. The specific calculation formula is: DSI = [∑j = 1k Size(Sj)] / max(Size(S1), Size(S2),..., Size(Sm)); Among them, k represents the number of subsets corresponding to the analysis task, Size(Sj) represents the data storage amount of the medical data of the j-th subset, and max(Size(S1), Size(S2),..., Size(Sm)) represents the maximum value of the data storage amounts of all subsets S1, S2,..., Sm; through the requirement information of the analysis task, obtain the calculation process of the corresponding analysis task, count the number C of each type of calculation operation, and the operation frequency F corresponding to each operation type, and calculate the calculation density index CDI of the analysis task. The specific calculation formula is: CDI = [∑g = 1G (Cg × Fg)] / T; where Cg represents the number of the g-th type of calculation operation, that is, the amount of calculation involved in this operation type (such as the number of executions, the amount of calculation, etc.); Fg represents the operation frequency corresponding to the g-th type of calculation operation, T represents the total duration of the analysis task, and G represents the total number of types of calculation operations of the analysis task; S202. According to the data scale index DSI and the calculation density index CDI corresponding to the analysis task, evaluate the complexity of the analysis task, and thus calculate the complexity evaluation coefficient R. The corresponding calculation formula is: R = α × DSI + β × CDI. In the calculation formula of the complexity evaluation coefficient R, the data values corresponding to the data scale index DSI and the calculation density index CDI participate in the operation, and the corresponding units are not considered; where α and β respectively represent the weight coefficients corresponding to the data scale index DSI and the calculation density index CDI, and α + β = 1.

[0008] Further, step S300 includes: S301. For each independent segment, summarize the complexity evaluation coefficient R of each analysis task in the corresponding analysis task list; compare the complexity evaluation coefficient R with a preset threshold R0, where the preset threshold R0 is obtained based on historical experience, business requirements or expert advice; if the complexity evaluation coefficient R is less than the threshold R0, then classify the corresponding analysis task as a simple analysis task; if the complexity evaluation coefficient R is greater than or equal to the threshold R0, then classify the corresponding analysis task as a complex analysis task; S302. For each simple analysis task, obtain the corresponding subset Sj, perform homomorphic encryption on the medical data of each corresponding subset Sj, and the encrypted medical data is represented as: E(Sj) = Encrypt(Sj), where Encrypt represents the homomorphic encryption operation, Sj represents the original data subset, and perform corresponding data processing on the encrypted data E(Sj), and store the data processing result in the database.

[0009] Further, step S400 includes: S401. For each complex analysis task, obtain the corresponding requirement information, obtain the analysis process of the corresponding medical data according to the requirement information, and decompose the complex analysis task into several subtasks; for each subtask, calculate the complexity evaluation coefficient R1 corresponding to the subtask according to the calculation process of the complexity evaluation coefficient R of each analysis task in the analysis task list. If the complexity evaluation coefficient R1 corresponding to the subtask is greater than or equal to the threshold R0, continue to decompose until the complexity evaluation coefficients R1 of all subtasks are less than the threshold R0; S402. For each complex analysis task, summarize all subtasks to obtain the corresponding subtask set K, and K = {k1, k2,..., kv}, where k1 represents the first subtask, k2 represents the second subtask, and so on, kv represents the vth subtask, and v represents the number of subtasks after the decomposition of the complex analysis task; according to the subtask set K, regard each subtask as a node. If subtask ka needs to be completed before subtask kb, then there is a directed edge from ka to kb; traverse each subtask in the subtask set K to form a directed acyclic graph DAG; according to the directed acyclic graph DAG, perform topological sorting to obtain the subtask execution order list Sk; S403. For each subtask in the subtask set K, perform homomorphic encryption processing on the medical data required by the subtask before execution; for subtask ka, the corresponding medical data Da is encrypted into ciphertext Wa, and Wa = Encrypt(Da, Keya), where Keya represents the key used to encrypt the medical data of subtask ka, and Wa is the encrypted medical data; after all subtasks complete the encryption processing, perform corresponding medical data processing according to the order of the subtask execution order list Sk; for each subtask, select a suitable parallel computing architecture for processing. Common parallel computing architectures include: Multi-core CPU architecture: Utilize the computing power of multi-core processors to execute multiple subtasks simultaneously.

[0010] Distributed computing architecture: Parallelly execute subtasks through a cluster composed of multiple computers, which is suitable for processing large-scale data.

[0011] Graphics Processing Unit (GPU) Architecture: Utilize the parallel computing power of the GPU to process complex computational tasks, especially suitable for tasks that require a large amount of parallel computing.

[0012] After the subtasks are executed, decrypt the ciphertext results Wa of each subtask, expressed as: Qa = Decrypt(Wa, Keya); where Qa represents the result of the decrypted subtask; according to the subtask execution order list Sk, starting from the first subtask, decrypt the results of each subtask one by one. After decrypting the results of all subtasks, summarize the corresponding subtask results in order to obtain the analysis result of the complex analysis task, and store the analysis result in the database.

[0013] A medical data analysis system based on homomorphic encryption technology, including: a data processing and segmentation module, a task complexity evaluation module, a task partitioning and storage module, and a task decomposition and parallel computing module; The data processing and segmentation module obtains the medical data stream, segments the medical data stream into independent segments according to a preset time window, and further divides each independent segment according to the business dimension to generate several subsets; obtains the analysis task list of the medical data, extracts the requirement information of each analysis task from the analysis task list, and associates each analysis task with the corresponding subset; The task complexity evaluation module extracts the corresponding task information from the corresponding analysis task for each subset, and combines the task information and the corresponding medical data to evaluate the complexity of the analysis task, thereby obtaining the complexity evaluation coefficient; The task partitioning and storage module divides the analysis tasks in the analysis task list of the medical data into simple analysis tasks and complex analysis tasks according to the complexity evaluation coefficient and the association relationship between different subsets; for simple analysis tasks, directly encrypt and store the medical data of the corresponding subset; For complex analysis tasks, the task decomposition and parallel computing module decomposes the complex analysis task into several subtasks according to the corresponding task information, analyzes the dependency relationship between the subtasks, thereby obtaining the subtask execution order list; for the decomposed subtasks, uses the parallel computing architecture to perform parallel encryption processing on each subtask. After the subtasks are executed, according to the order list and execution results of the subtasks, summarize the results of the subtasks by gradually decrypting.

[0014] Furthermore, the data processing and segmentation module includes a data acquisition unit and a data segmentation unit; The data acquisition unit retrieves the medical data stream from the database, slices the medical data stream according to a preset time window, and slices the medical data into independent segments; the data segmentation unit further divides each independent segment according to the business dimension to generate several subsets; obtains the analysis task list of the medical data, extracts the requirement information of each analysis task from the analysis task list, and corresponds each analysis task to the corresponding subset.

[0015] Furthermore, the task complexity evaluation module includes a data scale evaluation unit, a calculation density evaluation unit, and a complexity evaluation unit; The data scale evaluation unit calculates the data scale index of each analysis task according to the storage amount of the medical data of the subset corresponding to each analysis task; the calculation density evaluation unit calculates the calculation density index of each task according to the requirement information of each analysis task; the complexity evaluation unit evaluates the complexity of the analysis task and calculates the complexity evaluation coefficient according to the data scale index and the calculation density index.

[0016] Furthermore, the task division and storage module includes a task classification unit and an encrypted storage unit; The task classification unit divides the analysis tasks into simple analysis tasks and complex analysis tasks according to the complexity evaluation coefficient of the analysis task and a preset threshold; for the simple analysis tasks, the encrypted storage unit encrypts the data subsets involved in the analysis tasks using the homomorphic encryption technology.

[0017] Furthermore, the task decomposition and parallel computing module includes a task decomposition unit, a dependency analysis unit, and a parallel encryption and execution unit; For the complex analysis tasks, the task decomposition unit decomposes the analysis tasks into several subtasks according to the requirement information and calculates the complexity evaluation coefficient of each subtask; if the complexity evaluation coefficient of the subtask still exceeds the threshold, continue to decompose until the complexity evaluation coefficients of all subtasks are lower than the threshold; the dependency analysis unit analyzes the dependencies between the subtasks in the task, constructs a directed acyclic graph between the subtasks, performs topological sorting according to the directed acyclic graph, and generates an execution order list of each subtask; the parallel encryption and execution unit homomorphically encrypts the medical data of each subtask, processes the subtasks using the parallel computing architecture, decrypts the data and results of each subtask one by one after execution, finally merges the results in order, and stores the analysis results of the complex tasks.

[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: The analysis method proposed by the present invention can dynamically divide tasks according to the complexity of the tasks, distinguishing complex analysis tasks from simple analysis tasks; for simple analysis tasks, the data is directly encrypted and stored, while for complex tasks, a strategy of decomposition and parallel computing is adopted, effectively improving the task processing efficiency. For complex analysis tasks, through the decomposition of tasks and the analysis of the dependency relationships of subtasks, the analysis process can be accelerated through parallel computing. When executing subtasks, by gradually decrypting and summarizing the results, the data privacy during the calculation process is ensured, while optimizing the speed and efficiency of data processing. The present invention can not only handle simple operations such as basic addition and multiplication, but also overcome the limitations of homomorphic encryption in dealing with complex calculations through the strategies of subtask decomposition and dynamic analysis of tasks; by decomposing tasks, encrypting each subtask, and gradually decrypting and summarizing, it is ensured that even in the face of complex calculation requirements, effective analysis can still be carried out while protecting data privacy. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the accompanying drawings: Figure 1 is a schematic diagram of the modules of the medical data analysis system based on homomorphic encryption technology of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] Please refer to Figure 1 , the present invention provides the following technical solutions: A medical data analysis system based on homomorphic encryption technology, comprising: a data processing and segmentation module, a task complexity evaluation module, a task division and storage module, and a task decomposition and parallel computing module; The data processing and segmentation module obtains the medical data stream, segments the medical data stream into independent segments according to a preset time window, further divides each independent segment according to the business dimension to generate several subsets; obtains the analysis task list of the medical data, extracts the requirement information of each analysis task from the analysis task list, and corresponds each analysis task to the corresponding subset; For each subset, the task complexity evaluation module extracts corresponding task information from the corresponding analysis tasks, and combines the task information and the corresponding medical data to evaluate the complexity of the analysis tasks, so as to obtain the complexity evaluation coefficient; The task division and storage module divides the analysis tasks in the analysis task list of medical data into simple analysis tasks and complex analysis tasks according to the complexity evaluation coefficient and the association relationship between different subsets; for simple analysis tasks, the medical data of the corresponding subset is directly encrypted and stored; For complex analysis tasks, the task decomposition and parallel computing module decomposes the complex analysis tasks into several subtasks according to the corresponding task information, analyzes the dependency relationship between the subtasks, so as to obtain the subtask execution sequence list; for the decomposed subtasks, a parallel computing architecture is used to perform parallel encryption processing on each subtask. After the subtasks are completed, according to the sequence list and execution results of the subtasks, the results of the subtasks are summarized by means of step-by-step decryption.

[0022] The data processing and segmentation module includes a data acquisition unit and a data segmentation unit; The data acquisition unit acquires the medical data stream from the database, segments the medical data stream according to a preset time window, and segments the medical data into independent segments; the data segmentation unit further divides each independent segment according to the business dimension to generate several subsets; obtains the analysis task list of medical data, extracts the requirement information of each analysis task from the analysis task list, and corresponds each analysis task to the corresponding subset.

[0023] The task complexity evaluation module includes a data scale evaluation unit, a calculation density evaluation unit, and a complexity evaluation unit; The data scale evaluation unit calculates the data scale index of each analysis task according to the storage volume of the medical data of the subset corresponding to each analysis task; the calculation density evaluation unit calculates the calculation density index of each task according to the requirement information of each analysis task; the complexity evaluation unit evaluates the complexity of the analysis task and calculates the complexity evaluation coefficient according to the data scale index and the calculation density index.

[0024] The task division and storage module includes a task classification unit and an encryption storage unit; The task classification unit divides the analysis tasks into simple analysis tasks and complex analysis tasks according to the complexity evaluation coefficient of the analysis tasks and a preset threshold; for simple analysis tasks, the encryption storage unit uses homomorphic encryption technology to encrypt the data subset involved in the analysis tasks.

[0025] The task decomposition and parallel computing module includes a task decomposition unit, a dependency relationship analysis unit, and a parallel encryption and execution unit; The task decomposition unit decomposes a complex analysis task into several subtasks according to the requirement information, and calculates the complexity evaluation coefficient of each subtask; if the complexity evaluation coefficient of a subtask still exceeds the threshold, it continues to decompose until the complexity evaluation coefficients of all subtasks are lower than the threshold; the dependency analysis unit analyzes the dependencies between the subtasks in the task, constructs a directed acyclic graph between the subtasks, performs a topological sort according to the directed acyclic graph, and generates an execution order list for each subtask; the parallel encryption and execution unit performs homomorphic encryption processing on the medical data of each subtask, processes the subtasks using a parallel computing architecture, decrypts the data and results of each subtask one by one after execution, finally merges the results in order, and stores the analysis results of the complex task.

[0026] A medical data analysis method based on homomorphic encryption technology includes the following steps: Step S100. Obtain a medical data stream from a database, split the medical data stream into independent segments according to a preset time window, and further divide each independent segment according to the business dimension to generate several subsets; obtain a list of analysis tasks for medical data, extract the requirement information of each analysis task from the analysis task list, and correspond each analysis task to the corresponding subset. Step S200. Based on the correspondence between the analysis task and the subset, and combining the requirement information of the analysis task and the medical data of the corresponding subset, evaluate the complexity of the analysis task, so as to obtain a complexity evaluation coefficient. Step S300. According to the complexity evaluation coefficient of the analysis task, divide the analysis tasks in the analysis task list of medical data into simple analysis tasks and complex analysis tasks; for simple analysis tasks, directly encrypt and store the medical data of the corresponding subset. Step S400. For complex analysis tasks, according to the corresponding requirement information, decompose the complex analysis task into several subtasks, analyze the dependencies between the subtasks, so as to obtain a subtask execution order list; for the decomposed subtasks, use a parallel computing architecture to perform parallel encryption processing on each subtask, and after the subtasks are executed, according to the order list and execution results of the subtasks, summarize the results of the subtasks by means of step-by-step decryption.

[0027] Step S100 includes: S101. Obtain the medical data stream from the database, divide the medical data stream into independent segments according to a preset time window, and label them as {T1, T2,..., Tn}, where T1 represents the medical data segment corresponding to the 1st preset time window, T2 represents the medical data segment corresponding to the 2nd preset time window, and so on. Tn represents the medical data segment corresponding to the nth preset time window, and n represents the number of independent segments obtained by dividing the acquired medical data stream according to the preset time window; for each independent segment Ti, where i ranges from 1 to n; divide according to the business dimension of the medical data corresponding to the independent segment Ti, where the business dimensions include dimensions such as procurement, sales, and inventory; classify the medical data belonging to the same business dimension into one category, thereby generating several subsets, labeled as {S1, S2,..., Sm}, where S1 represents the set composed of the medical data of the 1st business dimension, S2 represents the set composed of the medical data of the 2nd business dimension, and so on. Sm represents the set composed of the medical data of the mth business dimension, and m represents the number of business dimensions; S102. Obtain the analysis task list of the medical data for each independent segment, extract the requirement information of each analysis task from the analysis task list, and obtain the business dimension of the medical data corresponding to each analysis task according to the requirement information; sequentially match the business dimension of the medical data corresponding to the analysis task with the business dimension corresponding to the subset Sj, where j ranges from 1 to m; according to the matching result, correspond each analysis task to the corresponding subset.

[0028] In this embodiment, assume that a continuous medical data stream is obtained from the database, and the data stream contains multi-dimensional information such as procurement, sales, and inventory. According to the preset time window (for example, 1 hour, 1 day, etc.), this data stream is divided into multiple independent time periods, and each time period is called an "independent segment".

[0029] For example, assume that the preset time window is 1 day, and the medical data stream obtained from the database is: Time range: from 00:00:00 on March 1, 2025 to 23:59:59 on March 7, 2025; the medical data stream contains information such as procurement, sales, and inventory.

[0030] Divide this data stream into multiple independent segments according to the preset time window: T1: from 00:00:00 on March 1, 2025 to 23:59:59 on March 1, 2025; T2: from 00:00:00 on March 2, 2025 to 23:59:59 on March 2, 2025; And so on until Tn, where n is the number of segments divided.

[0031] The data in each time period Ti will be divided according to business dimensions (such as procurement, sales, inventory, etc.). The pharmaceutical data of each business dimension will be grouped into a subset. For example, the data in segment T1 may include the following dimensions: Procurement: Procurement data 1, Procurement data 2...; Sales: Sales data 1, Sales data 2...; Inventory: Inventory data 1, Inventory data 2...; These data will be grouped into different subsets S1, S2, S3, etc. according to business dimensions: S1: Data in the procurement dimension (including Procurement data 1, Procurement data 2, etc.); S2: Data in the sales dimension (including Sales data 1, Sales data 2, etc.); S3: Data in the inventory dimension (including Inventory data 1, Inventory data 2, etc.); For each Ti (such as T2, T3, etc.), the business dimension division will also be carried out in the same way to generate the corresponding subsets S1, S2, S3, etc.

[0032] The system will extract the requirement information of each analysis task from the analysis task list. For example, assume the analysis task list is as follows: Task 1: Analyze the procurement data of a specific drug; Task 2: Analyze the sales data of a specific drug; Task 3: Analyze the procurement and sales trends of a specific drug and calculate the inventory turnover rate; The requirement information of each analysis task clarifies the business dimension to be analyzed. For Task 1, the requirement information indicates that the data in the "procurement" dimension needs to be analyzed; for Task 2, the requirement information indicates that the data in the "sales" dimension needs to be analyzed; for Task 3, the requirement information indicates that the data in the "procurement, sales, and inventory" dimensions needs to be analyzed.

[0033] According to the business dimension of the analysis task, it will be matched with the business dimension in subset Sj; Task 1 (procurement task) needs to analyze procurement data, so it will be matched with all subsets containing the procurement dimension; for example, S1 contains data in the procurement dimension, and Task 1 will correspond to S1. Task 2 (sales task) needs to analyze sales data, so it will be matched with all subsets containing the sales dimension. For example, S2 contains data in the sales dimension, and Task 2 will correspond to S2. Task 3 (inventory turnover rate) needs to analyze procurement, sales, and inventory data, so it will be matched with all subsets containing the procurement, sales, and inventory dimensions. For example, S1, S2, and S3 contain data in the procurement, sales, and inventory dimensions, and Task 3 will correspond to S1, S2, and S3.

[0034] Finally, analysis tasks 1, 2, and 3 are matched with the corresponding subsets to form the following mapping: Task 1 → S1 (purchase data); Task 2 → S2 (sales data); Task 3 → S1 (purchase data), S2 (sales data), and S3 (inventory data).

[0035] Step S200 includes: S201. Based on the correspondence between the analysis task and the subset, for each analysis task, obtain the data storage amount of the medical data of the corresponding subset, and calculate the data scale index DSI of the analysis task according to the data storage amount of the medical data of the subset. The specific calculation formula is: DSI = [∑k j = 1 Size(Sj)] / max(Size(S1), Size(S2),..., Size(Sm)); where k represents the number of subsets corresponding to the analysis task, Size(Sj) represents the data storage amount of the medical data of the j-th subset, and max(Size(S1), Size(S2),..., Size(Sm)) represents the maximum data storage amount among all subsets S1, S2,..., Sm; through the requirement information of the analysis task, obtain the calculation process of the corresponding analysis task, count the number C of each type of calculation operation, and the operation frequency F corresponding to each operation type, and calculate the calculation density index CDI of the analysis task. The specific calculation formula is: CDI = [∑G g = 1 (Cg × Fg)] / T; where Cg represents the number of the g-th type of calculation operation, that is, the amount of calculation involved in this operation type (such as the number of executions, the amount of calculation, etc.); Fg represents the operation frequency corresponding to the g-th type of calculation operation, T represents the total duration of the analysis task, and G represents the total number of types of calculation operations of the analysis task; S202. According to the data scale index DSI and the calculation density index CDI corresponding to the analysis task, evaluate the complexity of the analysis task, and thus calculate the complexity evaluation coefficient R. The corresponding calculation formula is: R = α × DSI + β × CDI. In the calculation formula of the complexity evaluation coefficient R, the data values corresponding to the data scale index DSI and the calculation density index CDI participate in the operation, and the corresponding units are not considered; where α and β respectively represent the weight coefficients corresponding to the data scale index DSI and the calculation density index CDI, and α + β = 1.

[0036] In this embodiment, taking Task 3 as an example, assume the following: Task 3 involves S1 (purchase data), S2 (sales data), and S3 (inventory data), and their storage amounts are 100GB, 200GB, and 50GB respectively; the sum of the storage amounts = 100GB + 200GB + 50GB = 350GB; The maximum storage capacity is max(100GB, 200GB, 50GB) = 200GB; DSI = (350GB) / (200GB) = 1.75.

[0037] The calculation operations for Task 3 are as follows: Operation 1: Execute 80 times, with a frequency of 0.1 times per second.

[0038] Operation 2: Execute 40 times, with a frequency of 0.15 times per second.

[0039] Operation 3: Execute 10 times, with a frequency of 0.3 times per second.

[0040] Total duration: 15 seconds.

[0041] CDI = [(80 × 0.1) + (40 × 0.15) + (10 × 0.3)] / 15 = (8 + 6 + 3) / 15 = 1.133.

[0042] According to the data scale index DSI and the calculation density index CDI, calculate the complexity evaluation coefficient R. Assume the weight coefficients α = 0.6 and β = 0.4.

[0043] R = (0.6 × 1.75) + (0.4 × 1.133) = 1.05 + 0.4532 = 1.5032.

[0044] Step S300 includes: S301. For each independent segment, summarize the complexity evaluation coefficient R of each analysis task in the corresponding analysis task list; compare the complexity evaluation coefficient R with a preset threshold R0, where the preset threshold R0 is obtained based on historical experience, business requirements, or expert advice; if the complexity evaluation coefficient R is less than the threshold R0, then classify the corresponding analysis task as a simple analysis task; if the complexity evaluation coefficient R is greater than or equal to the threshold R0, then classify the corresponding analysis task as a complex analysis task; S302. For each simple analysis task, obtain the corresponding subset Sj, perform homomorphic encryption on the medical data of each corresponding subset Sj, and the encrypted medical data is represented as: E(Sj) = Encrypt(Sj), where Encrypt represents the homomorphic encryption operation, Sj represents the original data subset, and perform corresponding data processing on the encrypted data E(Sj), and store the data processing result in the database.

[0045] Step S400 includes: S401. For each complex analysis task, obtain the corresponding requirement information, and based on the requirement information, obtain the analysis process of the corresponding medical data, and decompose the complex analysis task into several sub-tasks; for each sub-task, calculate the complexity evaluation coefficient R1 corresponding to the sub-task according to the calculation process of the complexity evaluation coefficient R of each analysis task in the analysis task list. If the complexity evaluation coefficient R1 corresponding to the sub-task is greater than or equal to the threshold R0, continue to decompose until the complexity evaluation coefficients R1 corresponding to all sub-tasks are less than the threshold R0. S402. For each complex analysis task, summarize all sub-tasks to obtain the corresponding sub-task set K, and K = {k1, k2,..., kv}, where k1 represents the first sub-task, k2 represents the second sub-task, and so on, kv represents the vth sub-task, and v represents the number of sub-tasks after the complex analysis task is decomposed; according to the sub-task set K, regard each sub-task as a node. If the sub-task ka needs to be completed before the sub-task kb, then there is a directed edge from ka to kb; traverse each sub-task in the sub-task set K to form a directed acyclic graph DAG; according to the directed acyclic graph DAG, perform topological sorting to obtain the sub-task execution order list Sk. In this embodiment, since topological sorting is a sorting algorithm for a directed acyclic graph (DAG), the purpose is to sort all nodes in the graph so that for each directed edge ka → kb, the node ka appears before the node kb in the sorting.

[0046] Basic steps of topological sorting: Initialization: First, find all sub-tasks without pre-dependencies; that is, nodes (sub-tasks) with an in-degree of 0, and they can be executed independently. Gradually remove nodes: Select a node with an in-degree of 0 to execute and remove it from the graph; after removing the node, the edges connected to it in the graph will also be deleted.

[0047] Update the in-degree: For the nodes connected to the deleted node, update their in-degrees. If the in-degree of a certain node becomes 0, it can also start to execute.

[0048] Repeat execution: Repeat this process until all sub-tasks are executed and sorted.

[0049] Among them, the specific implementation of topological sorting usually uses depth-first search (DFS) or breadth-first search (BFS) algorithms.

[0050] S403. For each subtask in the subtask set K, the medical data required by the subtask is homomorphically encrypted before execution; for the subtask ka, the corresponding medical data Da is encrypted into the ciphertext Wa, and Wa = Encrypt(Da, Keya), where Keya represents the key used to encrypt the medical data of the subtask ka, and Wa is the encrypted medical data; after all subtasks are encrypted, the corresponding medical data processing is performed according to the order of the subtask execution order list Sk; for each subtask, a suitable parallel computing architecture is selected for processing. Common parallel computing architectures include: Multi-core CPU architecture: Utilize the computing power of multi-core processors to execute multiple subtasks simultaneously.

[0051] Distributed computing architecture: Parallelly execute subtasks through a cluster composed of multiple computers, which is suitable for processing large-scale data.

[0052] Graphics processing unit (GPU) architecture: Utilize the parallel computing power of GPUs to process complex computing tasks, especially suitable for tasks that require a large amount of parallel computing.

[0053] After the subtasks are executed, decrypt the ciphertext result Wa of each subtask, expressed as: Qa = Decrypt(Wa, Keya); where Qa represents the result of the decrypted subtask; according to the subtask execution order list Sk, starting from the first subtask, decrypt the results of each subtask one by one. After the results of all subtasks are decrypted, summarize the corresponding subtask results in order to obtain the analysis result of the complex analysis task, and store the analysis result in the database.

[0054] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0055] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A medical data analysis method based on homomorphic encryption technology, characterized in that: The method comprises the following steps: Step S100. Obtain a medical data stream from a database, divide the medical data stream into independent segments according to a preset time window, further divide each independent segment according to a business dimension, and generate a number of subsets; obtain an analysis task list for medical data, extract the requirement information of each analysis task from the analysis task list, and correspond each analysis task to a corresponding subset; Step S200. Based on the correspondence between the analysis task and the subset, and in combination with the requirement information of the analysis task and the medical data of the corresponding subset, the complexity of the analysis task is evaluated to obtain a complexity evaluation coefficient; Step S300. According to the complexity evaluation coefficient of the analysis task, the analysis tasks in the analysis task list of the medical data are divided into simple analysis tasks and complex analysis tasks; for the simple analysis task, the medical data of the corresponding subset is directly encrypted and stored; Step S400. For complex analysis tasks, according to the corresponding demand information, the complex analysis tasks are decomposed into several subtasks, and the dependencies between the subtasks are analyzed to obtain the subtask execution sequence table; for the decomposed subtasks, each subtask is encrypted in parallel using a parallel computing architecture. After the subtasks are executed, the results of the subtasks are summarized by step-by-step decryption based on the subtask sequence table and execution results.

2. The medical data analysis method based on homomorphic encryption technology according to claim 1, characterized in that: The step S100 includes: S101. Obtain a medical data stream from a database, and divide the medical data stream into independent segments according to a preset time window, marked as {T1, T2, ..., Tn}, wherein T1 represents a medical data segment corresponding to the first preset time window, T2 represents a medical data segment corresponding to the second preset time window, and so on, Tn represents a medical data segment corresponding to the nth preset time window, and n represents the number of independent segments into which the obtained medical data stream is divided according to the preset time window; for each independent segment Ti, where i ranges from 1 to n; divide the medical data according to the business dimension corresponding to the independent segment Ti, and classify the medical data belonging to the same business dimension into one category, thereby generating several subsets, marked as {S1, S2, ..., Sm}, wherein S1 represents a set consisting of medical data of the first business dimension, S2 represents a set consisting of medical data of the second business dimension, and so on, Sm represents a set consisting of medical data of the mth business dimension, and m represents the number of business dimensions; S102. Obtain the analysis task list of the medical data for each independent segment, extract the requirement information of each analysis task from the analysis task list, and obtain the business dimension of the medical data corresponding to each analysis task based on the requirement information; match the business dimension of the medical data corresponding to the analysis task with the business dimension corresponding to the subset Sj in sequence, where j ranges from 1 to m; and match each analysis task with the corresponding subset based on the matching result.

3. The medical data analysis method based on homomorphic encryption technology according to claim 2 is characterized in that: The step S200 includes: S201. Based on the correspondence between the analysis task and the subset, for each analysis task, the data storage capacity of the corresponding subset of medical data is obtained, and the data scale index DSI of the analysis task is calculated according to the data storage capacity of the subset of medical data. The specific calculation formula is: DSI=[∑kj=1Size(Sj)] / max(Size(S1),Size(S2),...,Size(Sm)); Wherein, k represents the number of subsets corresponding to the analysis task, Size(Sj) represents the data storage capacity of the medical data of the jth subset, and max(Size(S1), Size(S2), ..., Size(Sm)) represents the maximum value of the data storage capacity in all subsets S1, S2, ..., Sm. By analyzing the demand information of the task, the calculation process of the corresponding analysis task is obtained, the number C of each type of calculation operation and the operation frequency F corresponding to each operation type are counted, and the calculation density index CDI of the analysis task is calculated. The specific calculation formula is: CDI=[∑G g=1(Cg×Fg)] / T; where Cg represents the number of the gth type of calculation operation; Fg represents the operation frequency corresponding to the gth type of calculation operation, T represents the total duration of the analysis task, and G represents the total number of calculation operation types of the analysis task. S202. Evaluate the complexity of the analysis task according to the data scale indicator DSI and the calculation density index CDI corresponding to the analysis task, and thus calculate the complexity evaluation coefficient R, and the corresponding calculation formula is: R=α×DSI+β×CDI, where α and β represent the weight coefficients corresponding to the data scale indicator DSI and the calculation density index CDI respectively, and α+β=1.

4. The medical data analysis method based on homomorphic encryption technology according to claim 3 is characterized in that: The step S300 includes: S301. For each independent segment, summarize the complexity evaluation coefficient R of each analysis task in the corresponding analysis task list; compare the complexity evaluation coefficient R with the preset threshold R0, if the complexity evaluation coefficient R is less than the threshold R0, the corresponding analysis task is classified as a simple analysis task; if the complexity evaluation coefficient R is greater than or equal to the threshold R0, the corresponding analysis task is classified as a complex analysis task; S302. For each simple analysis task, obtain the corresponding subset Sj, and perform homomorphic encryption on the medical data of each corresponding subset Sj. The encrypted medical data is expressed as: E(Sj)=Encrypt(Sj), where Encrypt represents the homomorphic encryption operation, Sj represents the original data subset, and the encrypted data E(Sj) is processed accordingly, and the data processing results are stored in the database.

5. The medical data analysis method based on homomorphic encryption technology according to claim 4 is characterized in that: The step S400 includes: S401. For each complex analysis task, obtain the corresponding demand information, obtain the corresponding analysis process of medical data according to the demand information, and decompose the complex analysis task into several subtasks; for each subtask, calculate the complexity evaluation coefficient R1 corresponding to the subtask according to the calculation process of the complexity evaluation coefficient R of each analysis task in the analysis task list; if the complexity evaluation coefficient R1 corresponding to the subtask is greater than or equal to the threshold R0, continue to decompose until the complexity evaluation coefficients R1 corresponding to all subtasks are less than the threshold R0; S402. For each complex analysis task, summarize all subtasks to obtain the corresponding subtask set K, and K={k1,k2,...,kv}, where k1 represents the first subtask, k2 represents the second subtask, and so on, kv represents the vth subtask, and v represents the number of subtasks after the complex analysis task is decomposed; according to the subtask set K, each subtask is regarded as a node, if the subtask ka needs to be completed before the subtask kb, then there is a directed edge from ka to kb; traverse to each subtask in the subtask set K, so as to form a directed acyclic graph DAG; according to the directed acyclic graph DAG, perform topological sorting to obtain the subtask execution order list Sk; S403. For each subtask in the subtask set K, the medical data required for the subtask is homomorphically encrypted before execution; for subtask ka, the corresponding medical data Da is encrypted into ciphertext Wa, and Wa=Encrypt(Da,Keya), where Keya represents the key used to encrypt the medical data of subtask ka, and Wa is the encrypted medical data; after all subtasks have completed the encryption processing, the corresponding medical data processing is performed according to the order of the subtask execution sequence list Sk; after the subtask is executed, the ciphertext result Wa of each subtask is decrypted, expressed as: Qa=Decrypt(Wa,Keya); where Qa represents the decrypted subtask result; according to the subtask execution sequence list Sk, starting from the first subtask, the result of each subtask is decrypted one by one. After the results of all subtasks are decrypted, the corresponding subtask results are summarized in order to obtain the analysis results of the complex analysis task, and the analysis results are stored in the database.

6. A medical data analysis system based on homomorphic encryption technology, applied to a medical data analysis method based on homomorphic encryption technology as described in any one of claims 1 to 5, characterized in that: The system includes: a data processing and segmentation module, a task complexity evaluation module, a task division and storage module, and a task decomposition and parallel computing module; The data processing and segmentation module obtains the medical data stream, segments the medical data stream into independent segments according to a preset time window, and further divides each independent segment according to a business dimension to generate a plurality of subsets; obtains an analysis task list of the medical data, extracts the requirement information of each analysis task from the analysis task list, and matches each analysis task with a corresponding subset; The task complexity evaluation module extracts corresponding task information from the corresponding analysis task for each subset, and evaluates the complexity of the analysis task by combining the task information and the corresponding medical data, thereby obtaining a complexity evaluation coefficient; The task division and storage module divides the analysis tasks in the analysis task list of the medical data into simple analysis tasks and complex analysis tasks according to the complexity evaluation coefficient and the correlation between different subsets; for the simple analysis task, the medical data of the corresponding subset is directly encrypted and stored; For complex analysis tasks, the task decomposition and parallel computing module decomposes the complex analysis task into several subtasks according to the corresponding task information, analyzes the dependencies between the subtasks, and obtains the subtask execution sequence table; for the decomposed subtasks, each subtask is encrypted in parallel using the parallel computing architecture, and after the subtasks are executed, the results of the subtasks are summarized by step-by-step decryption according to the subtask sequence table and execution results.

7. The medical data analysis system based on homomorphic encryption technology according to claim 6, characterized in that: The data processing and segmentation module includes a data acquisition unit and a data segmentation unit; The data acquisition unit acquires the medical data stream from the database, divides the medical data stream according to a preset time window, and divides the medical data into independent segments; the data segmentation unit further divides each independent segment according to the business dimension to generate a plurality of subsets; acquires the analysis task list of the medical data, extracts the requirement information of each analysis task from the analysis task list, and matches each analysis task with the corresponding subset.

8. The medical data analysis system based on homomorphic encryption technology according to claim 6, characterized in that: The task complexity evaluation module includes a data scale evaluation unit, a calculation density evaluation unit and a complexity evaluation unit; The data scale assessment unit calculates the data scale index of each analysis task according to the storage volume of medical data of the subset corresponding to each analysis task; the computational density assessment unit calculates the computational density index of each analysis task according to the demand information of each analysis task; the complexity assessment unit assesses the complexity of the analysis task according to the data scale index and the computational density index and calculates the complexity assessment coefficient.

9. The medical data analysis system based on homomorphic encryption technology according to claim 6, characterized in that: The task division and storage module includes a task classification unit and an encryption storage unit; The task classification unit divides the analysis tasks into simple analysis tasks and complex analysis tasks according to the complexity evaluation coefficient of the analysis tasks and a preset threshold; the encryption storage unit uses homomorphic encryption technology to encrypt the data subset involved in the simple analysis tasks.

10. The medical data analysis system based on homomorphic encryption technology according to claim 6, characterized in that: The task decomposition and parallel computing module includes a task decomposition unit, a dependency analysis unit and a parallel encryption and execution unit; The task decomposition unit decomposes the complex analysis task into several subtasks according to the demand information, and calculates the complexity evaluation coefficient of each subtask; If the complexity evaluation coefficient of the subtask still exceeds the threshold, continue to decompose until the complexity evaluation coefficients of all subtasks are lower than the threshold; The dependency analysis unit analyzes the dependencies between the subtasks in the task, constructs a directed acyclic graph between the subtasks, performs topological sorting based on the directed acyclic graph, and generates an execution order list for each subtask; the parallel encryption and execution unit performs homomorphic encryption on the medical data of each subtask, processes the subtasks using a parallel computing architecture, decrypts the data and results of each subtask one by one after execution, and finally merges the results in sequence, and stores the analysis results of complex tasks.

Citation Information

Patent Citations

  • Multi-agent coordination based management task dynamic decomposition method and system

    CN102063329A

  • Security detection method and device for computing task

    CN112487415A

  • Big data task scheduling method, device and system

    CN114116148A

  • Mass data summarization method and system based on task chain and divide-and-conquer method

    CN118672790A

  • Data processing method, system and device, computer equipment and storage medium

    CN119293815A