A smart retrieval system for health records of kidney disease patients
By using an ASIC parallel computing architecture and native homomorphic encryption circuits to perform feature extraction and pathological attention reasoning in ciphertext, generating hash codes and constructing inverted indexes, the privacy and efficiency issues in kidney disease health record retrieval are resolved, and efficient and secure pathological feature matching is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GENERAL HOSPITAL OF NUCLEAR IND
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-30
AI Technical Summary
In existing technologies, the retrieval of kidney disease health records suffers from problems such as high risk of privacy leakage, inaccurate pathological feature mining, low retrieval efficiency, and low pathological semantic matching degree.
Employing an ASIC parallel computing architecture and native homomorphic encryption circuit, long-term features of nephropathy and pathological attention inference are extracted without decrypting the ciphertext. This generates high-dimensional feature vectors of the ciphertext and pathological attention weights of the plaintext. A specific hash code is generated through a nephropathy semantic quantizer, and an inverted index is constructed to achieve extremely fast and accurate matching and fuzzy search.
This ensures the privacy and security of kidney disease medical data and the accuracy of feature mining, and significantly improves the efficiency of record retrieval and the degree of pathological semantic matching.
Smart Images

Figure CN121919274B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical data processing technology, specifically to an intelligent retrieval system for health records of kidney disease patients. Background Technology
[0002] Kidney disease is characterized by its long course, numerous complications, and the need for long-term follow-up monitoring. This results in the accumulation of massive amounts of heterogeneous, multi-source health record data in clinical and research settings, including electronic medical records, laboratory test data, long-term physiological indicators, medication records, and follow-up information. This data covers the entire process from initial diagnosis and treatment intervention to long-term follow-up, forming the core foundation for precise clinical diagnosis and treatment, disease monitoring, prognostic assessment, and kidney disease-related research analysis. The availability, security, and retrieval efficiency of this data directly impact the quality of kidney disease diagnosis and treatment and the progress of research.
[0003] With the rapid development of medical informatization, the amount of kidney disease health record data is growing exponentially. The multi-source and heterogeneous nature of this data (such as inconsistent data formats across different hospital information systems, differences in output data standards from laboratory equipment, and varying degrees of structuring between electronic medical records and follow-up records) further exacerbates the difficulty of record management and retrieval. However, existing technologies for retrieving kidney disease patient health records suffer from high privacy risks, inaccurate pathological feature mining, low retrieval efficiency, and low pathological semantic matching.
[0004] Based on this, the present invention provides an intelligent retrieval system for health records of kidney disease patients to solve the aforementioned technical problems. Summary of the Invention
[0005] The purpose of this invention is to provide an intelligent retrieval system for health records of kidney disease patients. This invention utilizes an ASIC parallel computing architecture and native homomorphic encryption circuit to efficiently extract long-term temporal features of kidney disease and perform pathological attention inference without decrypting the ciphertext. It accurately outputs high-dimensional feature vectors of the ciphertext and pathological attention weights of the plaintext. Based on these weights and vectors, a kidney disease semantic quantizer generates specific hash codes that preserve the pathological semantics of kidney disease and constructs an efficient inverted index. This enables extremely fast and accurate matching and fuzzy search of kidney disease health records, thereby ensuring the privacy and security of kidney disease medical data and the accuracy of feature mining, while significantly improving the efficiency of record retrieval and the degree of pathological semantic matching.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] This invention provides an intelligent retrieval system for health records of kidney disease patients, comprising a data acquisition and preprocessing unit, an ASIC dedicated computing unit, a high-speed retrieval and storage unit, and a security and privacy management unit, wherein:
[0008] The data acquisition and preprocessing unit is used to acquire raw kidney disease patient record data from multi-source heterogeneous medical devices and databases, and to preprocess the acquired record data.
[0009] The ASIC dedicated computing unit is used to utilize the ASIC parallel computing architecture and native homomorphic encryption circuit to perform long-term feature extraction and pathological attention inference tasks of kidney disease in parallel under ciphertext state, and output ciphertext high-dimensional feature vector and corresponding plaintext pathological attention weights.
[0010] The high-speed retrieval and storage unit is used to compress the encrypted high-dimensional feature vector into a specific hash code based on the plaintext pathological attention weight using a nephrology semantic quantizer, and to build an inverted index based on the specific hash code to perform extremely fast and accurate matching and fuzzy search on nephrology feature fields.
[0011] The security and privacy management unit is used to conduct compliant and secure management of the entire process of storing, transmitting, and retrieving kidney disease health records through a hardware encryption engine and dynamic permission control policies.
[0012] The data acquisition and preprocessing unit includes a multi-source data access module, a data cleaning and standardization module, and a data format unification module, wherein:
[0013] The multi-source data access module is used to connect to multi-source heterogeneous terminals such as hospital information systems, testing equipment, and electronic medical records to collect original file data of kidney disease patients.
[0014] The data cleaning and standardization module is used to clean, deduplicate, and complete missing, redundant, and abnormal kidney disease data.
[0015] The data format unification module is used to convert data with different structures into a standard data format that the system can process.
[0016] The ASIC dedicated computing unit includes an ASIC parallel computing module, a homomorphic encryption hardware module, a long-term feature extraction module, and a pathological attention inference module, wherein:
[0017] The ASIC parallel computing module is used to provide dedicated hardware parallel computing power to support high-concurrency medical data processing in dense conditions.
[0018] The homomorphic encryption hardware module is used to encrypt data at the hardware level and supports direct computation in ciphertext without decryption.
[0019] The long-term feature extraction module is used to solidify and deploy a time-series analysis algorithm specifically for kidney disease, and to mine long-term physiological change trends in patient history records in parallel within encrypted text.
[0020] The pathological attention reasoning module is used to calculate the contribution of different kidney disease indicators to the current pathological state based on the input temporal characteristics, output a high-dimensional feature vector of encrypted text, and output the calculated attention weights as plaintext pathological attention weights after on-chip decryption.
[0021] The homomorphic encryption hardware module performs data encryption at the hardware level, supporting direct computation in ciphertext without decryption. The specific operation is as follows:
[0022] A1: Encrypt the received kidney disease patient file data at the hardware level, generate and output ciphertext data;
[0023] A2: Call the homomorphic operation circuit to directly perform homomorphic addition and homomorphic multiplication operations on the received ciphertext data without decryption, and generate intermediate ciphertext operation results;
[0024] A3: Output the intermediate ciphertext operation results to the long-term feature extraction module and the pathological attention reasoning module respectively.
[0025] The long-term feature extraction module is equipped with a time-series analysis algorithm specifically designed for kidney disease. This algorithm mines long-term physiological trends in patient history records in parallel within encrypted data. The specific operation is as follows:
[0026] B1: Perform time-series fragmentation on the received confidential document archive data, dividing the patient's historical archive data into segments according to the time dimension;
[0027] B2: In encrypted state, homomorphic weighted accumulation circuits are used to process data of each time segment in parallel to extract the temporal variation characteristics of key information of patient physiological indicators and pathological diagnosis.
[0028] The specific expression for extracting the temporal features of renal physiological indicators in encrypted state is as follows:
[0029] ;
[0030] In the formula, This is the ciphertext time-series feature vector. Here, T is the homomorphic encryption function, and T is the total number of time windows. The weight of the t-th time window of the preset plaintext, The encrypted state of renal physiological indicators within the t-th time window;
[0031] B3: Using the homomorphic difference calculation unit, calculate the difference between ciphertext data segments in adjacent time windows and multiply it by the plaintext trend smoothing coefficient to generate a quantified value of the long-term trend of the ciphertext.
[0032] The specific expression for the long-term physiological trend quantification value of kidney disease patients is as follows:
[0033] ;
[0034] In the formula, This is a quantified value of the long-term trend in the encrypted state. This is the trend smoothing coefficient. This represents the change in renal physiological indicators within adjacent time windows;
[0035] B4: The long-term trend quantification value of the ciphertext is fused into the temporal feature vector of the ciphertext through homomorphic addition and then transmitted to the pathological attention inference module.
[0036] The pathological attention reasoning module calculates the contribution of different kidney disease indicators to the current pathological state based on the input temporal characteristics using hardware acceleration, outputs a high-dimensional ciphertext feature vector, and outputs the calculated attention weights as plaintext pathological attention weights after on-chip decryption. The specific operation is as follows:
[0037] C1: Receives the ciphertext time-series feature vector from the long-time-series feature extraction module and maps it into a ciphertext triplet structure of query vector, key vector and value vector through hardware mapping logic;
[0038] C2: Without performing decryption, the homomorphic dot product operation unit is called to calculate the ciphertext similarity score between the ciphertext query vector and the ciphertext key vector. The score is then approximated by a polynomial approximation circuit with nonlinear activation to simulate the normalization characteristics of the Softmax function and generate ciphertext pathological attention weights. Subsequently, the on-chip decryption unit converts them into plaintext pathological attention weights.
[0039] The specific expression for generating the attention weight in encrypted pathology is as follows:
[0040] ;
[0041] In the formula, For encrypted pathology attention weights, The Softmax function approximation algorithm is performed for hardware, where Q is the ciphertext query vector and K is the ciphertext key vector. For vector dimensions, To query the ciphertext dot product of the vector and the key vector;
[0042] C3: Using a homomorphic weighted accumulation array, perform batch homomorphic multiplication and accumulation operations on the plaintext pathological attention weights and the ciphertext value vectors to aggregate and generate ciphertext high-dimensional feature vectors;
[0043] The specific expression for the aggregation of high-dimensional feature vectors of encrypted text is as follows:
[0044] ;
[0045] In the formula, V is the generated high-dimensional feature vector of the ciphertext, and V is the ciphertext value vector. Add an accumulation operation to the homomorphic multiplication of the attention weights and value vectors in the encrypted pathology;
[0046] C4: Synchronously outputs the encrypted high-dimensional feature vector and the plaintext pathology attention weights.
[0047] The high-speed retrieval and storage unit includes a nephrology semantic quantizer module, an inverted index construction module, and a hybrid retrieval execution module, wherein:
[0048] The nephrology semantic quantizer module is used to perform weighted hashing on the high-dimensional feature vector of the ciphertext according to the input ciphertext pathology attention weight, and compress it into a compact, specific hash code that retains the semantic information of nephrology.
[0049] The inverted index construction module: Based on the generated specific hash codes, it establishes an inverted index table where the specific hash codes point to specific patient records;
[0050] The hybrid retrieval execution module is used to receive retrieval requests and supports both ultra-fast and precise matching based on specific hash codes and approximate fuzzy search based on feature vectors.
[0051] The nephrology semantic quantizer module performs weighted hashing on the high-dimensional feature vector based on the input encrypted pathological attention weights, compressing it into a compact, specific hash code that retains the semantic information of nephrology. The specific operation is as follows:
[0052] D1: Receive the plaintext pathological attention weights from the pathological attention reasoning module, use them as dynamic gating signals, adjust the preset plaintext initial projection matrix coefficients, and generate a plaintext dynamic projection matrix that adapts to the current pathological state.
[0053] D2: Using the plaintext dynamic projection matrix, without decryption, a homomorphic linear transformation is performed on the high-dimensional feature vector of the ciphertext to map the pathological semantic similarity in the high-dimensional space to linearly separable features in the low-dimensional space, generating the low-dimensional feature vector of the ciphertext.
[0054] D3: Using a polynomial fitting activation function, nonlinear activation and binarization truncation are performed on the low-dimensional feature vector of the ciphertext to compress the continuous ciphertext values into a 0 / 1 bit sequence and generate a specific hash code.
[0055] D4: Output the specific hash code to the inverted index building module.
[0056] The specific expressions involved in steps D1-D3 are as follows:
[0057] The specific expression for adjusting the weighted projection matrix in D1 is as follows:
[0058] ;
[0059] In the formula, This is the weighted projection matrix. The preset initial projection matrix of plaintext. For Hadama accumulation, This is the weighting adjustment coefficient. Let be the plaintext pathology attention weight vector as input, and 1 be the input... A vector consisting entirely of 1s with the same dimension;
[0060] The specific expression for the dimension-reduction linear transformation in D2 is as follows:
[0061] ;
[0062] In the formula, The reduced-dimensional feature vectors are the result of dimensionality reduction. The input is a high-dimensional feature vector of the ciphertext;
[0063] The specific expression for generating the binary specific hash code in D3 is as follows:
[0064] ;
[0065] In the formula, H is the final generated specific hash code. It is a non-linear activation function. This is the binarization truncation threshold. The sign function maps positive values to 1 and negative values to 0.
[0066] The security and privacy management unit includes a hardware encryption engine module, a dynamic access control module, and a full-process audit and tracing module, wherein:
[0067] The hardware encryption engine module is used to provide real-time encryption and decryption services for the archives and search results in the system without occupying the main CPU.
[0068] The dynamic access control module: based on the principle of minimum necessity, dynamically allocates data access granularity according to the role of the searcher, and controls the anonymization level and visibility range of the search results;
[0069] The full-process audit and tracking module is used to record all data access, retrieval, and permission change operation logs in an immutable manner.
[0070] Compared with the prior art, the beneficial effects of the present invention are:
[0071] This invention utilizes an ASIC parallel computing architecture and native homomorphic encryption circuit to efficiently extract long-term features of kidney disease and perform pathological attention inference without decrypting the ciphertext. It accurately outputs high-dimensional feature vectors of the ciphertext and pathological attention weights of the plaintext. Based on these weights and vectors, a kidney disease semantic quantizer generates specific hash codes that preserve the pathological semantics of kidney disease and constructs an efficient inverted index. This enables extremely fast and accurate matching and fuzzy search of kidney disease health records, thereby ensuring the privacy and security of kidney disease medical data and the accuracy of feature mining, while significantly improving the efficiency of record retrieval and the degree of pathological semantic matching. Attached Figure Description
[0072] Figure 1 This is a system diagram of an intelligent retrieval system for health records of kidney disease patients according to the present invention.
[0073] Figure 2 This is a timing flowchart of the internal ASIC dedicated computing unit in an intelligent retrieval system for health records of kidney disease patients according to the present invention.
[0074] Figure 3 This is a flowchart of the semantic quantification and retrieval construction process for kidney disease in an intelligent retrieval system for health records of kidney disease patients according to the present invention.
[0075] Explanation of icon numbers:
[0076] 100. Data Acquisition and Preprocessing Unit; 101. Multi-Source Data Access Module; 102. Data Cleaning and Standardization Module; 103. Data Format Unification Module; 200. ASIC Dedicated Computing Unit; 201. ASIC Parallel Computing Module; 202. Homomorphic Encryption Hardware Module; 203. Long-Term Feature Extraction Module; 204. Pathological Attention Inference Module; 300. High-Speed Retrieval and Storage Unit; 301. Nephropathy Semantic Quantizer Module; 302. Inverted Index Construction Module; 303. Hybrid Retrieval Execution Module; 400. Security and Privacy Management Unit; 401. Hardware Encryption Engine Module; 402. Dynamic Access Control Module; 403. Full-Process Audit and Tracking Module. Detailed Implementation
[0077] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0078] like Figures 1-3As shown, this embodiment provides an intelligent retrieval system for kidney disease patient health records, including a data acquisition and preprocessing unit 100, an ASIC dedicated computing unit 200, a high-speed retrieval and storage unit 300, and a security and privacy management unit 400. Specifically: the data acquisition and preprocessing unit 100 is used to collect raw kidney disease patient record data from multi-source heterogeneous medical devices and databases, and preprocess the collected record data; the ASIC dedicated computing unit 200 is used to utilize an ASIC parallel computing architecture and native homomorphic encryption circuit to perform long-term feature extraction and pathological attention inference tasks for kidney disease in parallel under encrypted state, outputting encrypted high-dimensional feature vectors and corresponding plaintext pathological attention weights; the high-speed retrieval and storage unit 300 is used to compress the encrypted high-dimensional feature vectors into specific hash codes based on the plaintext pathological attention weights using a kidney disease semantic quantizer, and construct an inverted index based on the specific hash codes to perform extremely fast and accurate matching and fuzzy search of kidney disease feature fields; the security and privacy management unit 400 is used to implement compliant and secure management of the entire process of storage, transmission, and retrieval of kidney disease health records through a hardware encryption engine and dynamic permission control strategies.
[0079] It should be noted that the data acquisition and preprocessing unit 100 is responsible for cleaning and standardizing the original heterogeneous kidney disease records from multiple sources, and then sending them to the ASIC dedicated computing unit 200. The ASIC dedicated computing unit 200 uses native homomorphic encryption circuits to perform long-term feature extraction and pathological attention inference in parallel in the encrypted state, and outputs encrypted high-dimensional feature vectors and corresponding plaintext pathological attention weights. The high-speed retrieval and storage unit 300 receives the weights and vectors, compresses them through a kidney disease semantic quantizer to generate specific hash codes and builds an inverted index to achieve extremely fast and accurate matching and fuzzy search. The security and privacy management unit 400 ensures the compliance and security of data in all stages of storage, transmission and retrieval through a hardware encryption engine and dynamic permission control strategy.
[0080] In this embodiment, it should also be noted that the data acquisition and preprocessing unit 100 includes a multi-source data access module 101, a data cleaning and standardization module 102, and a data format unification module 103. Specifically: the multi-source data access module 101 is used to connect to multi-source heterogeneous terminals such as hospital information systems, testing equipment, and electronic medical records to collect original data from kidney disease patients; the data cleaning and standardization module 102 is used to clean, deduplicate, and complete missing, redundant, and abnormal kidney disease data; and the data format unification module 103 is used to uniformly convert data of different structures into a standard data format that the system can process.
[0081] It should be noted that the multi-source data access module 101 is responsible for collecting raw archive data from various heterogeneous terminals, and then handing it over to the data cleaning and standardization module 102 for cleaning, deduplication and completion to improve data quality. Finally, the data format unification module 103 converts the processed data into a standard format that the system can recognize.
[0082] Furthermore, it should be noted that the data collected in the multi-source data access module 101 includes basic information such as the name, age, and course of disease of kidney disease patients;
[0083] Clinical laboratory data including renal function, urinalysis, and electrolytes;
[0084] It also includes pathology reports, medication records, follow-up records, and imaging data from ultrasound and CT scans.
[0085] In the data cleaning and standardization module 102, missing data completion is performed as follows: For missing values of key renal disease indicators such as serum creatinine and blood urea nitrogen, data interpolation methods such as linear interpolation or polynomial interpolation between adjacent time points are used to complete the missing values. Missing values of non-critical fields are marked as "not recorded" to avoid data distortion by blindly completing the missing values. Redundant data deduplication is performed by using the patient ID and data collection time as unique identifiers to remove duplicate test results and duplicate diagnosis records. Abnormal data cleaning is performed by using the 3σ principle to identify abnormal values of renal function indicators that exceed the clinically reasonable range. Clinical experience is used to determine whether they are data entry errors. Erroneous data is retained after correction, and data that cannot be corrected is marked as abnormal and stored separately, and is not included in subsequent feature calculations.
[0086] In this embodiment, it should also be noted that the ASIC dedicated computing unit 200 includes an ASIC parallel computing module 201, a homomorphic encryption hardware module 202, a long-term feature extraction module 203, and a pathological attention inference module 204. Specifically: the ASIC parallel computing module 201 provides dedicated hardware parallel computing power to support high-concurrency medical data processing in encrypted mode; the homomorphic encryption hardware module 202 performs data encryption at the hardware level, supporting direct computation in encrypted state without decryption. The specific operations are as follows: A1: Hardware-level encryption is performed on the received kidney disease patient file data to generate and output encrypted data; A2: The homomorphic operation circuit is called to directly perform homomorphic addition and homomorphic multiplication operations on the received encrypted data without decryption, generating intermediate encrypted operation results; A3: The intermediate encrypted operation results are output to the long-term feature extraction module 203 and the pathological attention inference module 204 respectively. Long-term feature extraction module 203: This module is used to solidify and deploy a time-series analysis algorithm specifically for kidney disease, enabling parallel mining of long-term physiological change trends in patient history records within encrypted data. The specific operations are as follows: B1: Perform time-series segmentation processing on the received encrypted file data, dividing the patient history record data into segments according to the time dimension; B2: In encrypted state, use a homomorphic weighted accumulation circuit to process the data of each time segment in parallel, extracting the time-series change features of key information in patient physiological indicators and pathological diagnoses; The specific expression for extracting the time-series features of kidney disease physiological indicators in encrypted state is as follows:
[0087] ;
[0088] In the formula, This is the ciphertext time-series feature vector. Here, T is the homomorphic encryption function, and T is the total number of time windows. The weight of the t-th time window of the preset plaintext, B3: Using the homomorphic difference calculation unit, calculate the difference between encrypted data segments in adjacent time windows and multiply it by the plaintext trend smoothing coefficient to generate the encrypted long-term trend quantification value; where the specific expression for the long-term physiological trend quantification value of kidney disease patients is as follows:
[0089] ;
[0090] In the formula, This is a quantified value of the long-term trend in the encrypted state. This is the trend smoothing coefficient. B4 represents the changes in renal physiological indicators within adjacent time windows; B5: The quantified long-term trend value of the encrypted text is fused into the encrypted temporal feature vector through homomorphic addition and transmitted to the pathological attention inference module 204. The pathological attention inference module 204 is used to calculate the contribution of different renal indicators to the current pathological state based on the input temporal features using hardware acceleration, output the encrypted high-dimensional feature vector, and synchronously output the calculated attention weights as plaintext pathological attention weights after on-chip decryption. The specific operation is as follows: C1: Receives the encrypted temporal feature vector from the long-term feature extraction module 203, and maps it into an encrypted triplet structure of query vector, key vector, and value vector through hardware mapping logic; C2: Without performing decryption, calls the homomorphic dot product operation unit to calculate the encrypted similarity score between the encrypted query vector and the encrypted key vector, and performs nonlinear activation approximation processing on the score through a polynomial approximation circuit to simulate the normalization characteristics of the Softmax function, generating encrypted pathological attention weights, which are then converted into plaintext pathological attention weights through the on-chip decryption unit; The specific expression for generating encrypted pathological attention weights is as follows:
[0091] ;
[0092] In the formula, For encrypted pathology attention weights, The Softmax function approximation algorithm is performed for hardware, where Q is the ciphertext query vector and K is the ciphertext key vector. For vector dimensions, To query the ciphertext dot product result of the vector and the key vector; C3: Using a homomorphic weighted accumulation array, perform batch homomorphic multiplication and accumulation operations on the plaintext pathological attention weights and the ciphertext value vectors, aggregating them to generate a high-dimensional ciphertext feature vector; the specific expression for the aggregation of the high-dimensional ciphertext feature vector is as follows:
[0093] ;
[0094] In the formula, V is the generated high-dimensional feature vector of the ciphertext, and V is the ciphertext value vector. C4: Performs a homomorphic multiplication and accumulation operation on the encrypted pathology attention weights and value vectors; C5: Synchronously outputs the encrypted high-dimensional feature vector and the plaintext pathology attention weights.
[0095] It should be noted that the ASIC parallel computing module 201 provides underlying hardware computing power support. The homomorphic encryption hardware module 202 performs hardware-level encryption on the input data and performs homomorphic addition and multiplication operations. The intermediate encrypted results are then sent to the long-term feature extraction module 203 and the pathological attention inference module 204, respectively. The long-term feature extraction module 203 performs time-series segmentation, weighted accumulation, and differential trend calculation on the patient's historical records in encrypted state to generate encrypted time-series feature vectors and transmits them to the pathological attention inference module 204. The pathological attention inference module 204 generates encrypted attention weights based on this vector through homomorphic dot product and Softmax approximation. After on-chip decryption, the plaintext weights are obtained. The plaintext weights are then used to perform homomorphic weighted accumulation with the encrypted value vector to generate encrypted high-dimensional feature vectors. Finally, the encrypted high-dimensional feature vectors and plaintext pathological attention weights are output synchronously.
[0096] Furthermore, it should be noted that the hardware-level encryption in the homomorphic encryption hardware module 202 adopts a hardware implementation based on ring learning homomorphic encryption, with the encryption key length set to 2048 bits to ensure the security of the ciphertext data, while also adapting to the efficiency requirements of subsequent homomorphic addition and multiplication operations.
[0097] The hardware implementation of "homomorphic addition and homomorphic multiplication operations" in A2 is as follows: it calls the on-chip dedicated homomorphic operation unit, adopts a pipelined parallel architecture, and processes the addition and multiplication operations of multiple sets of encrypted data simultaneously. The operation latency is controlled within 100ns to meet the requirements of high concurrency processing.
[0098] In the long-term feature extraction module 203, the total number of time windows T is set according to the kidney disease follow-up period, typically 12 windows (one window per month per year). For patients with a disease course exceeding 5 years, this can be dynamically adjusted to 24-60 windows to ensure complete capture of long-term physiological change trends. Trend smoothing coefficient. The value ranges from 0.1 to 0.3 and can be dynamically adjusted according to the length of the patient's disease: the longer the disease duration, the higher the value. The larger the value, the weaker the impact of short-term fluctuations on long-term trends.
[0099] Vector dimension in Pathological Attention Reasoning Module 204 Based on the number of kidney disease indicators, the standard values are 64 or 128 to ensure the rationality of attention weight calculation.
[0100] The specific implementation of the "polynomial approximation circuit" in C2 is as follows: a 3rd-order polynomial is used to approximate the Softmax function, and the fitting error is controlled within 5%, ensuring the accuracy of the calculation of the ciphertext attention weights, while reducing the hardware computational complexity.
[0101] In this embodiment, it should also be noted that the high-speed retrieval and storage unit 300 includes a nephrology semantic quantizer module 301, an inverted index construction module 302, and a hybrid retrieval execution module 303. Specifically, the nephrology semantic quantizer module 301 is used to perform weighted hashing on the high-dimensional feature vector of the encrypted text based on the input encrypted pathological attention weights, compressing it into a compact, specific hash code that retains the semantic information of the nephrology. The specific operation is as follows: D1 receives the plaintext pathological attention weights from the pathological attention inference module 204, uses them as a dynamic gating signal, adjusts the preset plaintext initial projection matrix coefficients, and generates a plaintext dynamic projection matrix adapted to the current pathological state. The specific expression for adjusting the weighted projection matrix is as follows:
[0102] ;
[0103] In the formula, This is the weighted projection matrix. The preset initial projection matrix of plaintext. For Hadama accumulation, This is the weighting adjustment coefficient. Let be the plaintext pathology attention weight vector as input, and 1 be the input... A vector consisting entirely of 1s with the same dimension;
[0104] D2: Using the plaintext dynamic projection matrix, without decryption, a homomorphic linear transformation is performed on the high-dimensional feature vector of the ciphertext, mapping the pathological semantic similarity in the high-dimensional space to linearly separable features in the low-dimensional space, generating the low-dimensional feature vector of the ciphertext; the specific expression of the dimensionality reduction linear transformation is as follows:
[0105] ;
[0106] In the formula, The reduced-dimensional feature vectors are the result of dimensionality reduction. D3: Using a multinomial fitting activation function, the low-dimensional feature vector of the ciphertext is subjected to nonlinear activation and binarization truncation, compressing the continuous ciphertext values into a 0 / 1 bit sequence to generate a specific hash code. The specific expression for generating the binarized specific hash code is as follows:
[0107] ;
[0108] In the formula, H is the final generated specific hash code. It is a non-linear activation function. This is the binarization truncation threshold. A sign function maps positive values to 1 and negative values to 0. D4: Outputs the specific hash code to the inverted index construction module 302. Inverted index construction module 302: Based on the generated specific hash code, it builds an inverted index table pointing to specific patient records using the specific hash code. Hybrid retrieval execution module 303: Receives retrieval requests and supports both ultra-fast, precise matching based on specific hash codes and approximate fuzzy search based on feature vectors.
[0109] It should be noted that the nephrology semantic quantizer module 301 dynamically adjusts the projection matrix according to the input plaintext pathology attention weights, performs homomorphic linear transformation and binarization truncation on the encrypted high-dimensional feature vector, generates a specific hash code that preserves the nephrology semantics, and outputs it to the inverted index construction module 302. The inverted index construction module 302 builds an inverted index table pointing to the patient's file based on the specific hash code, providing index support for the hybrid retrieval execution module 303. After receiving the retrieval request, the hybrid retrieval execution module 303 simultaneously supports ultra-fast and precise matching based on the specific hash code and approximate fuzzy search based on the feature vector.
[0110] Furthermore, it should be noted that the weight adjustment coefficient in the nephropathy semantic quantizer module 301... The value range is 0.5-0.8. The larger the value, the stronger the moderating effect of the plaintext pathology attention weight on the projection matrix, and the more effectively it highlights the semantic information of core pathological indicators. Binarization truncation threshold. The value ranges from 0 to 0.2, and is dynamically adjusted based on the mean of the low-dimensional feature vector of the encrypted text to ensure that the specific hash code can effectively distinguish patient files with different pathological states.
[0111] The construction rule of the inverted index table in the inverted index construction module 302 is as follows: each bit of the specific hash code is used as the index item to establish a mapping relationship of "specific hash code bit combination - patient file ID". The index table adopts a block storage method, with each block storing 1000 specific hash code mapping relationships to improve the index query efficiency.
[0112] In D3, "Specific Implementation of Polynomial Fitting Activation Function": A quadratic polynomial is used to fit the ReLU function, with a fitting interval of [-1, 1]. The expression of the fitted function is as follows: (when x≥0) (When x < 0), this ensures the accuracy of nonlinear activation while reducing the complexity of hardware implementation.
[0113] In this embodiment, it should also be noted that the security and privacy management unit 400 includes a hardware encryption engine module 401, a dynamic access control module 402, and a full-process audit and tracing module 403. Specifically: the hardware encryption engine module 401 provides real-time encryption and decryption services to the system's archives and search results without consuming the main CPU; the dynamic access control module 402 dynamically allocates data access granularity based on the minimum necessary principle and the searcher's role, controlling the anonymization level and visibility range of search results; and the full-process audit and tracing module 403 immutably records all data access, retrieval, and access control operation logs.
[0114] It should be noted that the hardware encryption engine module 401 provides real-time hardware-level encryption and decryption services for the static archive and dynamic search results in the system, ensuring the confidentiality of data throughout the entire process. The dynamic access control module 402, based on the principle of minimum necessity, dynamically allocates data access granularity according to the searcher role and controls the anonymization level and visibility of search results. The full-process audit and tracking module 403 records all data access, search operations, and permission change behaviors in an immutable log.
[0115] Furthermore, it should be noted that the hardware encryption engine module 401 uses an independent PCIe interface hardware encryption card, supports the AES-256 encryption algorithm, has an encryption / decryption rate of ≥10GB / s, does not occupy the system's main CPU resources, and ensures the real-time performance of data storage and transmission.
[0116] The specific rules for "role-based dynamic permission allocation" in the dynamic permission control module 402 are as follows:
[0117] ① Doctor role: Can only search the files of patients treated by the doctor. The search results show complete pathological indicators and diagnostic records, but do not show the patient's ID number or contact information.
[0118] ② Researcher role: Can only search patient files with anonymized and hidden patient names and ID numbers. Search results only show pathological features and time-series trends, and do not show specific diagnostic records;
[0119] ③ Administrator role: Can view all patient files, has permission assignment and log query functions, but cannot modify the original patient data;
[0120] ④ Nurse role: Can only retrieve basic information and follow-up records of patients under their care, and cannot view detailed pathological diagnoses and laboratory indicators.
[0121] The logs in the full-process audit tracking module 403 are stored using blockchain to ensure immutability. The log content includes the operator ID, operation time, and the operation type and content of collection / retrieval / permission change, which meets the compliance requirements of the "Medical Data Security Guidelines" and the "Personal Information Protection Law". The log retention period is no less than 3 years.
[0122] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0123] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A renal patient health profile intelligent retrieval system, characterized by, It includes a data acquisition and preprocessing unit (100), an ASIC dedicated computing unit (200), a high-speed retrieval and storage unit (300), and a security and privacy management unit (400), wherein: The data acquisition and preprocessing unit (100) is used to acquire original kidney disease patient file data from multi-source heterogeneous medical devices and databases, and to preprocess the acquired file data. The ASIC dedicated computing unit (200) is used to utilize the ASIC parallel computing architecture and native homomorphic encryption circuit to perform long-term feature extraction and pathological attention reasoning tasks of kidney disease in parallel under ciphertext state, and output ciphertext high-dimensional feature vector and corresponding plaintext pathological attention weights. The high-speed retrieval and storage unit (300) is used to compress the encrypted high-dimensional feature vector into a specific hash code by using a nephropathy semantic quantizer according to the plaintext pathological attention weight, and to build an inverted index based on the specific hash code to perform extremely fast and accurate matching and fuzzy search on the nephropathy feature field. The security and privacy management unit (400) is used to conduct compliant and secure management of the entire process of storing, transmitting and retrieving kidney disease health records through a hardware encryption engine and dynamic permission control policies. The ASIC dedicated computing unit (200) includes an ASIC parallel computing module (201), a homomorphic encryption hardware module (202), a long-term feature extraction module (203), and a pathological attention reasoning module (204). The pathological attention reasoning module (204) calculates the contribution of different kidney disease indicators to the current pathological state based on the input temporal characteristics, outputs a high-dimensional feature vector of encrypted text, and outputs the calculated attention weights as plaintext pathological attention weights after on-chip decryption. The specific operation is as follows: C1: Receives the ciphertext temporal feature vector from the long temporal feature extraction module (203) and maps it into a ciphertext triplet structure of query vector, key vector and value vector through hardware mapping logic; C2: Without performing decryption, the homomorphic dot product operation unit is called to calculate the ciphertext similarity score between the ciphertext query vector and the ciphertext key vector. The score is then approximated by a polynomial approximation circuit with nonlinear activation to simulate the normalization characteristics of the Softmax function and generate ciphertext pathological attention weights. Subsequently, the on-chip decryption unit converts them into plaintext pathological attention weights. C3: Using a homomorphic weighted accumulation array, perform batch homomorphic multiplication and accumulation operations on the plaintext pathological attention weights and the ciphertext value vectors to aggregate and generate ciphertext high-dimensional feature vectors; C4: Synchronously outputs the encrypted high-dimensional feature vector and the plaintext pathological attention weights; The high-speed retrieval and storage unit (300) includes a nephrology semantic quantizer module (301), an inverted index construction module (302), and a hybrid retrieval execution module (303). The nephrology semantic quantizer module (301) performs weighted hashing on the high-dimensional feature vector based on the input encrypted pathological attention weights, compressing it into a compact, specific hash code that retains the semantic information of nephrology. The specific operation is as follows: D1: Receive the plaintext pathological attention weights from the pathological attention reasoning module (204), use them as dynamic gating signals, adjust the preset plaintext initial projection matrix coefficients, and generate a plaintext dynamic projection matrix that adapts to the current pathological state. D2: Using the plaintext dynamic projection matrix, without decryption, a homomorphic linear transformation is performed on the high-dimensional feature vector of the ciphertext to map the pathological semantic similarity in the high-dimensional space to linearly separable features in the low-dimensional space, generating the low-dimensional feature vector of the ciphertext. D3: Using a polynomial fitting activation function, nonlinear activation and binarization truncation are performed on the low-dimensional feature vector of the ciphertext to compress the continuous ciphertext values into a 0 / 1 bit sequence and generate a specific hash code. D4: Output the specific hash code to the inverted index building module (302).
2. The intelligent retrieval system for health profile of a patient with kidney disease according to claim 1, wherein, The data acquisition and preprocessing unit (100) includes a multi-source data access module (101), a data cleaning and standardization module (102), and a data format unification module (103), wherein: The multi-source data access module (101) is used to connect to multi-source heterogeneous terminals such as hospital information systems, testing equipment, and electronic medical records to collect original file data of kidney disease patients. The data cleaning and standardization module (102) is used to clean, deduplicatize, and complete missing, redundant, and abnormal kidney disease data. The data format unification module (103) is used to convert data with different structures into a standard data format that the system can process.
3. The intelligent retrieval system for health records of kidney disease patients according to claim 1, characterized in that, The ASIC parallel computing module (201) is used to provide dedicated hardware parallel computing power to support high-concurrency medical data processing in dense conditions. The homomorphic encryption hardware module (202) is used to encrypt data at the hardware level and supports direct computation in the ciphertext state without decryption. The long-term feature extraction module (203) is used to solidify and deploy a time-series analysis algorithm specifically for kidney disease, and to mine long-term physiological change trends in patient history records in parallel within encrypted text. The pathological attention reasoning module (204) is used to calculate the contribution of different kidney disease indicators to the current pathological state based on the input temporal characteristics, output the encrypted high-dimensional feature vector, and output the calculated attention weights as plaintext pathological attention weights after on-chip decryption.
4. The intelligent retrieval system for health records of kidney disease patients according to claim 3, characterized in that, The homomorphic encryption hardware module (202) performs data encryption at the hardware level, supporting direct computation in ciphertext without decryption. The specific operation is as follows: A1: Encrypt the received kidney disease patient file data at the hardware level, generate and output ciphertext data; A2: Call the homomorphic operation circuit to directly perform homomorphic addition and homomorphic multiplication operations on the received ciphertext data without decryption, and generate intermediate ciphertext operation results; A3: Output the intermediate ciphertext operation results to the long-term feature extraction module (203) and the pathological attention reasoning module (204), respectively.
5. The intelligent retrieval system for health records of kidney disease patients according to claim 3, characterized in that, The long-term feature extraction module (203) is equipped with a time-series analysis algorithm specifically for kidney disease. This algorithm mines long-term physiological change trends in patient history records in parallel within the encrypted data. The specific operation is as follows: B1: Perform time-series fragmentation on the received confidential document archive data, dividing the patient's historical archive data into segments according to the time dimension; B2: In encrypted state, homomorphic weighted accumulation circuits are used to process data of each time segment in parallel to extract the temporal variation characteristics of key information of patient physiological indicators and pathological diagnosis. The specific expression for extracting the temporal features of renal physiological indicators in encrypted state is as follows: ; In the formula, This is the ciphertext time-series feature vector. Here, T is the homomorphic encryption function, and T is the total number of time windows. The weight of the t-th time window of the preset plaintext, The encrypted state of renal physiological indicators within the t-th time window; B3: Using the homomorphic difference calculation unit, calculate the difference between ciphertext data segments in adjacent time windows and multiply it by the plaintext trend smoothing coefficient to generate a quantified value of the long-term trend of the ciphertext. The specific expression for the long-term physiological trend quantification value of kidney disease patients is as follows: ; In the formula, This is a quantified value of the long-term trend in the encrypted state. This is the trend smoothing coefficient. This represents the change in renal physiological indicators within adjacent time windows; B4: The long-term trend quantification value of the ciphertext is fused into the temporal feature vector of the ciphertext through homomorphic addition and transmitted to the pathological attention inference module (204).
6. The intelligent retrieval system for health records of kidney disease patients according to claim 3, characterized in that, The specific expression for generating the attention weight in encrypted pathology is as follows: ; In the formula, For encrypted pathology attention weights, The Softmax function approximation algorithm is performed for hardware, where Q is the ciphertext query vector and K is the ciphertext key vector. For vector dimensions, To query the ciphertext dot product of the vector and the key vector; The specific expression for ciphertext high-dimensional feature vector aggregation is as follows: ; In the formula, V is the generated high-dimensional feature vector of the ciphertext, and V is the ciphertext value vector. Add an accumulation operation to the homomorphic multiplication of the attention weights and value vectors in the encrypted pathology.
7. The intelligent retrieval system for health records of kidney disease patients according to claim 1, characterized in that, The high-speed retrieval and storage unit (300) includes a nephrology semantic quantizer module (301), an inverted index construction module (302), and a hybrid retrieval execution module (303), wherein: The nephrology semantic quantizer module (301) is used to perform weighted hashing on the high-dimensional feature vector of the ciphertext according to the input ciphertext pathological attention weight, and compress it into a compact, specific hash code that retains the semantic information of nephrology. The inverted index construction module (302) builds an inverted index table based on the generated specific hash codes, pointing to specific patient records; The hybrid retrieval execution module (303) is used to receive retrieval requests and supports both high-speed, precise matching based on specific hash codes and approximation-based fuzzy search based on feature vectors.
8. The intelligent retrieval system for health records of kidney disease patients according to claim 7, characterized in that, The specific expressions involved in steps D1-D3 are as follows: The specific expression for adjusting the weighted projection matrix in D1 is as follows: ; In the formula, This is the weighted projection matrix. The preset initial projection matrix of plaintext. For Hadama accumulation, This is the weighting adjustment coefficient. Let 1 be the plaintext pathology attention weight vector as input, and 1 be the input pathology attention weight vector. A vector consisting entirely of 1s with the same dimension; The specific expression for the dimension-reduction linear transformation in D2 is as follows: ; In the formula, The reduced-dimensional feature vectors are the result of dimensionality reduction. The input is a high-dimensional feature vector of the ciphertext; The specific expression for generating the binarized specific hash code in D3 is as follows: ; In the formula, H is the final generated specific hash code. It is a non-linear activation function. This is the binarization truncation threshold. The sign function maps positive values to 1 and negative values to 0.
9. The intelligent retrieval system for health records of kidney disease patients according to claim 1, characterized in that, The security and privacy management unit (400) includes a hardware encryption engine module (401), a dynamic access control module (402), and a full-process audit and tracing module (403), wherein: The hardware encryption engine module (401) is used to provide real-time encryption and decryption services for the archives and search results in the system without occupying the main CPU. The dynamic permission control module (402) dynamically allocates data access granularity according to the searcher role based on the principle of minimum necessity, and controls the anonymization level and visibility range of search results. The full-process audit tracking module (403) is used to record all data access, retrieval, and permission change operation logs in an immutable manner.
Citation Information
Patent Citations
Cross-industry health insurance data sharing and privacy protection method based on homomorphic encryption
CN121580425A
Encryption processing device, encryption processing method, and encryption processing program
WO2022044464A1