A processing method for prescription data, an electronic device, and a storage medium

CN122598935APending Publication Date: 2026-08-18HUNAN YIHUITONG INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610781828.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0002]现有技术在处方数据隐私保护与可用性之间普遍存在难以平衡的问题

Benefits of technology

本发明提供的一种用于处方数据的处理方法,包括:对电子处方数据进行语义实体抽取,得到处方敏感数据,并对处方敏感数据进行风险分级标记,得到敏感特征标签;基于敏感特征标签对处方敏感数据进行随机差分扰动计算,得到扰动数值序列;基于处方敏感数据对扰动数值序列进行特征交叉与非线性映射,得到混淆处方数据;基于预设的药品知识图谱对混淆处方数据进行节点比对,得到图谱匹配节点;基于图谱匹配节点对混淆处方数据进行风险评估与数值更新,得到目标处方数据,解决了如何在保护处方数据隐私的前提下,保证数据的可用性以支持风险评估与知识图谱关联为主要问题的技术问题,实现了在有效保护处方数据隐私的同时,保留其语义结构与关联特性,从而保障了脱敏后数据在风险评估和知识图谱应用中的可用性与准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122598935A_ABST
    Figure CN122598935A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of processing method for prescription data, electronic equipment and storage medium, it is related to data processing technical field, including the following steps, to electronic prescription data is carried out semantic entity extraction, obtain prescription sensitive data, and prescription sensitive data is carried out risk classification mark, obtain sensitive feature label;Prescription sensitive data is carried out random difference disturbance calculation based on sensitive feature label, obtain disturbance numerical sequence;Based on prescription sensitive data, disturbance numerical sequence is carried out feature cross and non-linear mapping, obtain confused prescription data;Confused prescription data is carried out node comparison based on pre-set drug knowledge graph, obtain graph matching node;Confused prescription data is carried out risk assessment and numerical update based on graph matching node, obtain target prescription data, it solves how to protect the technical problem of the premise of prescription data privacy, guarantee the availability of data to support risk assessment and knowledge graph association as main problem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method for processing prescription data, an electronic device, and a storage medium. Background Technology

[0002] Existing technologies generally struggle to balance the protection of prescription data privacy with its usability. While traditional de-identification methods (such as static anonymization, generalization, or encrypted storage) can reduce the risk of privacy breaches to some extent, they often come at the cost of data granularity, relevance, and semantic integrity. This makes it difficult for risk assessment models to capture individualized medication patterns and disease progression, while the entity links, relational reasoning, and cross-source fusion required for knowledge graphs also face serious information gaps.

[0003] Some studies have attempted to employ differential privacy or federated learning, but these methods struggle to simultaneously address the source analysis of complex clinical pathways and the mining of knowledge associations based on prescription sequences. Therefore, how to preserve sufficient semantic dimensions and structured association capabilities while ensuring prescription data privacy to support computable modeling of medication risks and the construction of dynamic knowledge graphs has become a critical issue that urgently needs to be addressed. Summary of the Invention

[0004] The purpose of this invention is to at least partially solve one of the technical problems existing in the prior art.

[0005] To achieve the above objectives, the present invention provides a method for processing prescription data, comprising: Semantic entity extraction is performed on electronic prescription data to obtain prescription sensitive data, and risk classification and labeling are performed on the prescription sensitive data to obtain sensitive feature labels; Random difference perturbation calculation is performed on prescription sensitive data based on sensitive feature labels to obtain a perturbation value sequence; Based on prescription-sensitive data, feature crossing and nonlinear mapping are performed on the perturbation numerical sequence to obtain confused prescription data; Based on a pre-defined drug knowledge graph, the nodes of the confusing prescription data are compared to obtain graph matching nodes; Risk assessment and numerical updates of confused prescription data are performed based on graph matching nodes to obtain target prescription data.

[0006] Furthermore, semantic entity extraction is performed on the electronic prescription data to obtain prescription-sensitive data, and risk-level labeling is applied to this data to obtain sensitive feature tags, including: The electronic prescription data is split into text sequences to obtain prescription word sequences, and the text entities of the prescription word sequences are extracted based on a preset medical privacy lexicon to obtain sensitive prescription data. The privacy exposure value is calculated based on a preset sensitive entity weight table for prescription sensitive data. Sensitive feature labels are obtained by mapping threshold intervals to prescription sensitive data based on privacy exposure values.

[0007] Furthermore, based on the sensitive feature labels, random difference perturbation calculations are performed on the prescription sensitive data to obtain a perturbation value sequence, including: Numerical analysis of prescription sensitive data is performed based on sensitive feature labels to obtain the prescription feature matrix; Based on the sensitive feature labels, the perturbation amplitude mapping of the prescription feature matrix is ​​performed to obtain the noise amplitude curve; A Laplace probability density function is constructed based on the noise amplitude curve, and the position of each element in the prescription feature matrix is ​​independently and randomly sampled based on the Laplace probability density function to obtain the noise sampling sequence. The range of the noise sampling sequence is calculated and matrix transformation is performed based on the prescription feature matrix to obtain the difference noise matrix; The prescription feature matrix is ​​element-wise summed based on the difference noise matrix to obtain a noisy feature matrix. The noisy feature matrix is ​​then dimension-reduced and spliced ​​to obtain a perturbation numerical sequence.

[0008] Furthermore, based on the prescription feature matrix, the range of the noise sampling sequence is calculated and matrix transformed to obtain the difference noise matrix, including: The column vector analysis of the prescription feature matrix is ​​performed to obtain the feature extreme value sequence, and the range span is calculated based on the feature extreme value sequence to obtain the feature range value; The sampled values ​​of each dimension in the noise sampling sequence are scaled element-wise based on the feature range values ​​to obtain the scaled noise sequence. The discrete noise sequence is then scaled by numerical product to obtain the scaled noise sequence. Based on the prescription feature matrix, row and column scale analysis is performed on the scaled noise sequence to obtain sequence segmentation nodes; The scaling noise sequence is reconstructed into a matrix shape based on the sequence segmentation nodes to obtain the difference noise matrix.

[0009] Furthermore, based on the prescription feature matrix, the range of the noise sampling sequence is calculated and matrix transformed to obtain the difference noise matrix, including: The maximum and minimum values ​​of each column of the prescription feature matrix are extracted by iterating through the data to obtain the extreme value pairs of each column. The maximum and minimum values ​​of each extreme value pair are then subtracted pairwise to obtain the feature range vector. The boundary noise sequence is obtained by performing element-wise product scaling on the sampled values ​​of each dimension in the noise sampling sequence based on the range values ​​of each dimension in the feature range vector. Based on the number of rows and columns of the prescription feature matrix, the boundary noise sequence is padded by rows and aligned by columns to obtain the difference noise matrix.

[0010] Furthermore, based on prescription-sensitive data, feature crossing and nonlinear mapping are performed on the perturbation numerical sequence to obtain confused prescription data, including: Entity association parsing is performed on prescription-sensitive data to obtain a prescription semantic topology graph, and dimension reshaping calculation is performed on the perturbation numerical sequence to obtain the perturbation feature vector; Based on the prescription semantic topology graph, the node attributes of the perturbation feature vector are concatenated to obtain the cross feature matrix, and the cross feature matrix is ​​expanded by polynomial to obtain the high-dimensional feature matrix. A nonlinear numerical transformation is performed on the high-dimensional feature matrix based on a preset activation function to obtain a nonlinear mapping matrix; Based on the preset prescription field template, the nonlinear mapping matrix is ​​trunculated and the field position is mapped to obtain the field-level confusion vector. The prescription field template contains the dimension index range of the drug name field, the dosage value field, and the usage frequency field. The confused prescription data is obtained by performing structured assignment and filling based on the field-level confusion vector and the prescription field template.

[0011] Furthermore, based on the prescription semantic topology graph, the node attributes of the perturbation feature vectors are concatenated to obtain a cross-feature matrix, including: The semantic topology graph of the prescription is parsed by traversing nodes to obtain a node index sequence. Based on the node index sequence, the original attribute vectors corresponding to each node are extracted to obtain the node attribute matrix. Based on the node index sequence, the perturbation feature vector is dimensionally segmented and aligned with nodes to obtain the perturbation attribute sub-vectors corresponding to each node; The cross feature matrix is ​​obtained by horizontally concatenating each row of the node attribute matrix with the corresponding perturbation attribute sub-vector.

[0012] Furthermore, a polynomial expansion is performed on the cross-feature matrix to obtain a high-dimensional feature matrix, including: The cross feature matrix is ​​split into column vectors to obtain prescription entity vectors. The prescription entity vectors are then paired and multiplied element by element to obtain cross-product data blocks. Based on the cross-feature matrix, the cross-product data blocks are concatenated by feature dimensions to obtain a multi-level feature data table. The multi-level feature data table is then reorganized into matrix rows and columns to obtain a high-dimensional feature matrix.

[0013] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the above methods.

[0014] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any of the above methods.

[0015] Technical effect This invention provides a method for processing prescription data, comprising: extracting semantic entities from electronic prescription data to obtain prescription sensitive data, and marking the prescription sensitive data with risk classification to obtain sensitive feature labels; performing random difference perturbation calculation on the prescription sensitive data based on the sensitive feature labels to obtain a perturbation numerical sequence; performing feature cross-mapping and nonlinear mapping on the perturbation numerical sequence based on the prescription sensitive data to obtain obfuscated prescription data; performing node comparison on the obfuscated prescription data based on a preset drug knowledge graph to obtain graph matching nodes; and performing risk assessment and numerical update on the obfuscated prescription data based on the graph matching nodes to obtain target prescription data. This method solves the technical problem of ensuring data availability to support risk assessment and knowledge graph association while protecting prescription data privacy. It achieves effective protection of prescription data privacy while preserving its semantic structure and association characteristics, thereby ensuring the availability and accuracy of the desensitized data in risk assessment and knowledge graph applications. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the steps of a method for processing prescription data in one embodiment of the present invention; Figure 2 This is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.

[0018] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. The step numbers in the following embodiments are set only for ease of explanation, and there is no limitation on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0020] The following describes the prescription data processing system in an embodiment of the present invention. Please refer to [link / reference]. Figure 1 An embodiment of the present invention provides a method for processing prescription data, comprising: Step S1: Extract semantic entities from electronic prescription data to obtain prescription sensitive data, and perform risk classification and labeling on the prescription sensitive data to obtain sensitive feature labels.

[0021] Specifically, semantic entity extraction is performed on electronic prescription data to obtain sensitive prescription data. This sensitive data is then risk-classified and labeled to obtain sensitive feature tags. This step first uses a pre-trained medical domain BERT model to perform sequence annotation on the original electronic prescription text, identifying structured or unstructured fields such as patient name, ID number, diagnosis code, generic drug name, and dosage, thereby extracting the sensitive prescription data. Subsequently, based on the Personal Information Protection Law and the Medical Data Classification and Grading Guidelines, the extracted sensitive prescription data is classified into high, medium, and low levels according to the potential harm caused by leakage. For example, a combination of ID number and specific disease diagnosis is marked as high risk, while only the drug name is marked as low risk. Each piece of sensitive prescription data is then assigned a corresponding sensitive feature tag. In this embodiment, the above scheme achieves automated and accurate identification of sensitive prescription information and compliant risk labeling, providing a structured input foundation for subsequent differentiated desensitization processing.

[0022] Step S2: Perform random difference perturbation calculation on the prescription sensitive data based on sensitive feature labels to obtain the perturbation value sequence.

[0023] Specifically, based on sensitive feature labels, random differential perturbation calculations are performed on prescription sensitive data to obtain a perturbation numerical sequence. This process first sets a corresponding privacy budget ε according to the risk level identified by the sensitive feature labels, allocating smaller ε values ​​(e.g., ε=0.5) to high-risk data and larger ε values ​​(e.g., ε=2.0) to low-risk data. Then, a Laplace mechanism is introduced for numerical prescription sensitive data (e.g., medication dosage, frequency), generating noise according to the formula Lap(Δf / ε). Lap is an abbreviation for Laplace distribution, commonly used in differential privacy mechanisms, where Δf is the sensitivity of the field; for example, if the daily dosage range of an antihypertensive drug is 5–10 mg, then Δf is 5. For categorical data (e.g., drug names), an exponential mechanism is used to perturb the legitimate drug set according to probability. In this embodiment, the above scheme achieves refined perturbation matching the risk level while strictly satisfying differential privacy theory, effectively preventing the original prescription sensitive data from being accurately restored.

[0024] In the above formula Lap(Δf / ε), Lap(Δf / ε) represents sampling a noise value from a Laplace distribution with a mean of 0 and a scale parameter of Δf / ε, and adding this noise to the original data to achieve privacy protection.

[0025] For example: If Δf=5 and ε=1, then the noise follows the Laplace(0,5) law, and its probability density function is: Add this noise to the original dose (e.g., 7 mg) to obtain the perturbed value.

[0026] Step S3: Based on the prescription-sensitive data, perform feature crossing and nonlinear mapping on the perturbation numerical sequence to obtain confused prescription data.

[0027] Specifically, based on prescription-sensitive data, perturbed numerical sequences are subjected to feature cross-mapping and nonlinear mapping to obtain obfuscated prescription data. First, the perturbed numerical fields (such as dosage and frequency) are combined with categorical fields (such as drug name and diagnosis code) from the original prescription-sensitive data using a Cartesian product to construct high-dimensional cross-features. Then, these cross-features are input into a pre-trained multilayer perceptron (MLP) with LeakyReLU activation and a hidden layer dimension of 128. Nonlinear transformations break the linear correlation between the original data; for example, "Amlodipine 5mg qd" and "Hypertension I10" are cross-mapped into irreversible embedding vectors. In this embodiment, the above scheme effectively blocks attackers from reconstructing associations in desensitized data using statistical correlation or background knowledge, significantly improving the anti-inference capability of obfuscated prescription data.

[0028] Step S4: Based on the preset drug knowledge graph, perform node comparison on the confused prescription data to obtain graph matching nodes.

[0029] Specifically, based on a pre-defined drug knowledge graph, the confusing prescription data is compared node by node to obtain matching nodes. This process first standardizes the drug name field in the confusing prescription data and uses it as a query entity, inputting it into the constructed drug knowledge graph. This graph uses the generic drug name as a core node, associated with attributes such as indications, contraindications, and interactions. Then, a joint metric strategy of edit distance and semantic embedding is used to retrieve the nearest neighbor match in the graph node set. For example, the perturbed "Amlodipine Tablets (fuzzy code X73)" is mapped back to the standard node "Amlodipine". If the similarity exceeds a threshold of 0.85, it is confirmed as a matching node. In this embodiment, the above scheme effectively corrects the semantic shift introduced by the perturbation, ensuring the consistency and usability of the desensitized prescription data at the clinical logic level.

[0030] Step S5: Based on the graph matching nodes, perform risk assessment and numerical update on the confused prescription data to obtain the target prescription data.

[0031] Specifically, risk assessment and numerical updates are performed on the obfuscated prescription data based on graph matching nodes to obtain target prescription data. Specifically, the clinical risk attributes (such as contraindications and drug interaction levels) associated with the graph matching nodes are mapped back to the corresponding prescription records. If a "Amlodipine" node is matched and its knowledge graph indicates a severe interaction with "Ketoconazole," a high-risk marker is triggered. Subsequently, the dosage or frequency fields in the original obfuscated prescription are numerically corrected according to preset rules, for example, adjusting the daily dose from the perturbed 8.3 mg to the knowledge graph-recommended safe upper limit of 5 mg. The update process retains the privacy noise introduced by the perturbation, making only minor adjustments within the clinical safety boundaries. In this embodiment, the above scheme significantly reduces the potential medical risks of desensitized prescriptions in real-world medication scenarios while maintaining differential privacy protection.

[0032] In a specific embodiment, semantic entity extraction is performed on electronic prescription data to obtain prescription-sensitive data, and risk-level labeling is applied to the prescription-sensitive data to obtain sensitive feature tags, including: The electronic prescription data is split into text sequences to obtain prescription word sequences, and the text entities of the prescription word sequences are extracted based on a preset medical privacy lexicon to obtain sensitive prescription data. The privacy exposure value is calculated based on a preset sensitive entity weight table for prescription sensitive data. Sensitive feature labels are obtained by mapping threshold intervals to prescription sensitive data based on privacy exposure values.

[0033] Specifically, electronic prescription data is split according to natural delimiters in the text, including Chinese punctuation marks such as commas, periods, and colons, as well as English symbols such as semicolons, semicolons, and line breaks, thus forming a prescription vocabulary sequence composed of independent words or phrases. For example, an original prescription content "Patient: Li Moumou, Diagnosis: Type 2 Diabetes (E11.9), Medication: Metformin tablets 0.5g tid" will be segmented into multiple units such as "Patient", ":", "Li Moumou", ",", "Diagnosis", ":", "Type 2 Diabetes", "(", "E11.9", ")", ",", "Medication", ":", "Metformin tablets", "0.5g", and "tid". This segmentation method preserves the word order and local contextual structure of the original text, providing a foundation for subsequent accurate identification of sensitive information.

[0034] The prescription's word sequence is compared item by item using a pre-defined medical privacy thesaurus, and successfully matching words are extracted as sensitive prescription data. This thesaurus is constructed in accordance with regulations and standards such as the Personal Information Protection Law and the Electronic Medical Record Application Management Standards, and includes patient identification identifiers (such as name, fragments of ID number), disease codes (e.g., A64 in ICD-10 represents unspecified sexually transmitted diseases, F20 represents schizophrenia), high-risk drug names (e.g., "clozapine" and "tenofovir"), and specific diagnostic descriptions (e.g., "HIV infection" and "malignant tumor"). If any item in the prescription's word sequence completely matches an entry in the thesaurus, or conforms to a predefined regular expression pattern (e.g., "\d{6}****\d{4}" is used to match partially anonymized ID numbers), it is extracted. For example, in the above prescription, "Li Moumou" as the patient's name, "E11.9" as the diabetes code, and "metformin tablets" as a drug strongly associated with chronic diseases, could all be identified and included in the sensitive prescription data set.

[0035] After obtaining sensitive prescription data, the system assigns a corresponding numerical weight to each sensitive data item based on a pre-defined sensitive entity weight table to quantify its privacy leakage risk. This weight table, jointly developed by clinical experts and data compliance personnel, reflects the sensitivity of different entities in real-world medical scenarios. For example, the weight for "diagnosed with HIV (A64)" is set to 0.92, "schizophrenia (F20)" to 0.88, "type 2 diabetes (E11.9)" to 0.35, and "common cold" to only 0.08. For multiple sensitive entities extracted from a prescription, the overall privacy exposure value is obtained by adding up the weights of each item. If the sum exceeds 1.0, it is truncated to the upper limit of 1.0 to ensure consistency in the numerical range. For example, if a prescription contains "E11.9" (weight 0.35) and "metformin tablets" (weight 0.30), the privacy exposure value is 0.65.

[0036] Based on the calculated privacy exposure value, it is mapped to a preset threshold range to generate corresponding sensitive feature labels. Specifically, when the privacy exposure value is less than 0.3, it is labeled "low sensitivity"; when the value is between 0.3 and 0.7 (inclusive), it is labeled "medium sensitivity"; and when the value is greater than or equal to 0.7, it is labeled "high sensitivity". This label directly determines the strength of subsequent data processing strategies, such as whether generalization, perturbation, or complete masking is required.

[0037] In this embodiment, the above-mentioned scheme achieves automatic identification and risk classification of sensitive information in electronic prescriptions through structured lexicon matching, configurable weight calculation and explicit threshold mapping mechanism, providing reliable technical support for differentiated privacy protection of medical data in the process of sharing, analysis and publication.

[0038] In a specific embodiment, random difference perturbation calculation is performed on the prescription sensitive data based on sensitive feature labels to obtain a perturbation value sequence, including: Numerical analysis of prescription sensitive data is performed based on sensitive feature labels to obtain the prescription feature matrix; Based on the sensitive feature labels, the perturbation amplitude mapping of the prescription feature matrix is ​​performed to obtain the noise amplitude curve; A Laplace probability density function is constructed based on the noise amplitude curve, and the position of each element in the prescription feature matrix is ​​independently and randomly sampled based on the Laplace probability density function to obtain the noise sampling sequence. The range of the noise sampling sequence is calculated and matrix transformation is performed based on the prescription feature matrix to obtain the difference noise matrix; The prescription feature matrix is ​​element-wise summed based on the difference noise matrix to obtain a noisy feature matrix. The noisy feature matrix is ​​then dimension-reduced and spliced ​​to obtain a perturbation numerical sequence.

[0039] Specifically, based on sensitive feature labels, random difference perturbation calculations are performed on prescription sensitive data to generate a perturbation numerical sequence. The specific implementation process is as follows: First, numerical parsing is performed on the prescription sensitive data according to the determined sensitive feature labels (such as "low sensitivity," "medium sensitivity," or "high sensitivity"). This parsing operation converts the original textual sensitive information (e.g., "E11.9," "Metformin Tablets," or "Li Moumou") into fixed-dimensional real-number vectors through a pre-trained medical embedding model. Multiple vectors are stacked vertically according to their order of appearance in the prescription, forming a two-dimensional structure, namely the prescription feature matrix. The number of rows in this matrix is ​​equal to the number of sensitive entities, and the number of columns is determined by the embedding dimension, typically 128 or 256 dimensions.

[0040] Based on the sensitive feature label corresponding to the prescription, the corresponding perturbation intensity parameter is retrieved from the preset perturbation strategy configuration table, and this parameter is mapped to a noise amplitude curve with the same number of rows as the prescription feature matrix. For example, if the label is "high sensitivity", each row corresponds to a larger amplitude value (e.g., 1.2); if it is "medium sensitivity", a medium value is used (e.g., 0.6); and "low sensitivity" corresponds to a smaller value (e.g., 0.2). This curve is used to subsequently control the noise scale at different row positions, ensuring that the higher the risk level, the stronger the introduced perturbation.

[0041] Using each element in the noise amplitude curve as a scaling parameter, a corresponding Laplace probability density function is constructed. For each element position in the prescription feature matrix—i.e., the i-th row and j-th column—a random noise value is independently extracted from the Laplace distribution corresponding to its row. All extraction results are arranged according to their original matrix positions, forming a noise sampling sequence identical in shape to the prescription feature matrix. This process ensures that the perturbations of each element are independent, meeting the basic requirements of differential privacy.

[0042] The numerical distribution characteristics of the prescription feature matrix itself are used to adjust the noise sampling sequence. Specifically, the prescription feature matrix is ​​first analyzed column by column, and the difference between the maximum and minimum values ​​of each column is calculated to obtain the feature range values ​​of each column. Then, all elements of the corresponding column in the noise sampling sequence are uniformly multiplied by the range value of that column to achieve adaptive scaling, forming a scaled noise sequence. Afterwards, based on the row and column structure of the prescription feature matrix, the one-dimensional scaled noise sequence is re-segmented according to the original row width and reorganized into a two-dimensional form, thus obtaining the difference noise matrix.

[0043] The difference noise matrix is ​​added element-wise to the original prescription feature matrix to generate a noisy feature matrix. This matrix retains the original semantic structure but has been injected with random perturbations that meet privacy requirements. To further adapt to downstream processing needs, the noisy feature matrix is ​​dimensionality-reduced and concatenated: for example, the mean of each row vector is calculated or principal component projection is used to compress it into a single value, and then the results of each row are sequentially concatenated into a one-dimensional array, with the final output being a perturbed numerical sequence.

[0044] In this embodiment, the above scheme combines perturbation amplitude control driven by sensitive feature labels, adaptive noise scaling based on feature range, and matrix reconstruction mechanism that preserves structure, thereby effectively maintaining the availability of prescription data in statistical analysis or machine learning tasks while ensuring formal privacy security.

[0045] In a specific embodiment, the range calculation and matrix transformation of the noise sampling sequence are performed based on the prescription feature matrix to obtain the difference noise matrix, including: The column vector analysis of the prescription feature matrix is ​​performed to obtain the feature extreme value sequence, and the range span is calculated based on the feature extreme value sequence to obtain the feature range value; The sampled values ​​of each dimension in the noise sampling sequence are scaled element-wise based on the feature range values ​​to obtain the scaled noise sequence. The discrete noise sequence is then scaled by numerical product to obtain the scaled noise sequence. Based on the prescription feature matrix, row and column scale analysis is performed on the scaled noise sequence to obtain sequence segmentation nodes; The scaling noise sequence is reconstructed into a matrix shape based on the sequence segmentation nodes to obtain the difference noise matrix.

[0046] Specifically, the prescription feature matrix is ​​first analyzed using column vector parsing. Assume that after a certain electronic prescription is identified by sensitive feature tags, four sensitive entities are extracted (such as "Type 2 Diabetes", "Metformin Tablets", "Li Moumou", "2025-03-15"). These are then mapped to 4-dimensional real-number vectors using a medical embedding model (for simplicity, the embedding dimension is assumed to equal the number of entities, forming a 4×4 matrix). The resulting prescription feature matrix is ​​a 4x4 two-dimensional array. This matrix is ​​then parsed along the column direction, splitting it into four column vectors, each containing four values. For example, the first column might be [0.82, 0.75, 0.91, 0.68], the second column [-0.34, -0.29, -0.41, -0.30], and so on. For each column, its maximum and minimum values ​​are recorded, forming a feature extreme value sequence containing eight values—the first four are the maximum values ​​of each column, and the last four are the minimum values.

[0047] The range is calculated based on this feature extreme value sequence. Taking the first column as an example, its maximum value is 0.91, and its minimum value is 0.68, with a difference of 0.23, which is the feature range value of this column. The second column has a maximum value of -0.29 and a minimum value of -0.41, with a range of 0.12. This process continues, resulting in four feature range values, forming a feature range value sequence: [0.23, 0.12, 0.18, 0.20]. These values ​​reflect the original fluctuation range of the prescription feature matrix across each embedding dimension.

[0048] The noise sampling sequence is scaled element-wise using the aforementioned feature range values. The noise sampling sequence, independently generated based on a Laplace distribution in the previous step, has a structure identical to the prescription feature matrix, also 4×4. For example, its first column might be [0.15, -0.08, 0.22, -0.10]. Then, all elements in this column are multiplied by the corresponding feature range value of the first column, 0.23, to obtain the scaled first column: [0.0345, -0.0184, 0.0506, -0.0230]. Similarly, the noise values ​​in the second column are multiplied by 0.12, the third by 0.18, and the fourth by 0.20. Through this operation, the original noise is adaptively adjusted according to the actual data span in each dimension, forming a scaled noise sequence. The phrase "numerical product scaling of discrete noise sequences to obtain scaled noise sequences" mentioned here essentially refers to transforming the noise sampling sequence, which was originally a set of discrete numerical values, into a scaled version through the above column-by-column product operation. It is still called a scaled noise sequence, emphasizing the consistency between its source and transformation logic.

[0049] The scaling noise sequence is analyzed using row and column scaling based on the prescription feature matrix. Since the prescription feature matrix is ​​4×4, with 4 rows and 4 columns, and a total of 16 elements, the system determines the segmentation rule for the scaling noise sequence (currently a one-dimensional arrangement of 16 values): every 4 elements form a row, therefore the sequence segmentation nodes are located at positions 4, 8, and 12, for a total of 3 nodes. These nodes identify the boundary positions of each row of the original matrix within the sequence.

[0050] Based on these sequence segmentation nodes, the scaled noise sequence is reconstructed into a matrix structure. Specifically, the first four elements are taken as row 1, the fifth to eighth as row 2, the ninth to twelfth as row 3, and the thirteenth to sixteenth as row 4, reorganizing into a 4×4 two-dimensional structure. This structure is the differential noise matrix, where each element retains the randomness of the Laplace noise while incorporating the range information of the original prescription features in the corresponding dimension, achieving a perturbation intensity that matches the data distribution.

[0051] In a specific embodiment, the range calculation and matrix transformation of the noise sampling sequence are performed based on the prescription feature matrix to obtain the difference noise matrix, including: The maximum and minimum values ​​of each column of the prescription feature matrix are extracted by iterating through the data to obtain the extreme value pairs of each column. The maximum and minimum values ​​of each extreme value pair are then subtracted pairwise to obtain the feature range vector. The boundary noise sequence is obtained by performing element-wise product scaling on the sampled values ​​of each dimension in the noise sampling sequence based on the range values ​​of each dimension in the feature range vector. Based on the number of rows and columns of the prescription feature matrix, the boundary noise sequence is padded by rows and aligned by columns to obtain the difference noise matrix.

[0052] Specifically, the maximum and minimum values ​​of each column in the prescription feature matrix are extracted by iterating through the data. Assume that an electronic prescription, after sensitive entity recognition and embedding processing, forms a 5x5 prescription feature matrix (e.g., containing 5 sensitive fields: diagnosis code, drug name, dosage, patient ID, and consultation date; each field is mapped to a 5-dimensional vector, forming a square matrix to simplify processing). Each column of this matrix is ​​independently iterated through its 5 elements, recording the maximum and minimum values ​​respectively. For example, if the values ​​in column 1 are [0.73, 0.81, 0.69, 0.85, 0.77], then its maximum value is 0.85 and its minimum value is 0.69; if the values ​​in column 2 are [-0.42, -0.38, -0.45, -0.40, -0.39], then its maximum value is -0.38 and its minimum value is -0.45. After performing this operation on all 5 columns, 5 pairs of extreme values ​​are obtained: (0.85, 0.69), (-0.38, -0.45), (0.92, 0.88), (-0.21, -0.25), and (0.60, 0.55). These pairs of extreme values ​​fully characterize the numerical boundaries of the prescription features in each embedding dimension.

[0053] The feature range vector is obtained by subtracting the maximum and minimum values ​​of each pair of extreme values ​​in each column. Taking column 1 as an example, 0.85 minus 0.69 equals 0.16; in column 2, -0.38 minus -0.45 equals 0.07; in column 3, 0.92 - 0.88 = 0.04; in column 4, -0.21 - (-0.25) = 0.04; and in column 5, 0.60 - 0.55 = 0.05. The resulting feature range vector is [0.16, 0.07, 0.04, 0.04, 0.05]. The length of this vector is equal to the number of columns in the prescription feature matrix. Each dimension corresponds to the original fluctuation range of an embedding dimension; the larger the value, the more significant the difference between different prescriptions in that dimension.

[0054] The noise sampling sequence is scaled element-wise by multiplying the sampled values ​​in each dimension based on the range values ​​in each dimension of the feature range vector. The noise sampling sequence is a set of random numbers previously generated independently from the Laplace distribution, and its structure is consistent with the prescription feature matrix, also 5×5. For example, its first column might be [0.12, -0.09, 0.15, -0.07, 0.10]. Then, all five sampled values ​​in this column are multiplied by 0.16, the first dimension of the feature range vector, to obtain the scaled first column: [0.0192, -0.0144, 0.0240, -0.0112, 0.0160]. Similarly, the noise values ​​in the second column are multiplied by 0.07, the third column by 0.04, and so on. After this column-wise scaling operation, the original uniform-scale noise is given a perturbation intensity that matches the dynamic range of the prescription features, forming a boundary noise sequence. The term "boundary" here comes from the fact that the range reflects the data boundary span of each dimension. Therefore, the scaled noise is called the boundary noise sequence, emphasizing that its disturbance amplitude is constrained by the original data boundary.

[0055] The boundary noise sequence is padded and aligned by row and column based on the number of rows and columns of the prescription feature matrix. Since the prescription feature matrix is ​​5 rows and 5 columns, with a total of 25 elements, although the boundary noise sequence is arranged in one dimension, its internal order is consistent with the column-first or row-first storage method of the original matrix (usually expanded by row). Based on the 5 rows and 5 columns, the system divides the boundary noise sequence into groups of 5 elements each, which are then grouped into rows 1 through 5 of the matrix. For example, the first 5 scaled values ​​form row 1, the next 5 form row 2, and so on, until all 5 rows are filled. During this process, column alignment ensures that each column still corresponds to the same embedding dimension of the original prescription feature matrix after reconstruction, maintaining dimensional semantic consistency. The final 5×5 two-dimensional array is the difference noise matrix.

[0056] In a specific embodiment, based on prescription-sensitive data, feature crossing and nonlinear mapping are performed on the perturbation numerical sequence to obtain confused prescription data, including: Entity association parsing is performed on prescription-sensitive data to obtain a prescription semantic topology graph, and dimension reshaping calculation is performed on the perturbation numerical sequence to obtain the perturbation feature vector; Based on the prescription semantic topology graph, the node attributes of the perturbation feature vector are concatenated to obtain the cross feature matrix, and the cross feature matrix is ​​expanded by polynomial to obtain the high-dimensional feature matrix. A nonlinear numerical transformation is performed on the high-dimensional feature matrix based on a preset activation function to obtain a nonlinear mapping matrix; Based on the preset prescription field template, the nonlinear mapping matrix is ​​trunculated and the field position is mapped to obtain the field-level confusion vector. The prescription field template contains the dimension index range of the drug name field, the dosage value field, and the usage frequency field. The confused prescription data is obtained by performing structured assignment and filling based on the field-level confusion vector and the prescription field template.

[0057] Specifically, entity association parsing is performed on prescription-sensitive data to obtain a prescription semantic topology graph. For example, a real electronic prescription contains fields such as "Drug Name: Enteric-coated Aspirin Tablets," "Dosage: 100mg," "Frequency of Use: Once Daily," "Patient Name: Zhang Moumou," and "Diagnosis Code: I25.10." The system identifies these sensitive entities using Named Entity Recognition (NER) and relation extraction models, and further analyzes their semantic dependencies—for example, "Enteric-coated Aspirin Tablets" are commonly used to treat "I25.10 (atherosclerotic heart disease)," and its standard dose is usually "100mg," with usage typically "once daily." These entities and their relationships are modeled as a graph structure, namely the prescription semantic topology graph, where nodes represent entities (such as drugs, dosages, and diagnoses), and edges represent semantic or clinical logical relationships. This graph not only preserves the structural information of the original prescription but also encodes medical knowledge constraints.

[0058] Simultaneously, the perturbation numerical sequence undergoes dimensionality reshaping calculation to obtain a perturbation feature vector. Assume the perturbation numerical sequence output from the preceding differential privacy-enhancing step is a one-dimensional array containing 30 floating-point numbers (e.g., from 3 fields × 10-dimensional embeddings). Based on the preset embedding dimensions (e.g., 10 dimensions per field), the system reshapes this sequence into a one-dimensional perturbation feature vector of length 30, internally arranged by field order: the first 10 dimensions correspond to the drug name, the middle 10 dimensions to the dosage value, and the last 10 dimensions to the application frequency. Although this vector is in numerical form, it implicitly contains the semantic location of the corresponding prescription field.

[0059] Based on the prescription semantic topology graph, the system concatenates node attributes of perturbed feature vectors to obtain a cross-feature matrix. Specifically, the system uses the sub-vectors corresponding to each entity in the perturbed feature vector as node attributes and appends them to the corresponding nodes in the prescription semantic topology graph. For example, the drug node is appended with the first 10 dimensions, the dosage node with the middle 10 dimensions, and the usage node with the last 10 dimensions. Subsequently, the attribute vectors of all nodes are horizontally concatenated according to the node order in the graph to form a cross-feature matrix with the number of rows equal to the number of nodes (e.g., 3) and the number of columns equal to the total dimensions (e.g., 30). This matrix not only contains the perturbed numerical features but also integrates the semantic relationship structure between entities, laying the foundation for subsequent high-dimensional interactions.

[0060] A polynomial expansion is performed on the cross-feature matrix to obtain a high-dimensional feature matrix. For example, a second-order interactive expansion is performed on each row of the matrix (representing a combination of prescription fields): in addition to retaining the original 30 dimensions, all pairwise dimension product terms are added (such as 1st dimension × 2nd dimension, 1st dimension × 3rd dimension, ... up to 29th dimension × 30th dimension), for a total of 435 new terms (C(30,2) = 435), increasing the total dimension to 465. This operation explicitly models the nonlinear coupling relationship between different perturbation dimensions, such as the joint effect of drug embedding in one dimension and dose embedding in another dimension, thereby increasing the complexity and irreversibility of confusion.

[0061] A nonlinear numerical transformation is performed on the high-dimensional feature matrix based on a preset activation function to obtain a nonlinear mapping matrix. Here, activation functions such as ReLU or tanh are used to independently apply a nonlinear transformation to each element in the high-dimensional feature matrix. For example, if a product term has a result of -0.8, after tanh activation it becomes approximately -0.664; if it is 1.2, it becomes approximately 0.834. This step breaks linear separability, making it difficult to reconstruct the original perturbation value through reverse engineering, significantly improving privacy and security.

[0062] Based on a pre-defined prescription field template, the nonlinear mapping matrix is ​​truncated in dimensions and its field positions are mapped to obtain a field-level obfuscation vector. The prescription field template explicitly defines the dimensional index range of each key field in the vector; for example, the drug name field corresponds to dimensions 1-10, the dosage value field corresponds to dimensions 11-20, and the usage frequency field corresponds to dimensions 21-30. Although the high-dimensional feature matrix reaches 465 dimensions, the system only backtracks to extract the nonlinear mapping results corresponding to the original 30 dimensions (i.e., skipping higher-order interaction terms and retaining the base dimensions that correspond one-to-one with the original fields), forming a field-level obfuscation vector of length 30. This ensures that subsequent assignments can still correspond back to the standard prescription structure.

[0063] The system generates obfuscated prescription data by using a structured assignment and filling method based on a field-level obfuscation vector and a prescription field template. The field-level obfuscation vector is divided into three segments according to the template: the first 10 dimensions are input to a drug name decoder (e.g., a legitimate drug name embedded in the space through nearest neighbor matching), outputting "ibuprofen sustained-release capsules"; the middle 10 dimensions are input to a dosage value decoder, outputting "200mg"; and the last 10 dimensions are input to a usage frequency decoder, outputting "twice daily". These decoding results are filled into the structured template according to the field order of the original prescription, ultimately generating a complete obfuscated prescription data: "Drug Name: Ibuprofen Sustained-Release Capsules, Dosage Value: 200mg, Usage Frequency: Twice Daily". While semantically reasonable and formatted correctly, this data has no direct traceable connection to the original prescription.

[0064] In a specific embodiment, the perturbation feature vector is concatenated with node attributes based on the prescription semantic topology graph to obtain a cross-feature matrix, including: The semantic topology graph of the prescription is parsed by traversing nodes to obtain a node index sequence. Based on the node index sequence, the original attribute vectors corresponding to each node are extracted to obtain the node attribute matrix. Based on the node index sequence, the perturbation feature vector is dimensionally segmented and aligned with nodes to obtain the perturbation attribute sub-vectors corresponding to each node; The cross feature matrix is ​​obtained by horizontally concatenating each row of the node attribute matrix with the corresponding perturbation attribute sub-vector.

[0065] Specifically, in the actual processing flow of electronic prescription privacy protection, to achieve "concatenating node attributes of perturbation feature vectors based on the prescription semantic topology graph to obtain a cross-feature matrix," the specific operation must strictly rely on the precise alignment mechanism between the graph structure and the numerical vectors. The entire process begins with the parsing of the prescription semantic topology graph. This graph is generated by the preceding entity recognition and relation extraction steps, where each node represents a sensitive field in the prescription, such as "aspirin enteric-coated tablets" (drug name), "100mg" (dosage value), or "once daily" (usage frequency). The system performs a deterministic traversal of the graph—usually using a breadth-first strategy, and visits nodes according to the order in which the fields appear in the original prescription, thereby outputting a node index sequence, such as [drug name, dosage value, usage frequency], with corresponding indices [0, 1, 2].

[0066] With this index sequence, the next step is to extract the original attribute vectors corresponding to each node from the original embedding space. Assuming each field, after being encoded by the medical language model, outputs a fixed 10-dimensional real vector, then the drug name node corresponds to a 10-dimensional vector v0, the dosage value to v1, and the usage frequency to v2. These three vectors are arranged vertically in index order, forming a 3x10 node attribute matrix. Simultaneously, the perturbation feature vector—a 30-dimensional array generated by the preorder differential privacy module—needs to be decomposed and mapped one-to-one with the nodes in the graph. Since each field is known to occupy 10 dimensions, the system divides the perturbation feature vector into segments of 10 elements each, based on the field order implicit in the node index sequence: dimensions 0 to 9 are assigned to index 0 (drug name), dimensions 10 to 19 to index 1 (dosage value), and dimensions 20 to 29 to index 2 (usage frequency), thus obtaining three perturbation attribute sub-vectors, each 10 dimensions.

[0067] The crucial concatenation operation unfolds here: the i-th row of the node attribute matrix (i.e., the original semantic vector of the i-th node) is horizontally concatenated with the i-th perturbation attribute sub-vector along the dimensional direction. For example, if the original vector of the drug name is [0.73, 0.81, ..., 0.69], its corresponding perturbation sub-vector is [0.019, -0.014, ..., 0.016], concatenating them forms a new 20-dimensional row; the dosage and usage nodes are processed similarly. The final output cross-feature matrix is ​​3 rows and 20 columns, with each row integrating the original semantic information of a prescription field with its specific perturbation noise, preserving clinical logical constraints while injecting the randomness required for privacy protection.

[0068] In this embodiment, the above scheme establishes a strict mapping relationship between the prescription semantic topology graph and the perturbation feature vector through the node index sequence, ensuring that the original attributes and perturbation attributes are accurately aligned at the node granularity, and avoiding semantic distortion caused by field misalignment; at the same time, the horizontal splicing operation explicitly constructs a joint representation of the original features and perturbation features, providing structured input for subsequent high-order cross-computation, effectively supporting the generation quality and privacy strength of obfuscated prescription data.

[0069] In a specific embodiment, a high-dimensional feature matrix is ​​obtained by performing a polynomial expansion on the cross feature matrix, including: The cross feature matrix is ​​split into column vectors to obtain prescription entity vectors. The prescription entity vectors are then paired and multiplied element by element to obtain cross-product data blocks. Based on the cross-feature matrix, the cross-product data blocks are concatenated by feature dimensions to obtain a multi-level feature data table. The multi-level feature data table is then reorganized into matrix rows and columns to obtain a high-dimensional feature matrix.

[0070] Specifically, in the electronic prescription privacy protection process, multinomial expansion of the cross-feature matrix to generate a high-dimensional feature matrix is ​​a key step in achieving nonlinear obfuscation. This process strictly follows the steps outlined in the technical solution. First, the cross-feature matrix needs to be split into column vectors. Assuming the cross-feature matrix is ​​generated by the preceding node attribute concatenation step, its structure is 3 rows and 20 columns. The three rows correspond to the three prescription entity nodes: "Drug Name," "Dosage Value," and "Usage Frequency." Each 20-dimensional row is horizontally concatenated from the original attribute vector (10 dimensions) and the perturbation attribute sub-vector (10 dimensions). The system then cuts this matrix column by column, obtaining 20 column vectors of length 3. These column vectors are the prescription entity vectors, and the three elements of each vector represent the values ​​of a certain embedding dimension in the three prescription fields.

[0071] The 20 prescription entity vectors are paired up, and element-wise multiplication is performed on each pair. For example, the prescription entity vector in column 5 is [0.72, 0.85, 0.63], and the prescription entity vector in column 12 is [-0.11, 0.09, 0.14]. After pairing them up, the corresponding positions are multiplied to obtain [0.72×(-0.11), 0.85×0.09, 0.63×0.14] = [-0.0792, 0.0765, 0.0882], forming a 3D cross-product data block. Since the 20 vectors are paired up (without repetition or self-multiplication), a total of 190 such pairs are generated, thus generating 190 cross-product data blocks, each of which is a 3D column vector.

[0072] Based on the original cross-feature matrix, these cross-product data blocks are concatenated according to their feature dimensions. Specifically, the original 20 columns are retained as first-order features, and then 190 cross-product data blocks are sequentially concatenated horizontally to its right, forming a 3-row, 210-column multi-order feature data table. This table contains both the original linear features and all second-order interaction terms, fully expressing the coupling relationships within and across prescription fields. For example, in the prescription scenario of "Aspirin Enteric-coated Tablets - 100mg - Once Daily," a certain interaction term might reflect the joint perturbation effect between "drug embedding in the 3rd dimension" and "dosage embedding in the 7th dimension," a higher-order association that cannot be captured by a linear model.

[0073] The multi-level feature data table is then reorganized into a matrix. This reorganization does not alter the data content, but rather arranges it into a standard matrix form according to the input requirements of the subsequent nonlinear mapping module. In this embodiment, the 3-row, 210-column structure is maintained, and the output is a high-dimensional feature matrix. Each row still strictly corresponds to a prescription entity node, ensuring semantic alignment; each column represents an original or cross-feature dimension, expanding the total number of dimensions from 20 to 210, significantly improving the complexity of feature representation.

[0074] In this embodiment, the above scheme explicitly constructs a second-order interactive representation between prescription-sensitive fields by performing column splitting, pairwise multiplication, feature concatenation, and row and column reorganization on the cross-feature matrix, thereby deeply coupling the perturbation information with the original semantics. At the same time, the high-dimensional feature matrix retains the consistency of the node-level structure, providing a high-dimensional input with clinical logic constraints for subsequent nonlinear activation, effectively enhancing the irreversibility and business availability of obfuscated prescription data.

[0075] In a specific embodiment, the confusing prescription data is compared node by node based on a preset drug knowledge graph to obtain graph matching nodes, including: The node association of the pre-defined drug knowledge graph is parsed to obtain the graph node embedding matrix; The confused prescription data is projected into a vector space to obtain the confused projection vector; Based on the preset spatial alignment mapping matrix, the confused projection vector is linearly transformed to map the confused projection vector to the semantic vector space where the graph node embedding matrix is ​​located, thus obtaining the alignment feature vector. Based on the graph node embedding matrix, cosine similarity is calculated on the aligned feature vectors to obtain the node similarity matrix; Extreme value index extraction is performed based on the node similarity matrix to obtain the graph matching nodes.

[0076] Specifically, the system performs node association parsing on a pre-defined drug knowledge graph. This graph is constructed from an authoritative pharmaceutical database and contains tens of thousands of entity nodes related to drugs, indications, contraindications, and dosages. Nodes are connected by relational edges such as "treatment," "contraindication," and "inclusion." In this embodiment, only nodes directly related to the prescription field are considered, such as "aspirin enteric-coated tablets," "100mg," and "once daily." Each node is mapped to a fixed-dimensional real-valued vector using graph neural networks or embedding methods such as TransE. Assuming an embedding dimension of 128, when there are 50,000 related nodes in the graph, a 50,000-row, 128-column graph node embedding matrix is ​​obtained after parsing, where each row corresponds to the semantic representation of a graph node.

[0077] The system performs vector space projection on the confused prescription data. This confused prescription data is the output of the preceding high-dimensional feature matrix after nonlinear activation and dimensionality reduction, still retaining the structure of three prescription entities (drug name, dosage value, and usage frequency). The system independently projects these three entities to an intermediate vector space, for example, mapping them to a 128-dimensional vector through a fully connected layer, thus obtaining a 3×128 confused projection vector matrix. For simplicity, taking the confused entity corresponding to "Aspirin Enteric-coated Tablets" as an example, its projection result might be [0.63, -0.21, 0.87, …, 0.14], a total of 128 values.

[0078] The system performs a linear transformation on the obfuscated projection vector based on a pre-defined spatial alignment mapping matrix. This mapping matrix, with a size of 128×128, is pre-learned during model training through contrastive learning or adversarial alignment and is used to bridge the distributional differences between the obfuscated data space and the embedding space of the drug knowledge graph. The system left-multiplies each row of the obfuscated projection vector (i.e., the 128-dimensional vector of each prescription entity) by this mapping matrix, outputting a new 128-dimensional vector, the alignment feature vector. For example, after transformation, the numerical distribution of the 0th row of the original obfuscated projection vector more closely resembles the embedding style of real drug nodes in the graph, thus achieving cross-space semantic alignment.

[0079] Based on this, the system calculates cosine similarity between aligned feature vectors and the graph node embedding matrix. Specifically, it calculates the cosine similarity between each aligned feature vector (such as the 128-dimensional vector corresponding to the drug name) and all 50,000 rows of the graph node embedding matrix, resulting in a similarity score sequence of length 50,000. This process is performed on the three prescription entities, ultimately forming a 3×50,000 node similarity matrix. For example, the aligned vector of "Aspirin Enteric-coated Tablets" might have a similarity score of 0.92 with the node "Aspirin Enteric-coated Tablets (100mg)" in the graph, while it only has a similarity score of 0.31 with "Ibuprofen Extended-Release Capsules," demonstrating a significant semantic distinction.

[0080] Extreme value index extraction is performed based on the node similarity matrix. The system independently searches for the column index corresponding to the maximum similarity value in each row (i.e., each prescription entity). This index points to a specific row in the graph node embedding matrix, which is the graph matching node. For example, if the maximum value in row 0 appears in column 12,345, the graph matching node is "Aspirin Enteric-coated Tablets"; the maximum value in row 1 is in column 8,762, corresponding to "100mg"; and the maximum value in row 2 is in column 3,210, corresponding to "Once Daily". These three matching nodes together constitute the semantic verification result for confusing prescriptions.

[0081] In a specific embodiment, risk assessment and numerical updates are performed on the confused prescription data based on the graph matching nodes to obtain the target prescription data, including: Based on the graph matching nodes, the corresponding medication guidelines fields are extracted, and the drug dosage fields in the confused prescription data are compared within a range based on the drug guidelines fields to obtain dosage out-of-bounds markers; Based on dose out-of-bounds markers, the drug dosage field in the obfuscated prescription data is replaced with a safe value to obtain the calibrated dosage field; For non-dosage fields in the obfuscated prescription data that did not trigger out-of-bounds flags, retain their obfuscated field values ​​to obtain a set of retained fields; The prescription structure template based on the graph matching node aligns and splices the field order of the calibration dose field and the reserved field set to obtain the target prescription data.

[0082] Specifically, the system extracts corresponding medication guidelines based on the matching nodes in the drug knowledge graph. These matching nodes are three specific nodes obtained in the previous steps through comparison with the drug knowledge graph, such as "Aspirin Enteric-coated Tablets," "100mg," and "Once Daily." "Aspirin Enteric-coated Tablets" serves as the main drug node, associated with clearly defined medication guidelines in the pre-defined drug knowledge graph, including structured constraints such as the adult single-dose range (e.g., 50mg–300mg) and the maximum daily dose (e.g., 300mg / day). The system extracts these guidelines from the node's attribute slots to form the dose boundary conditions used for verification. Assume the extracted reasonable single-dose lower limit is 50mg and the upper limit is 300mg, in mg.

[0083] Based on these medication guidelines, the drug dosage fields in the obfuscated prescription data are compared within a range. Obfuscated prescription data is an intermediate result generated after high-dimensional perturbation and nonlinear transformation, and its dosage field may deviate from the clinically reasonable range due to random noise injection. For example, the original dosage is "100mg", but after perturbation it becomes "420mg". The system compares this value with the aforementioned standardized range [50, 300] and finds that 420 > 300, so it is judged as out of bounds, and a dosage out-of-bounds flag is generated with a value of "true"; if the perturbed dosage is "80mg", it falls within the reasonable range and is marked as "false".

[0084] The system replaces drug dosage fields in confused prescription data with safe values ​​based on dosage out-of-bounds flags. When the flag is "true", the system no longer retains the original perturbation value, but instead selects a safe replacement value from the medication guidelines field. In this embodiment, an upper limit truncation strategy is used, that is, out-of-bounds values ​​are replaced with the upper limit value of 300; mean replacement (such as 175mg) or nearest neighbor compliant values ​​can also be used, the specific strategy is determined by the configuration. For example, the original confused dose "420mg" is replaced with "300mg" to form the calibration dose field. If it does not exceed the limit (such as "80mg"), the original value is directly retained as the calibration dose field.

[0085] For non-dosage fields in the obfuscated prescription data that have not triggered out-of-bounds flags, retain their obfuscated values. In this scenario, non-dosage fields include "drug name" and "usage frequency". Since these two fields do not involve numerical safety boundaries and their semantics have been verified as reasonable in the preceding graph matching (e.g., matching the real "aspirin enteric-coated tablets" and "once daily"), they are retained regardless of whether they have been perturbed, as long as they are not judged as abnormal. For example, although the obfuscated drug name vector contains noise, its decoding result is still "aspirin enteric-coated tablets" and the usage frequency is "once daily", which together constitute the set of retained fields.

[0086] The prescription structure template based on graph matching nodes aligns and concatenates the calibration dose field and the reserved field set. Each drug node in the drug knowledge graph is associated with a standard prescription structure template, which specifies the output order of fields, such as [drug name, dosage value, frequency of use]. Following this template, the system places "Aspirin Enteric-coated Tablets" in the reserved field set at position 0, "Once Daily" in position 2, and then inserts the calibration dose field "300mg" in position 1, completing the alignment. Subsequently, the three are concatenated in sequence to form structurally complete, semantically compliant, and dosage-safe target prescription data, for example: "Aspirin Enteric-coated Tablets 300mg Once Daily".

[0087] In this embodiment, the above-mentioned scheme effectively intercepts the risk of dosage exceeding the limit due to privacy disturbance through the map-driven medication standard comparison mechanism; at the same time, while retaining the obfuscation effect of non-dosage fields, it only performs precise calibration on high-risk values ​​and achieves field-level alignment based on map structure templates, ensuring that the output target prescription data not only meets the strength of privacy protection, but also has clinical feasibility and regulatory compliance.

[0088] Reference Figure 2 This invention also provides a computer device whose internal structure can be as follows: Figure 2 As shown, the computer device includes a processor, memory, display screen, input device, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores the data corresponding to this embodiment. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the above-described method.

[0089] Those skilled in the art will understand that Figure 2 The structures shown are merely block diagrams of some structures related to the present invention and do not constitute a limitation on the computer devices on which the present invention is applied.

[0090] An embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0091] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the present invention and embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.

[0092] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0093] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for processing prescription data, characterized in that, include: Semantic entity extraction is performed on electronic prescription data to obtain prescription sensitive data, and risk classification and labeling are performed on the prescription sensitive data to obtain sensitive feature labels; Random difference perturbation calculation is performed on prescription sensitive data based on sensitive feature labels to obtain a perturbation value sequence; Based on prescription-sensitive data, feature crossing and nonlinear mapping are performed on the perturbation numerical sequence to obtain confused prescription data; Based on a pre-defined drug knowledge graph, the nodes of the confusing prescription data are compared to obtain graph matching nodes; Risk assessment and numerical updates of confused prescription data are performed based on graph matching nodes to obtain target prescription data.

2. The method for processing prescription data according to claim 1, characterized in that, Semantic entity extraction is performed on electronic prescription data to obtain prescription-sensitive data. This data is then risk-classified and labeled to obtain sensitive feature tags, including: The electronic prescription data is split into text sequences to obtain prescription word sequences, and the text entities of the prescription word sequences are extracted based on a preset medical privacy lexicon to obtain sensitive prescription data. The privacy exposure value is calculated based on a preset sensitive entity weight table for prescription sensitive data. Sensitive feature labels are obtained by mapping threshold intervals to prescription sensitive data based on privacy exposure values.

3. The method for processing prescription data according to claim 1, characterized in that, Random difference perturbation calculation is performed on prescription sensitive data based on sensitive feature labels to obtain a perturbation value sequence, including: Numerical analysis of prescription sensitive data is performed based on sensitive feature labels to obtain the prescription feature matrix; Based on the sensitive feature labels, the perturbation amplitude mapping of the prescription feature matrix is ​​performed to obtain the noise amplitude curve; A Laplace probability density function is constructed based on the noise amplitude curve, and the position of each element in the prescription feature matrix is ​​independently and randomly sampled based on the Laplace probability density function to obtain the noise sampling sequence. The range of the noise sampling sequence is calculated and matrix transformation is performed based on the prescription feature matrix to obtain the difference noise matrix; The prescription feature matrix is ​​element-wise summed based on the difference noise matrix to obtain a noisy feature matrix. The noisy feature matrix is ​​then dimension-reduced and spliced ​​to obtain a perturbation numerical sequence.

4. The method for processing prescription data according to claim 3, characterized in that, Based on the prescription feature matrix, the range of the noise sampling sequence is calculated and matrix transformed to obtain the difference noise matrix, including: The column vector analysis of the prescription feature matrix is ​​performed to obtain the feature extreme value sequence, and the range span is calculated based on the feature extreme value sequence to obtain the feature range value; The sampled values ​​of each dimension in the noise sampling sequence are scaled element-wise based on the feature range values ​​to obtain the scaled noise sequence. The discrete noise sequence is then scaled by numerical product to obtain the scaled noise sequence. Based on the prescription feature matrix, row and column scale analysis is performed on the scaled noise sequence to obtain sequence segmentation nodes; The scaling noise sequence is reconstructed into a matrix shape based on the sequence segmentation nodes to obtain the difference noise matrix.

5. A method for processing prescription data according to claim 3, characterized in that, Based on the prescription feature matrix, the range of the noise sampling sequence is calculated and matrix transformed to obtain the difference noise matrix, including: The maximum and minimum values ​​of each column of the prescription feature matrix are extracted by iterating through the data to obtain the extreme value pairs of each column. The maximum and minimum values ​​of each extreme value pair are then subtracted pairwise to obtain the feature range vector. The boundary noise sequence is obtained by performing element-wise product scaling on the sampled values ​​of each dimension in the noise sampling sequence based on the range values ​​of each dimension in the feature range vector. Based on the number of rows and columns of the prescription feature matrix, the boundary noise sequence is padded by rows and aligned by columns to obtain the difference noise matrix.

6. The method for processing prescription data according to claim 1, characterized in that, Based on prescription-sensitive data, feature crossing and nonlinear mapping are performed on the perturbed numerical sequences to obtain confused prescription data, including: Entity association parsing is performed on prescription-sensitive data to obtain a prescription semantic topology graph, and dimension reshaping calculation is performed on the perturbation numerical sequence to obtain the perturbation feature vector; Based on the prescription semantic topology graph, the node attributes of the perturbation feature vector are concatenated to obtain the cross feature matrix, and the cross feature matrix is ​​expanded by polynomial to obtain the high-dimensional feature matrix. A nonlinear numerical transformation is performed on the high-dimensional feature matrix based on a preset activation function to obtain a nonlinear mapping matrix; Based on the preset prescription field template, the nonlinear mapping matrix is ​​trunculated and the field position is mapped to obtain the field-level confusion vector. The prescription field template contains the dimension index range of the drug name field, the dosage value field, and the usage frequency field. The confused prescription data is obtained by performing structured assignment and filling based on the field-level confusion vector and the prescription field template.

7. A method for processing prescription data according to claim 6, characterized in that, Based on the prescription semantic topology graph, the node attributes of the perturbation feature vector are concatenated to obtain the cross-feature matrix, including: The semantic topology graph of the prescription is parsed by traversing nodes to obtain a node index sequence. Based on the node index sequence, the original attribute vectors corresponding to each node are extracted to obtain the node attribute matrix. Based on the node index sequence, the perturbation feature vector is dimensionally segmented and aligned with nodes to obtain the perturbation attribute sub-vectors corresponding to each node; The cross feature matrix is ​​obtained by horizontally concatenating each row of the node attribute matrix with the corresponding perturbation attribute sub-vector.

8. A method for processing prescription data according to claim 6, characterized in that, The cross-feature matrix is ​​expanded using a polynomial to obtain a high-dimensional feature matrix, including: The cross feature matrix is ​​split into column vectors to obtain prescription entity vectors. The prescription entity vectors are then paired and multiplied element by element to obtain cross-product data blocks. Based on the cross-feature matrix, the cross-product data blocks are concatenated by feature dimensions to obtain a multi-level feature data table. The multi-level feature data table is then reorganized into matrix rows and columns to obtain a high-dimensional feature matrix.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, The processor executes the steps of any one of claims 1 to 8 when executing a computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When a computer program is executed by a processor, it implements the steps of any one of claims 1 to 8.