Data verification method for clinical research data
By establishing potential feature space and introducing a noise injection mechanism in multi-center clinical trials, the privacy leakage problem caused by data differences is solved, and the stability of data analysis and privacy protection is balanced.
Patent Information
- Application Number
- CN202510153502.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-30
AI Technical Summary
In multicenter clinical trials, due to data differences among participants, the leakage of intermediate computing status may lead to the amplification of local data characteristics, thereby revealing the patient's privacy information.
By collecting data from each clinical research center, a standardized matrix is established and sparse non-negative matrix decomposition is carried out to establish a potential feature space. Then, quantification indicators of the distribution differences of potential characteristics between different research centers are calculated, the noise injection mechanism is activated, and the noise intensity control parameters are adjusted through orthogonal transformation and Laplace noise injection to ensure that the noise only affects the non-critical information part of the data.
Effectively identify and narrow the privacy risk points in the data, ensure the stability and accuracy of data analysis results, and protect patient privacy information without affecting the effectiveness of statistical results.
Smart Images

Figure CN120072349A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-party data security verification. More specifically, the present invention relates to a data verification method for clinical research data. Background Art
[0002] In multi-center clinical trials, multiple participants usually cooperate to calculate statistical indicators, such as verifying the treatment effect or drug safety of the sample population by analyzing the combined p-value of the efficacy difference. Due to the data differences between participants, especially the differences in population characteristics, the leakage of intermediate calculation states (such as gradient values, covariance matrices, etc.) may occur during the analysis of multi-party data sets. For example, there are significant differences in genomic characteristics between Asian and European populations. This data difference may cause non-uniformity during the intermediate calculation stage, resulting in the amplification of local data characteristics. If not prevented, these amplified local characteristics may be exploited by attackers, and through multiple iterative calculations and associations with intermediate results of different tasks, the sensitive data boundaries of a single center can be inversely inferred, thereby leaking the privacy information of patients.
[0003] Therefore, how to propose a data verification method for clinical research data that, on the premise of ensuring the statistical validity of the analysis results is not affected, verifies the differences in key data characteristics between participants through a special noise injection mechanism, reduces the differences in characteristics, and further protects the privacy information of patient groups in different regions, ethnic groups, or other classification methods is an urgently needed problem to be solved.
[0004] To solve the above problems, a technical solution is provided as follows. Summary of the Invention
[0005] To overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a data verification method for clinical research data to solve the problems raised in the above background art.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] S1: Collect the clinical research data of each clinical research center and establish a standardized matrix, perform sparse non-negative matrix factorization on the standardized matrix, and establish a latent feature space;
[0008] S2: Quantify the distribution difference index of the corresponding data of the latent features between different research centers, and perform a primary verification on the feature distribution differences between different research centers. If the distribution difference index between any two research centers is greater than or equal to the set distribution difference threshold, activate the noise injection mechanism;
[0009] S3: Perform a linear regression analysis on the potential features of each research center, iteratively extract the local gradient matrix and the global gradient tensor matrix of each research center, construct the within-class covariance matrix and between-class covariance matrix of the gradient distribution, and obtain the separability measure of the gradient distribution among research centers through matrix trace operation;
[0010] S4: Based on the deviation rate of the distribution difference quantification index and the separability measure of the gradient distribution, comprehensively calculate the noise intensity control parameter;
[0011] S5: Perform an orthogonal transformation on the local gradient matrix, construct an orthogonal projection matrix, generate noise through a Laplace distribution, and project the generated noise onto the direction adjusted by the orthogonal projection matrix to obtain the final noise direction;
[0012] S6: Inject the Laplace noise generated based on the noise intensity control parameter into the corresponding research center potential feature matrix, perform a secondary verification on the differences between the potential feature matrices of each research center, and adjust the noise intensity control parameter according to the distribution difference quantification index deviation of the secondary verification at the set scale step size;
[0013] S7: Re-perform a linear regression analysis on the potential features of each research center and calculate the gradient covariance matrix after noise enhancement, adjust the noise intensity control parameter based on the maximum element offset rate of the set covariance matrix, and turn off the noise injection mechanism after the adjustment is completed.
[0014] In a preferred embodiment, in S1, collecting the clinical research data of each clinical research center and establishing a standardized matrix, and performing sparse non-negative matrix factorization on the standardized matrix to establish a potential feature space specifically includes:
[0015] Collect the clinical research data of each clinical research center, uniformly process the format of the clinical research data of each research center, and establish a standardized matrix X i ∈R ni×d , where i is the research center number, n i is the number of samples of the i-th research center, and d is the number of features of each clinical research sample;
[0016] Perform feature selection through the L1 regularization method, and set the expression of the optimization objective of the minimum reconstruction error of matrix factorization for the standardized matrix X i to perform sparse non-negative matrix factorization, and the specific expression of the optimization objective is:
[0017]
[0018] In the formula, F i is the potential feature matrix, Q is the relationship matrix between the data in the potential feature matrix and the original data features, and λ is the preset regularization strength hyperparameter;
[0019] Set the constraint condition F i ≥0, Q≥0. Under the condition of meeting the constraint conditions, retain the potential feature matrix with a variance contribution degree greater than the set contribution degree value in the potential feature matrix, integrate the retained potential feature matrix, and establish a potential feature space;
[0020] Extract the mean and variance of the data corresponding to the retained potential features in the potential feature matrix based on high-dimensional statistical analysis.
[0021] In a preferred embodiment, in S2, by calculating the distribution difference quantization index of the data corresponding to the potential features between different research centers, conduct a primary verification of the feature distribution differences between different research centers. If there exists a distribution difference quantization index between any two research centers that is greater than or equal to the set distribution difference threshold, activating the noise injection mechanism specifically includes:
[0022] Obtain the mean and variance of the data corresponding to the potential features of each research center in the potential feature space, and calculate the weighted harmonic mean of the JS divergence and KL divergence of the data corresponding to the potential features between different research centers in parallel as the distribution difference quantization index;
[0023] Set a dynamic distribution difference threshold according to the research sample size of each research center, conduct a primary verification of the feature distribution differences between different research centers, compare the distribution difference quantization indexes of different research centers in the potential feature space pairwise. If there exists a distribution difference quantization index between any two research centers that is greater than or equal to the distribution difference threshold, activate the noise injection mechanism.
[0024] In a preferred embodiment, in S3, conduct a linear regression analysis on the potential features of each research center, iteratively extract the local gradient matrix and global gradient tensor matrix of each research center, construct the within-class covariance matrix and between-class covariance matrix of the gradient distribution, and obtain the separability measure of the gradient distribution between research centers through matrix trace operation specifically including:
[0025] Conduct a linear regression analysis on the potential features of each research center in the potential feature space, set the mean squared error as the loss function, and based on the SMPC protocol, iteratively extract the local gradient matrix G of each research center (t) , where t represents the corresponding iteration number;
[0026] According to the local gradient matrix extracted in each iteration, perform global mean processing on the local gradient matrix of each research center, calculate the global gradient tensor matrix of this iteration, and the calculation result is represented by g (t) ;
[0027] Calculate the local gradient matrix G in each iteration (t)The covariance matrix with the global gradient tensor matrix g (t) is marked as the within-class covariance matrix and denoted by ∑W;
[0028] Obtain the mean of the elements in all global gradient tensor matrices during the iteration process, and calculate the historical mean gradient matrix The covariance matrix with g (t) is marked as the between-class covariance matrix and denoted by ∑B;
[0029] Based on the within-class covariance matrix and the between-class covariance matrix, construct a separability measure for studying the gradient distribution between centers through matrix trace operation. The specific expression is:
[0030]
[0031] In the formula, β is the linear separability coefficient, and tr is the matrix trace operation identifier.
[0032] In a preferred embodiment, in S4, based on the deviation rate of the distribution difference quantization index and the separability measure of the gradient distribution, comprehensively calculate the noise intensity control parameter, which specifically includes:
[0033] Based on the deviation rate of the distribution difference quantization index between any two research centers and the distribution difference threshold, combined with the calculated linear separability coefficient, comprehensively calculate the noise intensity control parameter. The specific calculation expression is:
[0034]
[0035] In the formula, γ(p,q) is the noise intensity control parameter for controlling the generated noise for the potential feature matrix corresponding to research center p relative to q, D(p,q) is the distribution difference deviation rate between research centers p and q, ω is the sample-dimension ratio adjustment factor, n p 、n q are the sample numbers of research centers p and q respectively, c is the number of features in the potential feature space, and ∈ is the set smoothing parameter.
[0036] In a preferred embodiment, in S5, perform an orthogonal transformation on the local gradient matrix, construct an orthogonal projection matrix, generate noise through the Laplace distribution, and project the generated noise onto the direction adjusted by the orthogonal projection matrix to obtain the final noise direction.
[0037] In a preferred embodiment, in S6, inject the Laplace noise generated based on the noise intensity control parameter into the corresponding research center potential feature matrix, perform a secondary verification on the differences between the potential feature matrices of each research center, and adjust the noise intensity control parameter according to the distribution difference quantization index deviation of the secondary verification according to the set scale step length. Specifically include:
[0038] Inject Laplace noise generated based on the noise intensity control parameter into the matrix elements of the corresponding potential feature matrix of the research center;
[0039] Recalculate the distribution difference quantification index of each research center based on the potential feature matrix after noise injection, and conduct a secondary verification on the differences between the potential feature matrices of each research center;
[0040] Compare the distribution difference quantification indexes of different research centers after noise injection pairwise. If there is any distribution difference quantification index between any two research centers greater than or equal to the distribution difference threshold, increase the noise intensity control parameter by the set scale step;
[0041] If the distribution difference quantification indexes between all research centers are less than the distribution difference threshold, end the process of increasing the noise intensity control parameter.
[0042] In a preferred embodiment, in S7, re - conduct a linear regression analysis on the potential features of each research center and calculate the gradient covariance matrix after noise enhancement. Adjust the noise intensity control parameter based on the set maximum element deviation rate of the covariance matrix. After the adjustment is completed, turn off the noise injection mechanism, which specifically includes:
[0043] Re - conduct a linear regression analysis on the potential features of each research center and iteratively extract the local gradient matrix and the corresponding global gradient tensor matrix of each research center, calculate the covariance matrix between the local gradient matrix and the global gradient tensor matrix during the iterative process, and mark it as the gradient covariance matrix after noise enhancement;
[0044] Set the maximum element deviation rate of the covariance matrix. Compare the elements of the within - class covariance matrix without noise injection with the gradient covariance matrix after noise enhancement. If there are elements with an element deviation rate greater than the maximum element deviation rate, proportionally decrease the noise intensity control parameter according to the ratio exceeding the maximum element deviation rate, and regenerate the noise injection;
[0045] Repeatedly verify the deviation rate of the elements in the gradient covariance matrix after noise enhancement until the deviation rate of all elements is less than the maximum element deviation rate, then end the adjustment of the noise intensity control parameter and turn off the noise injection mechanism simultaneously.
[0046] The technical effects and advantages of a data verification method for clinical research data of the present invention:
[0047] By standardizing and sparsifying non - negative matrix factorization of multi - center clinical research data, a latent feature space is established. When evaluating the distribution differences of data from different research centers, quantitative indicators are used to verify the differences of latent features, effectively identifying privacy risk points in the data. By introducing linear regression analysis based on the gradient matrix, within - class and between - class covariance matrices are constructed, and the separability measure of the gradient distribution is calculated using matrix trace operation to ensure the stability and accuracy of data analysis results. Combining the deviation rate of the distribution difference quantitative indicator and the separability measure of the gradient distribution, the noise intensity control parameter is dynamically calculated to precisely adjust the degree of noise injection, ensuring that the injection of noise does not affect the validity of statistical results. Through orthogonal transformation and Laplace noise injection, the noise only affects the non - critical information part of the data, effectively avoiding privacy leakage. The secondary verification after noise injection further ensures that the differences in latent features between research centers are not overly interfered by noise, maintaining the consistency and credibility of the data. Finally, the optimized gradient covariance matrix adjusts the noise intensity control parameter through the maximum element offset rate of the covariance matrix to ensure that the impact of noise on the analysis results is minimized, achieving a balance between privacy protection and data analysis efficiency. Brief Description of the Drawings
[0048] Figure 1 Schematic diagram of a data verification method for clinical research data according to the present invention. Detailed Embodiments
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0050] Embodiment 1
[0051] Figure 1 A data verification method for clinical research data according to the present invention is given, which includes the following steps:
[0052] S1: Collect clinical research data from each clinical research center and establish a standardized matrix, perform sparsifying non - negative matrix factorization on the standardized matrix, and establish a latent feature space;
[0053] S2: By calculating the distribution difference quantitative indicators of the corresponding data of latent features between different research centers, conduct a primary verification of the feature distribution differences between different research centers. If the distribution difference quantitative indicator between any two research centers is greater than or equal to the set distribution difference threshold, activate the noise injection mechanism;
[0054] S3: Perform a linear regression analysis on the potential features of each research center, iteratively extract the local gradient matrix and the global gradient tensor matrix of each research center, construct the within-class covariance matrix and between-class covariance matrix of the gradient distribution, and obtain the separability measure of the gradient distribution between research centers through matrix trace operation;
[0055] S4: Based on the deviation rate of the distribution difference quantization index and the separability measure of the gradient distribution, comprehensively calculate the noise intensity control parameter;
[0056] S5: Perform an orthogonal transformation on the local gradient matrix, construct an orthogonal projection matrix, generate noise through a Laplace distribution, and project the generated noise onto the direction adjusted by the orthogonal projection matrix to obtain the final noise direction;
[0057] S6: Inject the Laplace noise generated based on the noise intensity control parameter into the corresponding research center potential feature matrix, perform a secondary verification on the differences between the potential feature matrices of each research center, and adjust the noise intensity control parameter according to the distribution difference quantization index deviation of the secondary verification at the set scale step size;
[0058] S7: Re-perform a linear regression analysis on the potential features of each research center and calculate the gradient covariance matrix after noise enhancement, adjust the noise intensity control parameter based on the maximum element offset rate of the set covariance matrix, and turn off the noise injection mechanism after the adjustment is completed.
[0059] In S1, collect the clinical research data of each clinical research center and establish a standardized matrix, perform sparse non-negative matrix factorization on the standardized matrix, and establish a potential feature space.
[0060] Collect the clinical research data of each clinical research center, perform a unified format processing on the clinical research data of each research center, and establish a standardized matrix where i is the research center number, n i is the number of samples of the i-th research center, d is the number of features of each clinical research sample, and the features are the expressions of the parameter items of the basic information, clinical indicators, gene data, treatment information, and experimental results of the sample patients.
[0061] Perform feature selection through the L1 regularization method, and set the expression of the optimization objective of the minimum reconstruction error of matrix factorization for the standardized matrix X i Perform sparse non-negative matrix factorization, and the specific expression of the optimization objective is:
[0062]
[0063] In the formula, F iLet be the latent feature matrix, Q be the relationship matrix between the data in the latent feature matrix and the features of the original data, and λ be a regularization intensity hyperparameter preset comprehensively according to the feature intensity and data density of the samples. The larger the value of the hyperparameter, the stronger the sparsity of the model, that is, more features will be zeroed out (i.e., features are removed). However, an overly large value will lead to excessive sparsification, possibly losing useful features and resulting in underfitting. On the contrary, a smaller value may retain too many features, leading to overfitting.
[0064] Set the constraint condition F i ≥0, Q≥0. Under the condition of satisfying the constraint condition, retain the latent feature matrix with a variance contribution degree greater than the set contribution degree value (default set to 85%) in the latent feature matrix, integrate the retained latent feature matrix, and establish a latent feature space.
[0065] Extract the mean and variance of the data corresponding to the retained latent features in the latent feature matrix based on high-dimensional statistical analysis.
[0066] In S2, quantify the distribution difference of features among different research centers by calculating the distribution difference quantification index of the data corresponding to the latent features among different research centers. If the distribution difference quantification index between any two research centers is greater than or equal to the set distribution difference threshold, activate the noise injection mechanism.
[0067] Obtain the mean and variance of the data corresponding to the latent features of each research center in the latent feature space, and calculate the JS divergence and KL divergence of the data corresponding to the latent features among different research centers in parallel. The JS divergence is symmetric and suitable for comparing probability distributions, while the KL divergence mainly quantifies the difference between two Gaussian distributions through the difference in mean and variance. The specific calculation expression is:
[0068]
[0069] In the formula, D KL (p||q) and D JS (p||q) are the KL divergence and JS divergence of research center p and research center q respectively, m is the mixed distribution of the sample data of research center p and q, k is the number of latent features in the latent feature space, j is the subscript count of k, and the value range is [0, k], μ pj 、μ qj are the means of the data corresponding to the j-th latent feature of research center p and q respectively, are the variances of the data corresponding to the j-th latent feature of research center p and q respectively. Perform a weighted (the weights are set comprehensively based on the similarity of the data distributions between p and q) harmonic operation on the KL divergence and JS divergence, and use the harmonic mean as the distribution difference quantification index.
[0070] Set a dynamic distribution difference threshold according to the research sample size of each research center (default setting is 5%), conduct a verification on the distribution differences of characteristics among different research centers, compare pairwise the quantification indicators of distribution differences among different research centers in the potential feature space. If there is any quantification indicator of distribution difference between any two research centers greater than or equal to the distribution difference threshold, activate the noise injection mechanism.
[0071] In S3, perform a linear regression analysis on the potential features of each research center, iteratively extract the local gradient matrix and global gradient tensor matrix of each research center, construct the within-class covariance matrix and between-class covariance matrix of the gradient distribution, and obtain the separability measure of the gradient distribution among research centers through matrix trace operation.
[0072] Perform a linear regression analysis on the potential features of each research center in the potential feature space, set the mean squared error as the loss function, and based on the SMPC protocol, iteratively extract the local gradient matrix G of each research center (t) , where t represents the corresponding iteration number;
[0073] According to the local gradient matrix extracted in each iteration, perform a global mean processing on the local gradient matrix of each research center, calculate the global gradient tensor matrix of this iteration, and the calculation result is represented by g (t) ;
[0074] Calculate the covariance matrix of the local gradient matrix G (t) in each iteration and the global gradient tensor matrix g (t) , mark it as the within-class covariance matrix, represented by ∑W, specifically:
[0075]
[0076] In the formula, N is the total number of iterations.
[0077] Obtain the mean of the elements in all global gradient tensor matrices during the iteration process, calculate the historical mean gradient matrix and the covariance matrix of g (t) , mark it as the between-class covariance matrix, represented by ∑B, specifically:
[0078]
[0079] Based on the within-class covariance matrix and between-class covariance matrix, construct the separability measure of the gradient distribution among research centers through matrix trace operation, and the specific expression is:
[0080]
[0081] In the formula, β is the linear separability coefficient, and tr is the matrix trace operation identifier.
[0082] In S4, based on the deviation rate of the distribution difference quantization index and the separability measure of the gradient distribution, the noise intensity control parameter is comprehensively calculated.
[0083] Based on the deviation rate between the distribution difference quantization index and the distribution difference threshold between any two research centers, combined with the calculated linear separability coefficient, the noise intensity control parameter is comprehensively calculated. The specific calculation expression is:
[0084]
[0085] In the formula, γ(p,q) is the noise intensity control parameter for controlling the generated noise with respect to the potential feature matrix corresponding to research center p relative to q, D(p,q) is the distribution difference deviation rate between research centers p and q, ω is the sample-dimension ratio adjustment factor, n p 、n q are the sample numbers of research centers p and q respectively, c is the number of features in the potential feature space, ∈ is the set smoothing parameter, and the default setting is 0.01.
[0086] In S5, the local gradient matrix is orthogonally transformed to construct an orthogonal projection matrix. Noise is generated through the Laplace distribution, and the generated noise is projected onto the direction adjusted by the orthogonal projection matrix to obtain the final noise direction. The expression for constructing the noise projection matrix orthogonal to the main gradient direction is:
[0087] P = I - g (N) (g (N) ) T ;
[0088] where I is the identity matrix, g (N) is the global gradient tensor matrix extracted in the last iteration, and P is the orthogonal projection matrix.
[0089] A basic noise is set (comprehensively set based on the size of the data distribution difference of the potential features), and then the noise distribution scale is controlled by the noise intensity control parameter γ(p,q).
[0090] In S6, the Laplace noise generated based on the noise intensity control parameter is injected into the corresponding research center's potential feature matrix, the differences between the potential feature matrices of each research center are secondarily verified, and the noise intensity control parameter is adjusted according to the distribution difference quantization index deviation of the secondary verification at the set scale step.
[0091] Using the noise intensity control parameter, Laplace noise matching the dimension of the latent feature matrix is generated. The generation of Laplace noise is based on the range of the data distribution and the noise amplitude. Through this noise, the elements of the feature matrix can be perturbed. The amount of perturbation is determined according to the noise intensity control parameter and the value of each element in the feature matrix. The injected latent feature matrix will carry noise information, while ensuring privacy, still maintaining the original structural features.
[0092] Based on the latent feature matrix after noise injection, recalculate the distribution difference quantization index of each research center, and use the aforementioned JS divergence and KL divergence to conduct a secondary verification of the differences between the latent feature matrices of each research center. To ensure that after adding noise, the distribution difference between research centers does not exceed the preset threshold, thus ensuring the privacy protection effect during the model training process.
[0093] Compare the distribution difference quantization indexes of different research centers after noise injection pairwise. If there is any distribution difference quantization index between any two research centers greater than or equal to the distribution difference threshold, increase the noise intensity control parameter by the set scale step (default set to 1%). Gradually enhancing the perturbation effect of the noise can ensure that the system will not immediately over-perturb the data.
[0094] If the distribution difference quantization indexes between all research centers are less than the distribution difference threshold, end the process of increasing the noise intensity control parameter.
[0095] In S7, re-conduct a linear regression analysis on the latent features of each research center and calculate the gradient covariance matrix after noise enhancement. Based on the set maximum element offset rate of the covariance matrix, adjust the noise intensity control parameter, and turn off the noise injection mechanism after adjustment.
[0096] Re-conduct a linear regression analysis on the latent features of each research center and iteratively extract the local gradient matrix and the corresponding global gradient tensor matrix of each research center, calculate the covariance matrix of the local gradient matrix and the global gradient tensor matrix during the iterative process, and mark it as the gradient covariance matrix after noise enhancement;
[0097] Set the maximum element offset rate of the covariance matrix (default set to 3%, used to quantify the change amplitude of the gradient matrix elements during the noise injection process. The purpose of setting the maximum element offset rate is to ensure that during the noise enhancement process, the change of the matrix will not be too large, thus maintaining the effectiveness of the gradient and the analysis results of the data model).
[0098] Compare the elements of the within-class covariance matrix without injected noise with the gradient covariance matrix with enhanced noise. If there are elements with an element offset rate greater than the maximum element offset rate, proportionally downscale the noise intensity control parameter according to the ratio exceeding the maximum element offset rate, and regenerate the noise injection. The adjustment of the noise intensity control parameter will affect the intensity of subsequent noise generation, ensuring moderate perturbation so that privacy protection does not overly affect model analysis.
[0099] Repeatedly verify the offset rate of the elements in the gradient covariance matrix with enhanced noise until the offset rate of all elements is less than the maximum element offset rate, then end the adjustment of the noise intensity control parameter and simultaneously turn off the noise injection mechanism.
[0100] In the collaborative analysis of the multi-party research center, after the data of each party is adjusted by the noise injection mechanism, the data and model contributions of each research center remain private, ensuring that the data privacy of each party is not leaked and enabling effective joint data analysis and calculation.
[0101] The above formulas are all dimensionless and take their numerical calculations. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.
[0102] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center containing one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0103] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0104] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0105] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.
[0106] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0107] In addition, the functional modules in each embodiment of this application can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0108] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0109] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0110] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A data verification method for clinical research data, characterized in that: The steps include: S1: Collect clinical research data from each clinical research center and establish a standardized matrix, perform sparse non-negative matrix factorization on the standardized matrix, and establish a latent feature space; S2: By calculating the distribution difference quantitative index of the corresponding data of the potential features between different research centers, the feature distribution difference of different research centers is verified. If the distribution difference quantitative index between any two research centers is greater than or equal to the set distribution difference threshold, the noise injection mechanism is activated; S3: Perform linear regression analysis on the potential features of each research center, iteratively extract the local gradient matrix and global gradient tensor matrix of each research center, construct the intra-class covariance matrix and inter-class covariance matrix of the gradient distribution, and obtain the separability measure of the gradient distribution between research centers through matrix trace operation; S4: Based on the deviation rate of the distribution difference quantification index and the separability measure of the gradient distribution, the noise intensity control parameter is comprehensively calculated; S5: Perform an orthogonal transformation on the local gradient matrix, construct an orthogonal projection matrix, generate noise through Laplace distribution, project the generated noise to the direction adjusted by the orthogonal projection matrix, and obtain the final noise direction; S6: Inject the Laplace noise generated by the noise intensity control parameter into the corresponding research center potential feature matrix, perform a secondary verification on the difference between the potential feature matrices of each research center, quantify the indicator deviation based on the distribution difference of the secondary verification, and adjust the noise intensity control parameter according to the set scale step size; S7: Perform linear regression analysis on the potential features of each research center again and calculate the gradient covariance matrix after noise enhancement. Adjust the noise intensity control parameters based on the set maximum element offset rate of the covariance matrix. After the adjustment is completed, turn off the noise injection mechanism.
2. A data verification method for clinical research data according to claim 1, characterized in that: In S1, clinical research data from each clinical research center is collected and a standardized matrix is established. The standardized matrix is subjected to sparse non-negative matrix decomposition to establish a latent feature space, which specifically includes: Collect clinical research data from each clinical research center, unify the format of the clinical research data from each research center, and establish a standardized matrix where i is the research center number, n i is the number of samples in the i-th research center, and d is the number of characteristics of each clinical research sample; The feature selection is performed by L1 regularization method, and the expression of the optimization objective of minimum reconstruction error of matrix decomposition is set to the standardized matrix X i Perform sparse non-negative matrix decomposition, and the optimization target expression is as follows: In the formula, F i is the latent feature matrix, Q is the relationship matrix between the data in the latent feature matrix and the original data features, and λ is the preset regularization strength hyperparameter; Set the constraint F i ≥0, Q≥0, retain the potential feature matrix whose variance contribution is greater than the set contribution value in the potential feature matrix while satisfying the constraint conditions, integrate the retained potential feature matrix, and establish the potential feature space; The mean and variance of the data corresponding to the latent features retained in the latent feature matrix are extracted based on high-dimensional statistical analysis.
3. A data verification method for clinical research data according to claim 2, characterized in that: In S2, by calculating the distribution difference quantitative index of the corresponding data of the potential features between different research centers, the feature distribution difference of different research centers is verified. If the distribution difference quantitative index between any two research centers is greater than or equal to the set distribution difference threshold, the noise injection mechanism is activated, specifically including: Obtain the mean and variance of the latent feature corresponding data of each research center in the latent feature space, and calculate the weighted harmonic mean of the JS divergence and KL divergence of the latent feature corresponding data between different research centers in parallel as a quantitative indicator of distribution difference; A dynamic distribution difference threshold is set according to the research sample size of each research center, and the feature distribution differences of different research centers are verified. The distribution difference quantitative indicators of different research centers in the potential feature space are compared pairwise. If the distribution difference quantitative indicator between any two research centers is greater than or equal to the distribution difference threshold, the noise injection mechanism is activated.
4. A data verification method for clinical research data according to claim 3, characterized in that: In S3, linear regression analysis is performed on the potential features of each research center, the local gradient matrix and global gradient tensor matrix of each research center are iteratively extracted, the intra-class covariance matrix and inter-class covariance matrix of the gradient distribution are constructed, and the separability measure of the gradient distribution between research centers is obtained through matrix trace operation. Specifically, it includes: A linear regression analysis is performed on the latent features of each research center in the latent feature space, and the mean error variance is set as the loss function. The local gradient matrix G of each research center is iteratively extracted based on the SMPC protocol with the goal of minimizing the loss function. (t) , t represents the corresponding number of iterations; According to the local gradient matrix extracted in each iteration, the local gradient matrix of each research center is processed by global mean, and the global gradient tensor matrix of this iteration is calculated. The calculation result is expressed as g (t) express; Calculate the local gradient matrix G at each iteration (t) With the global gradient tensor matrix g (t) The covariance matrix of is marked as the intra-class covariance matrix and is represented by ∑W; Get the mean of all elements in the global gradient tensor matrix during the iteration process and calculate the historical mean gradient matrix With g (t) The covariance matrix of is marked as the between-class covariance matrix and is represented by ∑B; Based on the intra-class covariance matrix and the inter-class covariance matrix, the separability measure of the gradient distribution between research centers is constructed through matrix trace operation. The specific expression is: Where β is the linear separability coefficient and tr is the matrix trace operation identifier.
5. A data verification method for clinical research data according to claim 4, characterized in that: In S4, based on the deviation rate of the distribution difference quantification index and the separability metric of the gradient distribution, the noise intensity control parameters are comprehensively calculated, including: Based on the deviation rate between the distribution difference quantitative index and the distribution difference threshold between any two research centers, combined with the calculated linear separability coefficient, the noise intensity control parameter is comprehensively calculated. The specific calculation expression is: Where γ(p,q) is the intensity control parameter of the noise generated by the potential feature matrix corresponding to the research center p relative to q, D(p,q) is the distribution difference deviation rate between the research centers p and q, ω is the sample-dimension ratio adjustment factor, and n p 、n q are the number of samples in the research center p and q respectively, c is the number of features in the latent feature space, and ∈ is the set smoothing parameter.
6. A data verification method for clinical research data according to claim 5, characterized in that: In S5, the local gradient matrix is orthogonally transformed to construct an orthogonal projection matrix, noise is generated through Laplace distribution, and the generated noise is projected to the direction adjusted by the orthogonal projection matrix to obtain the final noise direction.
7. A method for verifying clinical research data according to claim 6, characterized in that: In S6, the Laplace noise generated by the noise intensity control parameter is injected into the corresponding research center potential feature matrix, and the difference between the potential feature matrices of each research center is verified twice. Based on the distribution difference quantification indicator deviation of the secondary verification, the noise intensity control parameter is adjusted according to the set scale step size, which specifically includes: Injecting Laplace noise generated by a noise intensity control parameter based on the corresponding matrix element of the research center latent feature matrix; Recalculate the quantitative index of distribution difference of each research center based on the latent feature matrix after noise injection, and conduct a second verification on the difference between the latent feature matrices of each research center; The distribution difference quantitative indicators of different research centers after noise injection are compared pairwise. If the distribution difference quantitative indicators between any two research centers are greater than or equal to the distribution difference threshold, the noise intensity control parameter is increased according to the set scale step size; If the quantitative indicators of distribution differences among all research centers are less than the distribution difference threshold, the process of increasing the noise intensity control parameters is terminated.
8. A data verification method for clinical research data according to claim 7, characterized in that: In S7, linear regression analysis is performed on the potential features of each research center again and the gradient covariance matrix after noise enhancement is calculated. The noise intensity control parameter is adjusted based on the maximum element offset rate of the covariance matrix. After the adjustment is completed, the noise injection mechanism is turned off, which specifically includes: Re-perform linear regression analysis on the potential features of each research center and iteratively extract the local gradient matrix of each research center and the corresponding global gradient tensor matrix. Calculate the covariance matrix of the local gradient matrix and the global gradient tensor matrix during the iteration process, which is marked as the noise-enhanced gradient covariance matrix. Set the maximum element offset rate of the covariance matrix, compare the elements of the intra-class covariance matrix without noise injection with the gradient covariance matrix with noise enhancement, and if there are elements with an element offset rate greater than the maximum element offset rate, then adjust the noise intensity control parameter proportionally according to the ratio exceeding the maximum element offset rate, and regenerate the noise injection; Repeatedly check the offset rate of the elements in the noise-enhanced gradient covariance matrix until the offset rates of all elements are less than the maximum element offset rate, then end the adjustment of the noise intensity control parameters and turn off the noise injection mechanism.