Data verification method, electronic device, and computer-readable storage medium
By receiving and processing intermediate sample data, calculating global statistical data and verification values, the accurate verification of user data in the privacy industry is achieved, and the data sharing challenges brought about by privacy and confidentiality is solved. It is suitable for data verification in medicine, economics and other fields.
Patent Information
- Application Number
- CN202510138506.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-02-08
AI Technical Summary
In the privacy industry, the verification of user data of all parties cannot share original data due to privacy and confidentiality issues, which makes it difficult for traditional centralized t-testing methods to achieve effective data verification.
By receiving intermediate sample data sent by each participant device, calculating local statistical data and local quantity, constructing a t distribution, calculating global quantity and global statistical data, obtaining the test value, realizing t-test, determining whether the user data meets expectations, and avoiding sharing of original data.
Without sharing original data, accurate data verification is achieved, user privacy is protected, and data privacy protection and security problems in traditional methods are solved. It is suitable for user data verification in more scenarios.
Smart Images

Figure CN119577400B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing. Specifically, it relates to a data verification method, an electronic device, and a computer-readable storage medium. Background Art
[0002] Currently, for some privacy-sensitive industries, such as hospitals, banks, etc., due to the privacy and confidentiality of data, when verifying the data of each party's users, in order to protect the privacy of users, the original data of each participating party cannot be shared with other participating parties, which poses a challenge to the verification of user data. Summary of the Invention
[0003] The purpose of this application is to provide a data verification method, an electronic device, and a computer-readable storage medium, which can achieve data verification without sharing the original user data.
[0004] In a first aspect, an embodiment of this application provides a data verification method, including: receiving intermediate sample data sent by each participating party device, where the intermediate sample data includes local statistical data and local quantity obtained based on the sample data included in the participating party device, and the sample data includes user data; calculating a global quantity according to the local quantity of each participating party device; calculating global statistical data according to the global quantity and the local statistical data of each participating party device; calculating a verification value according to the global quantity and the global statistical data, where the verification value is used to determine whether the user data meets the expectation.
[0005] Through the above embodiments, it is possible to verify user data without sharing the original user data, improving the privacy of user data. Further, by performing local calculations in each participating party device to determine the local statistical data and local quantity, and then aggregating the local data calculated by each participating party device, since the result is determined based on each participating party, the data verification result can be made more objective and the obtained verification result can be more accurate.
[0006] In an optional embodiment, the calculating a verification value according to the global quantity and the global statistical data includes: constructing a t-distribution that satisfies the degrees of freedom determined by the global quantity based on the global quantity; calculating a verification value based on the global statistical data and the t-distribution.
[0007] In the above embodiment, by using the characteristics of the t-distribution, a hypothesis testing method of t-test can be implemented to compare whether there is a significant difference in the means of two types of samples and determine the user data and the compared parameters.
[0008] In an alternative embodiment, the global statistical data includes the global sample mean and the global sample standard deviation; calculating the test value based on the global statistical data and the t-distribution includes: substituting the set sample pseudo-mean, the global sample mean, and the global sample standard deviation into the t-distribution to calculate the test value.
[0009] In the above embodiment, in the case of a single sample, it can be compared with the set sample pseudo-mean to determine whether the user data meets the test criteria.
[0010] In an alternative embodiment, substituting the set sample pseudo-mean, the global sample mean, and the global sample standard deviation into the t-distribution to calculate the test value is determined by the following formula: ; where represents the global sample mean; represents the global sample standard deviation; D represents the global quantity; represents the set sample pseudo-mean.
[0011] In an alternative embodiment, the global statistical data includes the mean of the global paired sample differences and the standard deviation of the global paired sample differences; calculating the test value based on the global statistical data and the t-distribution includes: substituting the mean of the global paired sample differences and the standard deviation of the global paired sample differences into the t-distribution to calculate the test value.
[0012] In an alternative embodiment, substituting the mean of the global paired sample differences and the standard deviation of the global paired sample differences into the t-distribution to calculate the test value is determined by the following formula: ; where represents the mean of the global paired sample differences; represents the standard deviation of the global paired sample differences; D represents the global quantity.
[0013] In an alternative embodiment, the global statistical data includes a global mean and a global standard deviation; the local statistical data includes a local sum of squares and a local sum; calculating the global statistical data based on the global quantity and the local statistical data of each participating device includes: calculating the global sum based on the local sum in the local statistical data, where the global sum is the sum of global sample sums or global paired sample differences; obtaining the global mean based on the global quantity and the global sum, where the global mean is the global sample mean or the mean of global paired sample differences; calculating the global sum of squares based on the local sum of squares in the local statistical data; obtaining the global standard deviation based on the global sum, the global mean, the global sum of squares, and the global quantity, where the global standard deviation includes the global sample standard deviation or the standard deviation of global paired sample differences.
[0014] In an alternative embodiment, the intermediate sample data includes first intermediate sample data and second intermediate sample data; the first intermediate sample data includes first local statistical data and a first local quantity obtained based on first sample data included in the participating device, where the first sample data is user data in a first state of the user; the second intermediate sample data includes second local statistical data and a second local quantity obtained based on second sample data included in the participating device, where the second sample data is user data in a second state of the user; the global quantity includes a first global quantity and a second global quantity; the global statistical data includes first global statistical data and second global statistical data; calculating the global quantity based on the local quantity of each participating device includes: calculating the first global quantity based on the first local quantity of each participating device, and calculating the second global quantity based on the second local quantity of each participating device; calculating the global statistical data based on the global quantity and the local statistical data of each participating device includes: calculating the first global statistical data based on the first global quantity and the first local statistical data of each participating device, and calculating the second global statistical data based on the second global quantity and the second local statistical data of each participating device; calculating the test value based on the global statistical data and the t-distribution includes: calculating the test value based on the first global statistical data, the second global statistical data, the first global quantity, the second global quantity, and the t-distribution.
[0015] In the above embodiment, in the case of two samples, a comparison can be formed between the two types of samples to determine whether the user data meets the requirements.
[0016] In an alternative embodiment, the first global statistical data includes a first global sample mean and a first global sample variance; the second global statistical data includes a second global sample mean and a second global sample variance; calculating a test value according to the first global statistical data, the second global statistical data, the first global quantity, the second global quantity, and the t-distribution includes: substituting the first global sample mean, the first global sample variance, the second global sample mean, the second global sample variance, the first global quantity, and the second global quantity into the t-distribution to calculate the test value.
[0017] In an alternative embodiment, substituting the first global sample mean, the first global sample variance, the second global sample mean, the second global sample variance, the first global quantity, and the second global quantity into the t-distribution to calculate the test value is implemented by the following formula: ; where represents the first global sample mean; represents the second global sample mean; represents the common variance; represents the first global quantity; represents the second global quantity; the first global sample variance is equal to the second global sample variance and is equal to the common variance.
[0018] In an alternative embodiment, substituting the first global sample mean, the first global sample variance, the second global sample mean, the second global sample variance, the first global quantity, and the second global quantity into the t-distribution to calculate the test value is implemented by the following formula: ; where represents the first global sample mean; represents the second global sample mean; represents the first global quantity; represents the second global quantity; represents the first global sample variance; represents the second global sample variance.
[0019] In a second aspect, an embodiment of the present application provides an electronic device, including: a processor and a memory, where the memory stores machine-readable instructions executable by the processor, and when the electronic device runs, the machine-readable instructions are executed by the processor to perform the steps of the data verification method described in any one of the above.
[0020] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the data verification method described in any one of the above.
[0021] In a fourth aspect, an embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the data verification method described in any one of the above. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0023] Figure 1 It is a schematic diagram of the interaction of the operating system for the data verification method provided by the embodiment of the present application;
[0024] Figure 2 It is a block diagram of an electronic device provided by the embodiment of the present application;
[0025] Figure 3 It is a flowchart of the data verification method provided by the embodiment of the present application;
[0026] Figure 4 It is an alternative flowchart of step 340 of the data verification method provided by the embodiment of the present application;
[0027] Figure 5 It is an alternative flowchart of step 330 of the data verification method provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The following will describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application.
[0029] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.
[0030] Currently, for some privacy-related industries, since the original data of the industry is private and confidential, it is not possible to share the original data with other participating parties. However, if it is necessary to verify the books in the privacy industry, it may lead to difficulties in data verification. For example, each participating party in a t-test (such as hospitals, banks, etc.) cannot share its original data with other participating parties, which poses challenges to the traditional centralized t-test method.
[0031] In statistics, the t-test is a hypothesis testing method used to compare whether there are significant differences in the means of two types of samples. The t-test is widely used in various fields, including medicine, economics, social sciences, etc. It can assist in making decisions or inferences by determining whether the differences between sample means are statistically significant. Through the t-test, it can be determined whether the average value of a set of data is significantly different from the average value of another set of data, thereby identifying the phenomena and laws behind the data.
[0032] Based on the above research, an embodiment of the present application can provide a data verification method, an electronic device, and a computer-readable storage medium, which can achieve data verification without sharing the original data.
[0033] For ease of understanding of this embodiment, first, a detailed introduction to the operating environment for executing a data verification method disclosed in an embodiment of the present application is provided.
[0034] As Figure 1 shown, it is a schematic diagram of the operating environment of the data verification method provided by an embodiment of the present application. The operating environment of this data verification method may include a central server and multiple participating party devices.
[0035] The central server communicates with one or more participating party devices through a network for data communication or interaction. The central server can be a network server, a database server, etc. The participating party devices can be personal computers (PCs), tablets, smartphones, personal digital assistants (PDAs), etc. The participating party can also be a network server, a database server, etc.
[0036] The participating party devices can be internal devices of each participating party in the application field. For example, each participating party can be a hospital, a bank, a social organization, etc. The participating party devices can then be internal devices of organizations such as hospitals, banks, and social organizations.
[0037] Taking the participating party as a hospital as an example, if ten hospitals are participating parties, the participating party devices of the ten hospitals can be used for the original data within their corresponding hospitals. The original data can be the medical data, physical parameters, medication data, etc. of patients. The medical data, physical parameters, and medication data of patients are the privacy data of patients and cannot be transmitted to other participating party devices. In the embodiments of the present application, the data sent out by the participating party device is the data after processing the medical data, physical parameters, medication data, etc. of patients.
[0038] As Figure 2 shown, it is a block diagram of an electronic device. The electronic device 200 may include a memory 211 and a processor 213. Those of ordinary skill in the art can understand that Figure 2 the structure shown is only illustrative and does not limit the structure of the electronic device 200. For example, the electronic device 200 may further include more or fewer components than Figure 2 shown, or have a different configuration from Figure 2 shown.
[0039] The above-mentioned memory 211 and processor 213 are directly or indirectly electrically connected to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The above-mentioned processor 213 is used to execute the executable module stored in the memory.
[0040] Among them, the memory 211 can be, but is not limited to, a random access memory (Random Access Memory, abbreviated as RAM), a read-only memory (Read Only Memory, abbreviated as ROM), a programmable read-only memory (Programmable Read-Only Memory, abbreviated as PROM), an erasable programmable read-only memory (Erasable Programmable Read-Only Memory, abbreviated as EPROM), an electrically erasable programmable read-only memory (Electric Erasable Programmable Read-Only Memory, abbreviated as EEPROM), etc. Among them, the memory 211 is used to store a program, and after the processor 213 receives an execution instruction, it executes the program. The method executed by the electronic device 200 defined by the process disclosed in any embodiment of the embodiments of the present application can be applied to the processor 213 or implemented by the processor 213.
[0041] The above-mentioned processor 213 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 213 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0042] In this embodiment, Figure 1 the shown central server may include components of the electronic device 200.
[0043] The electronic device 200 in this embodiment can be used to execute each step in the various methods provided in the embodiments of the present application. The implementation process of the data verification method will be described in detail through several embodiments below.
[0044] The data verification method provided in the embodiments of the present application can be applied to an electronic device, and the steps in the data verification method are executed through the electronic device. Exemplarily, the electronic device may be Figure 1 the shown central server. The central server can obtain the data of each participating device, and the data obtained by the central server from each participating device can be the data after being processed by the participating device. The central server does not obtain the original data from each participating device.
[0045] Please refer to Figure 3 , which is a flowchart of the data verification method provided in the embodiments of the present application. The following will elaborate on Figure 3 the specific process shown.
[0046] Step 310, receive the intermediate sample data sent by each participating device.
[0047] Among them, the intermediate sample data includes local statistical data and local quantity obtained based on the sample data included in the participating device, and the sample data includes user data.
[0048] This intermediate sample data can be the data obtained after the participating party device processes the original sample data. Exemplarily, the local statistical data can represent the result of statistical processing on the sample data contained in the corresponding participating party device. The statistical processing can include processing such as summation, mean calculation, variance calculation, standard deviation calculation, etc. The local quantity can represent the number of samples in the sample data contained in the corresponding participating party device.
[0049] This data verification method can be applied in fields such as medicine and social population distribution prediction.
[0050] Taking the medical field as an example, this data verification method can be used to verify whether the physical state after taking medicine meets the expectations. Exemplarily, the user data can include the user's physical parameters. The physical parameters can include but are not limited to parameters such as blood pressure, weight, body temperature, etc. of the subject before and after the test.
[0051] For social population distribution prediction, this data verification method can be used to verify the change in the number of people flowing in a certain activity area, etc. Exemplarily, the user data can include data such as the number of users and the age segments of users.
[0052] Step 320, calculate the global quantity according to the local quantities of each participating party device.
[0053] Exemplarily, the global quantity can be directly obtained by summing the received local quantities.
[0054] Step 330, calculate the global statistical data according to the global quantity and the local statistical data of each participating party device.
[0055] Regarding the different information represented by the local statistical data, the calculation methods can be different.
[0056] Exemplarily, when the local statistical data is the mean, it can be directly summed and then the global mean can be determined based on the global quantity.
[0057] Exemplarily, when the local statistical data is the variance, the global variance is calculated through a transformation formula based on the global quantity and the local statistical data.
[0058] When the local statistical data contains multiple types of numerical values, the global data can be calculated for each type of data.
[0059] Step 340, calculate the test value according to the global quantity and the global statistical data.
[0060] Among them, the test value is used to determine whether the user data meets the expectations.
[0061] The t-test is a test method that can compare whether there are significant differences in the means of two types of samples. In one implementation, the t-test can be achieved by constructing a t-distribution. As Figure 4 shown, step 340 described above may include: step 341, constructing a t-distribution that satisfies the degrees of freedom determined by the global quantity; step 342, calculating a test value based on the global statistical data and the t-distribution.
[0062] Exemplarily, the test value may be a t-test value.
[0063] In this embodiment, the global statistical data includes the global mean and the global standard deviation; the local statistical data includes the local sum of squares and the local sum. As Figure 5 shown, step 330 described above may include: steps 331 to 333.
[0064] Step 331, calculating the global sum based on the local sum in the local statistical data.
[0065] Exemplarily, the global sum can be directly calculated from all the local sums.
[0066] Step 332, obtaining the global mean based on the global quantity and the global sum.
[0067] Exemplarily, the global mean can be directly obtained by dividing the global sum by the global quantity.
[0068] Step 333, calculating the global sum of squares based on the local sum of squares in the local statistical data.
[0069] Exemplarily, the global sum of squares can be obtained by adding up the local sums of squares sent by all participating device.
[0070] Step 334, obtaining the global standard deviation based on the global sum, the global mean, the global sum of squares, and the global quantity.
[0071] Exemplarily, the relationship for obtaining the global standard deviation based on the global sum, the global mean, the global sum of squares, and the global quantity can be determined by transforming the standard deviation calculation formula.
[0072] Exemplarily, the standard deviation calculation formula can be expressed as:
[0073] ;
[0074] where is the i-th value in the data sequence to be calculated, and n represents the number of parameters used when calculating the standard deviation.
[0075] The calculation formula for s can be transformed to obtain:
[0076] ;
[0077] Among them, represents the global sum; represents the global mean; represents the global sum of squares.
[0078] Through the transformation of the above formula, the central server can calculate the global mean and the global standard deviation without obtaining the original sample data.
[0079] In one embodiment, the global statistical data may include the global sample mean and the global sample standard deviation.
[0080] Exemplarily, the above step 330 may include calculating the global sample sum according to the local sum in the local statistical data; obtaining the global sample mean according to the global quantity and the global sample sum; calculating the global sum of squares according to the local sum of squares in the local statistical data; and obtaining the global sample standard deviation according to the global sample sum, the global sample mean, the global sum of squares, and the global quantity.
[0081] The above step 342 may include: substituting the set sample pseudo-mean, the global sample mean, and the global sample standard deviation into the t-distribution to calculate the test value.
[0082] The set sample pseudo-mean may be a value pre-set by the user, and this value may be a value obtained through experience.
[0083] Taking the sample data as the blood pressure data of doctors and patients as an example, the set sample pseudo-mean may be a value determined according to the blood pressure standard of normal users. Taking the data verification method for realizing new drug testing as an example, the sample data may be the physical parameters of doctors and patients, and the set sample pseudo-mean may be a value determined according to the physical parameter standard of patients in the diseased state for which the new drug is used for diagnosis and treatment.
[0084] The following describes the calculation of the above parameters through some formulas.
[0085] Exemplarily, substituting the set sample pseudo-mean, the global sample mean, and the global sample standard deviation into the t-distribution to calculate the test value is determined by the following formula:
[0086] ;
[0087] Among them, represents the global mean; represents the global standard deviation; D represents the global quantity; represents the set sample pseudo-mean.
[0088] In this embodiment, each participating party's device can first perform computational processing on the sample data it contains and count the local quantity of the sample data it contains. , calculate the local sum and the local sum of squares, and transmit these three pieces of information to the central server. Taking the participating party as an example, the local sum calculated by its participating party's device locally is as follows:
[0089] ;
[0090] The local sum of squares can be achieved in the following way:
[0091] ;
[0092] The central server collects the local quantities, calculates the local sums, and calculates the local sums of squares uploaded by each participating party, and sums them respectively to obtain the global quantity, the global sum, and the global sum of squares. Among them, the global quantity is achieved in the following way:
[0093] ;
[0094] The global sum is calculated in the following way:
[0095] ;
[0096] The calculation process of the global data sum of squares is as follows:
[0097] ;
[0098] The central server calculates the global test value based on the above global data, where the global mean can be achieved through the following calculation method:
[0099] ;
[0100] The global standard deviation s can be achieved through the following calculation method:
[0101] .
[0102] In the above embodiment, a single sample can be tested. For example, if a new drug needs to be tested, it is necessary to check whether the average blood pressure of doctors and patients using the new drug is significantly different from the blood pressure values of normal people. In this example, the single sample can be the blood pressure of doctors and patients using the new drug, and each participating party can be multiple hospitals participating in the test of the new drug. The participating party is the j-th hospital participating in the new drug test, is the blood pressure sequence of the patients using the new drug in the j-th hospital.
[0103] Exemplarily, if the average blood pressure of doctors and patients using the new drug is significantly different from that of normal people, it indicates that the effect of the new drug is lacking; if the average blood pressure of doctors and patients using the new drug is not significantly different from that of normal people, it indicates that the effect of the new drug is better.
[0104] Taking the t-test value calculated above as an example, a relatively large or small (very large absolute value) t-test value indicates a significant difference between the average blood pressure of doctors and patients using the new drug and that of normal people. In this case, the effect of the new drug may be quite different from the expectation. A t-test value close to zero indicates a very small difference between the average blood pressure of doctors and patients using the new drug and that of normal people, which may be caused by random error. In this case, it can be considered that the effect of the new drug may meet the expectation.
[0105] In another embodiment, the global statistical data includes the mean of the global paired sample differences and the standard deviation of the global paired sample differences.
[0106] The paired samples can be the data of two times for the same object. Taking the sample data as the blood pressure data of doctors and patients as an example, the paired samples can represent the blood pressure data of the same doctor and patient in two states.
[0107] The above step 330 may include: calculating the sum of the global paired sample differences according to the local sum in the local statistical data; obtaining the mean of the global paired sample differences according to the global quantity and the sum of the global paired sample differences; calculating the global sum of squares according to the local sum of squares in the local statistical data; obtaining the standard deviation of the global paired sample differences according to the sum of the global paired sample differences, the mean of the global paired sample differences, the global sum of squares, and the global quantity.
[0108] The above step 342 may include: substituting the mean of the global paired sample differences and the standard deviation of the global paired sample differences into the t-distribution to calculate the test value.
[0109] The calculation of the above parameters will be described below through some formulas.
[0110] Exemplarily, substituting the mean of the global paired sample differences and the standard deviation of the global paired sample differences into the t-distribution, the calculated test value is determined by the following formula:
[0111] ;
[0112] where represents the mean of the global paired sample differences; represents the standard deviation of the global paired sample differences; D represents the global quantity.
[0113] In this embodiment, each participating party device can first perform calculation processing on the sample data it contains, count the local quantity of the sample data it contains, calculate the sum of local paired sample differences, and the sum of squares of local paired sample differences, and transmit these three pieces of information to the central server. Taking the participating party as an example, the sum of its local paired sample differences can be achieved in the following way:
[0114] ;
[0115] The difference between the two paired samples and is: is: .
[0116] The sum of squares of local paired sample differences can be achieved in the following way:
[0117] .
[0118] The central server collects the local quantity, the sum of local paired sample differences, and the sum of squares of local paired sample differences uploaded by each participating party. Among them, the global quantity is achieved in the following way:
[0119] ;
[0120] where m represents the number of participating parties.
[0121] The global paired sample difference is achieved in the following way:
[0122] ;
[0123] The sum of squares of global paired sample differences is achieved in the following way:
[0124] ;
[0125] The mean value of the global paired sample difference can be achieved in the following way:
[0126] ;
[0127] The calculation process of
[0128] is as follows:
[0129] In the above embodiments, paired samples can be tested. For example, when testing a new drug, it is necessary to check whether there are significant differences in the blood pressure changes of doctors and patients using the new drug before and after use. In this example, the paired samples can be the blood pressure of doctors and patients before using the new drug and the blood pressure of doctors and patients after using the new drug. Each participating party can be multiple hospitals participating in the test of the new drug. The participating party is the j-th hospital participating in the new drug test, and is the blood pressure sequence of the patients using the new drug in the j-th hospital.
[0130] As shown by the t-test value calculated above, a larger or smaller t-test value (a very large absolute value) indicates a significant difference between the mean blood pressure of doctors and patients before using the new drug and the mean blood pressure of doctors and patients after using the new drug. In this case, the new drug may have a therapeutic effect on doctors and patients. A t-test value close to zero indicates that the difference between the mean blood pressure of doctors and patients before using the new drug and the mean blood pressure of doctors and patients after using the new drug is very small, which may be caused by random errors. In this case, it can be considered that the improvement effect of the new drug on doctors and patients is small.
[0131] The above embodiments are all based on the processing methods of single samples or paired samples. In some implementation modes, the samples provided by the participating parties can also be two relatively independent types of samples.
[0132] The intermediate sample data includes first intermediate sample data and second intermediate sample data.
[0133] The first intermediate sample data includes first local statistical data and first local quantity obtained based on the first sample data included in the participating party's device. The first sample data is the user data of the user in the first state.
[0134] The second intermediate sample data includes second local statistical data and second local quantity obtained based on the second sample data included in the participating party's device. The second sample data is the user data of the user in the second state.
[0135] For different application scenarios, the information represented by the above first state and second state is different. Still taking the medical field as an example, to verify a new substance, the function of this new drug is to reduce high blood pressure, and it is hoped to compare the effects of this new drug (Drug A) with the existing standard treatment drug for reducing high blood pressure (Drug B). In this example, the first state can be the state of doctors and patients who have taken the new drug Drug A, and the first sample data is the body parameters of doctors and patients who have taken the new drug Drug A; the first state can be the state of doctors and patients who have taken the standard treatment drug B, and the second sample data can be the body parameters of doctors and patients who have taken the standard treatment drug B.
[0136] The global quantities include a first global quantity and a second global quantity; the global statistical data include a first global statistical data and a second global statistical data.
[0137] The above step 320 may include: calculating the first global quantity according to the first local quantities of the participating devices, and calculating the second global quantity according to the second local quantities of the participating devices.
[0138] Exemplarily, the first global quantity may be the sum of all the first local quantities, and the second global quantity may be the sum of all the second local quantities.
[0139] The above step 330 may include: calculating the first global statistical data according to the first global quantity and the first local statistical data of the participating devices, and calculating the second global statistical data according to the second global quantity and the second local statistical data of the participating devices.
[0140] The above step 342 may include: calculating a test value according to the first global statistical data, the second global statistical data, the first global quantity, the second global quantity, and the t-distribution.
[0141] Exemplarily, the first global statistical data may include a first global sample mean and a first global sample variance; the second global statistical data may include a second global sample mean and a second global sample variance.
[0142] The above calculating a test value according to the first global statistical data, the second global statistical data, the first global quantity, the second global quantity, and the t-distribution may include: substituting the first global sample mean, the first global sample variance, the second global sample mean, the second global sample variance, the first global quantity, and the second global quantity into the t-distribution to calculate the test value.
[0143] When using a common variance or different variances for two types of samples, different calculation logics can be adopted, which will be introduced separately in the following two embodiments.
[0144] In one embodiment, the two types of samples may use a common variance, that is, the first global sample variance and the second global sample variance take the same value.
[0145] The above substituting the first global sample mean, the first global sample variance, the second global sample mean, the second global sample variance, the first global quantity, and the second global quantity into the t-distribution to calculate the test value is implemented by the following formula:
[0146] ;
[0147] Wherein, represents the first global sample mean; represents the second global sample mean; represents the common variance; represents the first global quantity; represents the second global quantity, where the first global sample variance is equal to the second global sample variance and both are equal to the common variance.
[0148] For the two types of samples, the participating device also needs to calculate intermediate sample data for the two types of samples. Exemplarily, the intermediate sample data can also include information such as local sum, local sum of squares, local quantity, etc. Among them, for the two types of samples, it is necessary to calculate the local sum, local sum of squares, local quantity, etc. of the first type of samples respectively; calculate the local sum, local sum of squares, local quantity, etc. of the second type of samples.
[0149] Among them, the calculation methods for the local sum, local sum of squares, and local quantity can be similar to those for calculating the local sum, local sum of squares, and local quantity of a single sample in the previous implementation. The difference is that for two types of samples, it is necessary to calculate the local sum, local sum of squares, and local quantity for each type of sample respectively.
[0150] Taking the participating party as an example, the information sent by the participating device to the central server includes: the first local sample quantity and the second local sample quantity of the two types of samples and , the first local sum and the second local sum of the two types of samples and , the first local sum of squares and the second local sum of squares of the two types of samples and . Among them, and 's calculation method can refer to the calculation method of the local quantity in the previous embodiment ; and 's calculation can refer to the calculation method of the local sum in the previous embodiment ; and 's calculation can refer to the calculation method of the local sum of squares in the previous embodiment .
[0151] The central server can receive the first local sample quantity and the second local sample quantity, the first local sum and the second local sum, and the first local sum of squares and the second local sum of squares of the two types of samples sent by each participating device.
[0152] First, based on the received data, the first global sample mean and the second global sample mean can be calculated:
[0153] ; ;
[0154] Calculate the common variance :
[0155] .
[0156] In this embodiment, the degrees of freedom of the t-distribution are .
[0157] In another embodiment, the two types of samples can use different variances, that is, the values of the first global sample variance and the second global sample variance are not necessarily the same.
[0158] The above first global sample mean, first global sample variance, second global sample mean, second global sample variance, first global quantity, and second global quantity are substituted into the t-distribution to calculate the test value, which is achieved through the following formula:
[0159] ;
[0160] where represents the first global sample mean; represents the second global sample mean; represents the first global quantity; represents the second global quantity; represents the first global sample variance; represents the second global sample variance.
[0161] In this embodiment, except for the first global sample variance and the second global sample variance, the calculation methods of other parameters are the same as those in the previous embodiment.
[0162] The first global sample variance and the second global sample variance and are the variances of the two types of samples.
[0163] Among them, the first global sample variance is achieved through the following calculation method:
[0164] ;
[0165] The second global sample variance is achieved through the following method:
[0166] ;
[0167] This test value t, under the null hypothesis: - = 0 is true, follows a degrees of freedom of:
[0168] 。
[0169] Since the central server cannot directly obtain the original sample data, the first global sample variance and the second global sample variance can be calculated based on the received first local sample quantity, second local sample quantity, first local sum, second local sum, first local sum of squares, and second local sum of squares.
[0170] Exemplarily, it can be calculated through the following formula:
[0171] The first global sample variance can be achieved through the following formula:
[0172] ;
[0173] The second global sample variance can be achieved through the following formula:
[0174] 。
[0175] Then the t-test value is:
[0176]
[0177] The degrees of freedom are:
[0178] 。
[0179] Subsequently, the p-value can be calculated based on the degrees of freedom and the t-test value, which will not be elaborated here.
[0180] In the t-test for two independent samples, the calculation of the t-test value is to evaluate whether there is a significant difference between the means of the two independent samples. The magnitude and sign of the t-test value can be used to judge the existence and direction of this difference.
[0181] A relatively large or small t-test value (large absolute value) indicates a significant difference between the two means. In this case, the object being tested may be significantly different from the expectation, and it is considered that there is indeed a significant difference between the two types of sample data. If the t-test value is positive, it indicates that the mean of the first group is significantly higher than the mean of the second group. If the t-test value is negative, it indicates that the mean of the first group is significantly lower than the mean of the second group.
[0182] A t-test value close to zero indicates that the difference between the two means is very small and may be caused by random error. In this case, it can be considered that the object being tested may meet the expectation, and it is considered that there is no significant difference between the means of the two types of sample data.
[0183] Specifically, in an actual application scenario, such as when comparing the therapeutic effects of two drugs, if the t-test value is large, it indicates that the efficacy of one drug is significantly higher than that of the other; if the t-test value is close to zero, it indicates that the effects of the two drugs are similar.
[0184] The data verification method provided by the embodiments of the present application completes the calculation of the t-test through secure distributed computing on the premise of protecting the data privacy of all parties, avoiding the sharing of original sample data, solving the data privacy protection and security problems existing in traditional t-test methods, and realizing the t-test while ensuring the security of the original sample data.
[0185] Furthermore, the embodiments of the present application provide t-tests for single samples, two independent samples, and paired samples, which can better adapt to the verification of user data in more scenarios.
[0186] In addition, the embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the data verification method described in the above method embodiments.
[0187] The computer program product of the data verification method provided by the embodiments of the present application includes a computer-readable storage medium storing program codes. The instructions included in the program codes can be used to execute the steps of the data verification method described in the above method embodiments. For details, please refer to the above method embodiments and will not be elaborated here.
[0188] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and the module, program segment, or part of the code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0189] In addition, in each embodiment of the present application, each functional module can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.
[0190] If the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes. It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprises", "comprising", or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, the elements defined by the statement "comprising..." do not exclude the existence of additional identical elements in the process, method, article, or device comprising the said elements.
[0191] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application. It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0192] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A data verification method, characterized in that, Including: Receiving intermediate sample data sent by each participating party device, where the intermediate sample data includes local statistical data and local quantity obtained based on the sample data included in the participating party device, the sample data includes user data, the user data includes the user's physical parameters, and the physical parameters include blood pressure; Calculate the global quantity according to the local quantities of the devices of each participating party, including: , where D represents the global quantity, is the blood pressure sequence of the patients using the new drug in the j-th hospital, also represents the local quantity; Calculating global statistical data according to the global quantity and the local statistical data of each participating party device; where the global statistical data includes the global sample mean and the global sample standard deviation; Constructing a t-distribution that satisfies the degrees of freedom determined by the global quantity based on the global quantity; Calculating a test value based on the global statistical data and the t-distribution, including substituting a set sample pseudo-mean, the global sample mean, and the global sample standard deviation into the t-distribution to calculate the test value; the set sample pseudo-mean includes a value determined based on the blood pressure standard of normal users; Wherein, the test value is used to determine whether the user data meets the expectation, and the test value includes verifying whether the user's physical state after taking the medicine meets the expectation; the larger the absolute value of the test value, the greater the difference between the average blood pressure of the users using the new drug and the blood pressure value of normal people; the smaller the absolute value of the test value, the smaller the difference between the average blood pressure of the users using the new drug and the blood pressure value of normal people.
2. The method according to claim 1, characterized in that The calculation of the test value by substituting the set sample pseudo-mean, the global sample mean, and the global sample standard deviation into the t-distribution is determined by the following formula: ; Among them, represents the global sample mean; represents the global sample standard deviation; D represents the global quantity; represents the set sample quasi-mean.
3. The method according to claim 1, wherein The global statistical data includes the mean of the global paired sample differences and the standard deviation of the global paired sample differences; The calculation of the test value based on the global statistical data and the t-distribution includes: Substituting the mean of the global paired sample differences and the standard deviation of the global paired sample differences into the t-distribution to calculate the test value; where the paired samples include the blood pressure of the user before using the new drug and the blood pressure of the user after using the new drug; The smaller the absolute value of the test value, the smaller the difference between the mean blood pressure of the user before using the new drug and the mean blood pressure of the user after using the new drug.
4. The method according to claim 3, characterized in that, The calculation of the test value by substituting the mean of the global paired sample differences and the standard deviation of the global paired sample differences into the t-distribution is determined by the following formula: ; Among them, represents the mean of the global paired sample differences; represents the standard deviation of the global paired sample differences; D represents the global quantity.
5. The method according to any one of claims 2 to 4, characterized in that, The global statistical data includes the global mean and the global standard deviation; the local statistical data includes the local sum of squares and the local sum; The calculation of the global statistical data according to the global quantity and the local statistical data of each participating party device includes: Calculating the global sum according to the local sum in the local statistical data, where the global sum is the global sample sum or the sum of the global paired sample differences; Obtaining the global mean according to the global quantity and the global sum, where the global mean is the global sample mean or the mean of the global paired sample differences; Calculating the global sum of squares according to the local sum of squares in the local statistical data; A global standard deviation is obtained based on the global sum, the global mean, the global sum of squares, and the global quantity, where the global standard deviation includes a global sample standard deviation or a standard deviation of global paired sample differences.
6. The method according to claim 1, characterized in that, The intermediate sample data includes first intermediate sample data and second intermediate sample data; the first intermediate sample data includes first local statistical data and a first local quantity obtained based on first sample data included in the participating party device, where the first sample data is user data in a first state; the second intermediate sample data includes second local statistical data and a second local quantity obtained based on second sample data included in the participating party device, where the second sample data is user data in a second state; the first state represents a state where the user uses a first new drug, and the first sample data includes physical parameters of the user using the first new drug; the second state represents a state where the user uses a second new drug, and the second sample data includes physical parameters of the user using the second new drug; The global quantity includes a first global quantity and a second global quantity; the global statistical data includes first global statistical data and second global statistical data; The calculating of the global quantity according to the local quantities of the participating party devices includes: calculating the first global quantity according to the first local quantities of the participating party devices, and calculating the second global quantity according to the second local quantities of the participating party devices; The calculating of the global statistical data based on the global quantity and the local statistical data of the participating party devices includes: calculating the first global statistical data according to the first global quantity and the first local statistical data of the participating party devices, and calculating the second global statistical data according to the second global quantity and the second local statistical data of the participating party devices; The calculating of the test value based on the global statistical data and the t-distribution includes: Calculating a test value according to the first global statistical data, the second global statistical data, the first global quantity, the second global quantity, and the t-distribution; where the test value characterizes the difference between the means of two independent samples, and the smaller the absolute value of the test value, the smaller the difference between the first new drug and the second new drug.
7. The method according to claim 6, characterized in that The first global statistical data includes a first global sample mean and a first global sample variance; the second global statistical data includes a second global sample mean and a second global sample variance; The calculating of the test value according to the first global statistical data, the second global statistical data, the first global quantity, the second global quantity, and the t-distribution includes: Substituting the first global sample mean, the first global sample variance, the second global sample mean, the second global sample variance, the first global quantity, and the second global quantity into the t-distribution to calculate the test value.
8. The method according to claim 7, wherein Substituting the first global sample mean, the first global sample variance, the second global sample mean, the second global sample variance, the first global quantity, and the second global quantity into the t-distribution to calculate the test value is achieved through the following formula: ; Wherein, represents the first global sample mean; represents the second global sample mean; represents the common variance; represents the first global quantity; represents the second global quantity; the first global sample variance is equal to the second global sample variance and is equal to the common variance.
9. The method according to claim 7, wherein Substituting the first global sample mean, the first global sample variance, the second global sample mean, the second global sample variance, the first global quantity, and the second global quantity into the t-distribution to calculate the test value is achieved through the following formula: ; Among them, represents the first global sample mean; represents the second global sample mean; represents the first global quantity; represents the second global quantity; represents the first global sample variance; represents the second global sample variance.
10. An electronic device, characterized in that, Including: A processor and a memory, where the memory stores machine-readable instructions executable by the processor. When the electronic device runs, the machine-readable instructions are executed by the processor to perform the steps of the method according to any one of claims 1 to 9.
11. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by the processor, it performs the steps of the method according to any one of claims 1 to 9.
12. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by the processor, it implements the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Cross-sample feature selection method in federated learning system and federated learning system
CN113537361A
Prediction model determination method, S-wave travel time curve prediction method and related equipment
CN115545356A