Apparatus and method for providing statistical data

KR102998547B1Active Publication Date: 2026-08-03POSTECH ACADEMY INDUSTRY FOUNDATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
KR1020240143784
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-10-21
Publication Date
2026-08-03
Estimated Expiration
2044-10-21

Smart Images

  • Figure 112024114165413-PAT00008_ABST
    Figure 112024114165413-PAT00008_ABST
Patent Text Reader

Abstract

A statistical data providing device according to the present invention includes a memory storing a data providing program that provides a response value corresponding to a query; and a processor that executes the data providing program, wherein the data providing program receives a query, generates a raw response value corresponding to the query, calculates the sparsity of the raw data from the raw response value, optimizes a noise value for the raw data based on the sparsity and optimization conditions, and generates a query response value by applying an optimized noise value corresponding to the sparsity to the raw data.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to an apparatus and method for providing statistical data for a query. Specifically, the present invention relates to an apparatus and method for providing statistical data that provides a query response value to which differential privacy is applied based on the scarcity of raw data. Background Technology

[0003] The content described in this section merely provides background information regarding the present embodiment and does not constitute prior art.

[0004] Differential Privacy defines the level of data security as a precise value and is receiving significant attention today as the most practical data security technology used in artificial intelligence, given the active use of data in AI training and big data analysis.

[0005] The guarantee of differential personal information protection is achieved by adding a predetermined noise value to raw data based on a pre-set security level, whereas conventional methods generate query response values ​​by applying the same noise value extracted from a probability distribution to all raw data.

[0006] However, the level of security guaranteed by default varies depending on the sparsity within the dataset for each data value, and conventional methods consider only the worst-case scenario, adding excessive noise values ​​to some data, thereby reducing the usability of the data.

[0007] Here, the noise value is a value added to or subtracted from the raw data, which can partially alter the raw data. The magnitude of the noise value may vary depending on the preset security level. If the preset security level is relatively high, a relatively larger noise value may be applied to the raw data.

[0008] In other words, if noise values ​​based on a pre-set security level are applied to all data, highly sparsity data can be effectively protected, but less sparsity data will differ significantly from its original value and fail to reflect the statistical characteristics of the raw dataset.

[0009] FIG. 1 shows the statistical characteristics of a raw response value (1) and the statistical characteristics of a query response value (2) to which conventional differential personal information protection is applied. As shown in FIG. 1, the query response value (2) to which conventional differential personal information protection is applied may show a large gap (3) with the statistical characteristics of the raw response value (1).

[0010] This occurs because the same noise value is applied to all data, and the same noise value is applied to data placed in the first section (4), which has relatively low scarcity, and the second section (5) or third section (6), which has relatively high scarcity. When the same noise value is applied, the data in the second section (5) or third section (6) can be protected, but the values ​​of the data in the first section (4) are significantly distorted and fail to reflect the statistical characteristics of the low response value (1).

[0011] Therefore, technology is required to solve such problems. Prior art literature

[0013] Korean Patent Publication No. 10-2024-0051538 (Title of Invention: Method and Apparatus for Generating Task-Adaptive Differential Privacy for Privacy-Preserving Machine Learning) The problem to be solved

[0014] The objective of the present invention is to provide a differential personal information protection technology that applies noise values ​​differentially according to the sparsity of raw data in a raw dataset.

[0015] In addition, the objective of the present invention is to provide query data that satisfies a pre-set security level by differentially applying noise values ​​according to the sparsity of raw data and reflects the statistical characteristics of the raw dataset.

[0016] The objects of the present invention are not limited to those mentioned above, and other unmentioned objects and advantages of the present invention may be understood from the following description and will be more clearly understood by the embodiments of the present invention. Furthermore, it will be readily apparent that the objects and advantages of the present invention can be realized by the means and combinations thereof set forth in the claims. means of solving the problem

[0018] In one embodiment of the present invention, a statistical data providing device comprises: a memory storing a data providing program that provides a response value corresponding to a query; and a processor that executes the data providing program, wherein the data providing program receives a query, generates a raw response value corresponding to the query, calculates the sparsity of the raw data from the raw response value, optimizes a noise value for the raw data based on the sparsity and optimization conditions, and generates a query response value by applying the optimized noise value corresponding to the sparsity to the raw data.

[0019] In addition, the above optimization condition may be such that the security level of the query response value satisfies a preset security level, and the difference between the statistical distribution of the query response value and the statistical distribution of the row response value is minimized.

[0020] In addition, the data providing program generates the query response value by matching a noise value optimized for the raw data to satisfy the optimization conditions based on the scarcity and noise distribution of the raw data, and the noise distribution may be generated based on the sensitivity of the query and the preset security level.

[0021] In addition, the data providing program can generate a first temporary response value by applying a first noise value to the raw data based on the scarcity of the raw data and the probability of a noise value occurring in the noise distribution, and calculate an optimization score of the first temporary response value for the optimization condition.

[0022] In addition, the data providing program may generate a second temporary response value by applying either the first noise value or a second noise value that is larger or smaller than the first noise value to the raw data, and calculate an optimization score of the second temporary response value for the optimization condition.

[0023] In addition, the data providing program can set the temporary response value corresponding to the largest optimization score as the query response value.

[0024] A method for providing statistical data according to an embodiment of the present invention may include: receiving a query; generating a raw response value corresponding to the query; calculating the sparsity of the raw data in the raw response value; matching an optimized noise value for the sparsity to satisfy an optimization condition; and applying the optimized noise value corresponding to the sparsity to the raw data to generate a query response value.

[0025] In addition, the above optimization condition may be such that the security level of the query response value satisfies a preset security level, and the difference between the statistical distribution of the query response value and the statistical distribution of the row response value is minimized.

[0026] In addition, the step of matching the optimized noise value may match the noise value optimized for the raw data to satisfy the optimization condition based on the sparsity and noise distribution of the raw data, and the noise distribution may be generated based on the sensitivity of the query and the preset security level.

[0027] In addition, the step of matching the optimized noise value may generate a first temporary response value by applying a first noise value to the raw data based on the sparsity of the raw data and the probability of a noise value occurring in the noise distribution, and calculate an optimization score of the first temporary response value for the optimization condition.

[0028] Additionally, the step of matching the optimized noise value may generate a second temporary response value by applying either the first noise value or a second noise value that is larger or smaller than the first noise value to the raw data, and calculate an optimization score of the second temporary response value for the optimization condition.

[0029] In addition, the step of generating the above query response value may set the temporary response value corresponding to the largest optimization score as the above query response value. Effects of the invention

[0031] The statistical data providing apparatus and method of the present invention generate query data by differentially applying noise values ​​according to the sparsity of raw data in a raw dataset, thereby providing query data that preserves the statistical characteristics of the raw dataset.

[0032] In addition to the above, the specific effects of the present invention are described together with the specific details for implementing the invention below. Brief explanation of the drawing

[0034] Figure 1 is an example diagram illustrating conventional differential personal information protection technology. FIG. 2 is a conceptual diagram schematically illustrating a statistical data providing device according to an embodiment of the present invention. Figure 3 is a block diagram schematically showing the configuration of the statistical data providing device illustrated in Figure 2. FIGS. 4 to 8 are illustrative diagrams for explaining the operation of a statistical data providing device according to an embodiment of the present invention. FIG. 9 is a flowchart illustrating a method for providing statistical data according to an embodiment of the present invention. Specific details for implementing the invention

[0035] Terms and words used in this specification and claims shall not be interpreted as being limited to their general or dictionary meanings. In accordance with the principle that an inventor may define the concept of a term or word to best describe their invention, they shall be interpreted in a meaning and concept consistent with the technical spirit of the invention. Furthermore, since the embodiments described in this specification and the configurations illustrated in the drawings are merely one embodiment of the invention and do not represent the entire technical spirit of the invention, it should be understood that various equivalents, modifications, and applicable examples capable of replacing them may exist at the time of filing this application.

[0036] The terms first, second, A, B, etc., as used in this specification and claims may be used to describe various components, but said components should not be limited by said terms. These terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.

[0037] The terms used in this specification and claims are used merely to describe specific embodiments and are not intended to limit the invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "having" should be understood as not precluding the existence or addition of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification.

[0038] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which this invention pertains.

[0039] Terms such as those defined in commonly used dictionaries should be interpreted as having meanings consistent with their meanings in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this application.

[0040] In addition, each component, process, procedure, or method included in each embodiment of the present invention may be shared within a scope that is not technically contradictory to one another.

[0041] Hereinafter, with reference to FIGS. 2 to 9, a statistical data providing apparatus and method according to embodiments of the present invention will be described in detail.

[0043] First, a statistical data providing device will be described with reference to FIGS. 2 to 8.

[0044] FIG. 2 is a conceptual diagram schematically showing a statistical data providing device according to an embodiment of the present invention, FIG. 3 is a block diagram schematically showing the configuration of the statistical data providing device shown in FIG. 2, and FIG. 4 to 8 are illustrative diagrams for explaining the operation of a statistical data providing device according to an embodiment of the present invention.

[0045] Referring to FIGS. 2 and 3, a statistical data providing device (100) receives a query requesting data having statistical characteristics and can provide a query response value (200) for the query based on a raw dataset. To schematically describe the operation of the statistical data providing device (100) providing the query response value (200), when the statistical data providing device (100) receives a query, it can generate an unprocessed raw response value using multiple raw data included in the raw dataset, calculate the scarcity of each raw data in the raw response value, and generate the query response value (200) by applying different noise values ​​to the raw data according to the scarcity. That is, the query response value (200) can be generated and provided through differential privacy protection, which applies noise values ​​differentially according to the scarcity of the raw response value in the raw dataset.

[0046] To perform such an operation, the data providing device (100) may include a memory (110) and a processor (120).

[0047] Memory (110) may store a data providing program that generates a query response value by differentially applying a noise value to each raw data based on the sparsity of each raw data in the raw response value. Memory (110) may be interpreted as a general term for a non-volatile storage device that continues to maintain stored information even when power is not supplied, and a volatile storage device that requires power to maintain stored information. Additionally, memory (110) may perform the function of temporarily or permanently storing data processed by the processor (120). Memory (110) may include non-volatile storage devices such as magnetic storage media or flash storage media in addition to volatile storage devices that require power to maintain stored information, but the scope of the present invention is not limited thereto.

[0048] The processor (120) can execute a data provision program stored in memory (110) to receive a query, generate a raw response value corresponding to the query, calculate sparsity for each raw data in the raw response value, match an optimized noise value for each raw data based on sparsity and optimization conditions, and apply the optimized noise value according to the sparsity to the raw data to generate a query response value.

[0049] Here, sparsity is a numerical representation of how common or rare the corresponding raw data is among the raw response values. If a specific raw data value is relatively uncommon among the raw response values, it can be measured as having relatively high sparsity. Accordingly, since high sparsity implies that the raw data is uncommon and thus has a higher probability of being inferred, a relatively larger noise value can be applied to protect the raw data. For reference, the noise value is a value added to or subtracted from the raw data.

[0050] In addition, the optimization condition is to ensure that the security level of the query response value satisfies a preset security level and that the difference between the statistical distribution of the query response value and the statistical distribution of the row response value is minimized. The data provider program can generate query response values ​​by matching an optimized noise value according to the sparsity of each row data to satisfy these optimization conditions, and by applying the optimized noise to each row data.

[0051] Next, the operation of the data providing program will be explained in detail with reference to FIGS. 4 to 8.

[0052] Referring to FIG. 4, an example is described in which a request for a height distribution for a raw dataset (10) is received via a query. The dataset (10) contains height data for 1,000 people. When a data provider receives such a query, it can generate a raw response value (20) for the query using a number of raw data (11) key values, and calculate the sparsity for each raw data (11) in the raw response value (20).

[0053] Here, scarcity is a numerical representation of how common or rare the corresponding raw data is in the raw response value (10). In the raw response value (20), the second raw data (11-2) may be scarcer than the first raw data (11-1). The third raw data (11-3) may be scarcer than the second raw data (11-2), and the 1000th raw data (11-1000) may be scarcer than the first raw data (11-1). Scarcity may be calculated based on frequency, but is not limited thereto, and any criterion capable of representing scarcity may be applied.

[0054] Additionally, the data provider program can generate a query response value by matching an optimized noise value for each raw data (11) that satisfies the optimization condition based on the sparsity of each raw data (11) and the noise distribution for the query. The optimization condition is to minimize the difference between the statistical distribution of the query response value and the statistical distribution of the raw response value while the security level of the query response value satisfies a preset security level.

[0055] Here, the security level refers to the strength of privacy protection and can determine the magnitude of the noise value added to the raw data to protect the raw data. The higher the security level, the larger the noise value applied to the raw data, which can reduce the possibility of the actual raw data being leaked or identified.

[0056] Before describing the operation of matching the optimized noise value for each raw data (11) that satisfies the optimization condition, the noise distribution is described.

[0057] A noise distribution for a query can be generated based on the query's sensitivity and a preset security level. Query sensitivity is a value that measures how much of a change occurs in the row response value (20) for a query when a single row data (11) is added or deleted, and is a value that combines the sensitivity for each row data (11).

[0058] Referring to FIG. 4, for example, the difference between the raw response value of the dataset excluding the first raw data (11-1) and the raw response value (20) of the original dataset (10) is measured, the difference between the raw response value of the dataset excluding the second raw data (11-2) and the raw response value (20) of the original dataset (10) is measured, and the same operation is performed up to the 1000th raw data (11-1000). Then, the sensitivity of the query can be calculated by combining the difference between the raw response value of the dataset excluding each raw data (11) and the raw response value (20) of the original dataset (10).

[0059] For queries with relatively high sensitivity, since each row data (11) has a greater impact on the row response value (20), it is necessary to protect the row data (11) by adding a relatively larger noise value. Accordingly, different noise distributions may be generated for each query.

[0060] Representative examples of noise distributions include the Laplace distribution and the Gaussian distribution, but they are not limited to these. Figure 5 is an example diagram illustrating the Laplace distribution. Figure 5 shows the Laplace distribution (L) according to changes in query sensitivity when the preset security level (ε) is 1. The x-axis represents the noise value, and the y-axis represents the probability of each noise value being selected. In the Laplace distribution (L), when the security level (ε) is fixed, the smaller the query sensitivity, the relatively smaller the noise value is likely to be selected. Looking at the Laplace distribution with a sensitivity of 1 (L-1), the distribution becomes narrower as the noise value approaches 0, increasing the probability that smaller noise values ​​will be selected. On the other hand, looking at the Laplace distribution with a sensitivity of 3 (L-3), the distribution is wider than that of the Laplace distribution with a sensitivity of 1 (L-1), increasing the probability that relatively larger noise values ​​will be selected.

[0061] For reference, a smaller value of the security level (ε) indicates a relatively higher level of security. In other words, a security level (ε) value of 1 means that the security level is higher than a value of 2.

[0062] Based on the noise distribution (L) of the query generated in this way and the sparsity of each raw data (11), the data providing program can generate a query response value that satisfies the optimization condition by matching a noise value optimized for each raw data (11).

[0063] Here, the optimization condition can be expressed by mathematical formula 1.

[0064] [Mathematical Formula 1]

[0065]

[0066] In mathematical formula 1 represents the security level of the query response value, and ε represents a preset security level, represents the minimum statistical distribution difference (U) between the row response value (D) and the query response value (M(D)). That is, Equation 1 is to find a noise value for each row data such that the security level of the query response value (M(D)) is greater than or equal to a preset security level (ε), and the statistical distribution difference (U) with respect to the row response value (D) is minimized.

[0067] Here, the difference in statistical distribution (U) (or difference in utility values) can be a criterion for evaluating the accuracy or usefulness of the query response value. A small difference in statistical distribution (U) from the row response value may indicate that the accuracy of the query response value is relatively high. The difference in statistical distribution (U) can be measured using methods such as absolute error, relative error, mean squared error (MSE), and bias.

[0068] Next, referring to FIGS. 6 to 8, an operation to generate a query response value by applying a noise value optimized for each raw data (11) based on the scarcity of the raw data (11) is described. Referring to FIG. 6, the data providing program can generate a first temporary data by applying a first noise value to each raw data (11) based on the scarcity of each raw data (11) and the probability that each noise value occurs in the noise distribution (L) of the query. The first noise value may be applied differentially according to the scarcity of each raw data (11), and the first temporary data may be a value obtained by adding or subtracting the first noise value from the raw data (11).

[0069] For example, if 3 is applied as the first noise value for the first raw data (11-1) in FIG. 4, 2 may be applied as the first noise value for the second raw data (11-2) which is less rare than the first raw data (11-1), 1 may be applied as the first noise value for the third raw data (11-3) which is less rare than the second raw data (11-2), and 3.5 may be applied as the first noise value for the 1000 raw data (11-1000) which is more rare than the first raw data (11-1).

[0070] And, the data provider program can generate a first temporary response value (30-1) using the first temporary data for each raw data (11) and calculate an optimization score for the first temporary response value (30-1) for the optimization condition. At this time, if the security level of the first temporary response value (30-1) does not satisfy a pre-set security level, the data provider program may not calculate the optimization score.

[0071] The data providing program generates second temporary data by applying either a first noise value or a second noise value that is larger or smaller than the first noise value to each raw data (11). Then, based on the second temporary data, a second temporary response value (30-2) is generated as shown in FIG. 7, and the optimization score of the second temporary response value (30-2) for the optimization condition can be calculated.

[0072] At this time, the optimization score of the temporary response value (30) may be the difference in statistical distribution (40) between the raw response value (10) and the temporary response value (30), and the smaller the difference, the relatively higher the optimization score may be. Since the difference in statistical distribution (40-1) between the first temporary response value (30-1) and the raw response value (10) shown in FIG. 6 is greater than the difference in statistical distribution (40-2) between the second temporary response value (30-2) and the raw response value (10) shown in FIG. 7, the optimization score of the second temporary response value (30-2) may be greater than that of the first temporary response value (30-1). In FIG. 6 and FIG. 7, the difference in statistical distribution (40) was calculated as a gap for the center value, but it is not limited thereto.

[0073] The data providing program may repeat the operation of modifying the noise value for each raw data a predetermined number of times, and provide the temporary response value (30) having the largest optimization score as the query response value (50) as shown in FIG. 8. That is, for each raw data (11), the noise value of the previous stage is maintained, or the noise value that is added or subtracted by a predetermined amount from the noise value of the previous stage is repeated to find the temporary response value (30) having the largest optimization score and set it as the query response value (50).

[0074] In this way, the data provider program can generate query response values ​​by applying different noise values ​​to each raw data based on the sparsity of the raw data in the query response values. That is, query response values ​​can be generated by applying relatively large noise values ​​to raw data with relatively high sparsity and relatively small noise values ​​to raw data with relatively low sparsity. The query response values ​​generated in this way can protect each raw data according to a pre-set security level while maintaining the statistical distribution of the raw response values, thereby increasing the usability of the raw dataset.

[0075] Additionally, the communication module (130) may include a device comprising hardware and software required to transmit and receive signals, such as control signals or data signals, through a wired or wireless connection with another network device in order to perform data communication regarding signal data with an external device. The database (140) may store various data for the operation of a data provision program.

[0077] FIG. 9 is a flowchart illustrating a method for providing statistical data according to an embodiment of the present invention.

[0078] Referring to FIGS. 2 and FIG. 9, a statistical data provision method using a statistical data provision device (100) is described. The statistical data provision method (S100) receives a query (step S110) and generates a raw response value corresponding to the query (step S120). Then, the sparsity of the raw data is calculated from the raw response value (step S130), and a noise value optimized based on the sparsity is matched to each raw data to satisfy the optimization condition (step S140), and the query response value (200) is generated by applying the optimized noise value corresponding to the sparsity to the raw data (step S150).

[0079] Next, the process of the statistical data provision method (S100) will be explained in detail with reference to FIGS. 4 to 8.

[0080] The case in which a request for a key distribution for a raw dataset is received via a query is described as an example. Steps S110 to S130 are described with reference to FIG. 4. The raw dataset (10) contains key data for 1,000 people. When the statistical data providing device (100) receives a query, it generates a raw response value (20) for the query using a plurality of height values ​​of raw data (11), and can calculate the sparsity for each raw data (11) in the raw response value (20).

[0081] Here, scarcity is a numerical representation of how common or rare the corresponding raw data is in the raw response value (10). In the raw response value (20), the second raw data (11-2) may be scarcer than the first raw data (11-1). The third raw data (11-3) may be scarcer than the second raw data (11-2), and the 1000th raw data (11-1000) may be scarcer than the first raw data (11-1). Scarcity may be calculated based on frequency, but is not limited thereto, and any criterion capable of representing scarcity may be applied.

[0082] Next, with reference to FIGS. 4 to 8, the process of matching an optimized noise value to each raw data (step S140) and the process of generating a query response value by applying an optimized noise value corresponding to sparsity to the raw data (step S150) are described.

[0083] The statistical data providing device (100) can generate a query response value by matching an optimized noise value for each raw data (11) that satisfies an optimization condition based on the sparsity of each raw data (11) and the noise distribution (L) for the query. The optimization condition is to minimize the difference between the statistical distribution of the query response value and the statistical distribution of the raw response value while the security level of the query response value satisfies a preset security level.

[0084] A statistical data providing device (100) can generate a query response value that satisfies optimization conditions by matching a noise value optimized for each raw data (11) based on the noise distribution (L) of the query and the sparsity of each raw data (11).

[0085] Here, the optimization condition can be expressed by mathematical formula 1.

[0086] [Mathematical Formula 1]

[0087]

[0088] In mathematical formula 1 represents the security level of the query response value, and ε represents a preset security level, represents the minimum statistical distribution difference (U) between the row response value (D) and the query response value (M(D)). That is, Equation 1 is to find a noise value for each row data such that the security level of the query response value (M(D)) is greater than or equal to a preset security level (ε), and the statistical distribution difference (U) with respect to the row response value (D) is minimized.

[0089] Referring to FIG. 5, a statistical data providing device (100) can generate first temporary data by applying a first noise value to each raw data (11) based on the scarcity of each raw data (11) and the probability that each noise value occurs in the noise distribution (L) of the query. The first noise value may be applied differentially according to the scarcity of each raw data (11), and the first temporary data may be a value obtained by adding or subtracting the first noise value from the raw data (11).

[0090] For example, if 3 is applied as the first noise value for the first raw data (11-1) in FIG. 2, 2 may be applied as the first noise value for the second raw data (11-2) which is less rare than the first raw data (11-1), 1 may be applied as the first noise value for the third raw data (11-3) which is less rare than the second raw data (11-2), and 3.5 may be applied as the first noise value for the 1000 raw data (11-1000) which is more rare than the first raw data (11-1).

[0091] And, the statistical data providing device (100) can generate a first temporary response value (30-1) using the first temporary data for each raw data (11) and calculate an optimization score of the first temporary response value (30-1) for the optimization condition. At this time, the data providing program may not calculate the optimization score if the security level of the first temporary response value (30-1) does not satisfy a pre-set security level.

[0092] Referring to FIG. 6, the statistical data providing device (100) generates second temporary data by applying either a first noise value or a second noise value that is larger or smaller than the first noise value to each raw data (11). Then, based on the second temporary data, a second temporary response value (30-2) is generated, and an optimization score of the second temporary response value (30-2) for the optimization condition can be calculated.

[0093] At this time, the optimization score of the temporary response value (30) may be the difference in statistical distribution (40) between the raw response value (10) and the temporary response value (30), and the smaller the difference, the relatively higher the optimization score may be. Since the difference in statistical distribution (40-1) between the first temporary response value (30-1) and the raw response value (10) shown in FIG. 5 is greater than the difference in statistical distribution (40-2) between the second temporary response value (30-2) and the raw response value (10) shown in FIG. 6, the optimization score of the second temporary response value (30-2) may be greater than that of the first temporary response value (30-1). In FIG. 5 and FIG. 6, the difference in statistical distribution (40) was calculated as a gap for the center value, but it is not limited thereto.

[0094] The statistical data providing device (100) can generate a query response value (50) by repeating the operation of modifying the noise value for each raw data as a predetermined number of times and setting the temporary response value (30) having the largest optimization score as the query response value (50). That is, for each raw data (11), the operation of maintaining the noise value of the previous stage or applying a noise value that is added or subtracted by a predetermined amount from the noise value of the previous stage to the raw data (11) is repeated to find the temporary response value (30) having the largest optimization score and set it as the query response value (50).

[0096] The above description is merely an illustrative explanation of the technical concept of the present embodiment, and a person skilled in the art to which the present embodiment belongs would be able to make various modifications and variations within the scope of the essential characteristics of the present embodiment. Accordingly, the present embodiments are intended to explain, not limit, the technical concept of the present embodiment, and the scope of the technical concept of the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment shall be interpreted by the claims below, and all technical concepts within an equivalent scope shall be interpreted as being included within the scope of rights of the present embodiment.

Claims

Claim 1 A statistical data providing device comprising: a memory storing a data providing program that provides a response value corresponding to a query; and a processor that executes the data providing program, wherein the data providing program receives a query, generates a raw response value corresponding to the query, calculates the sparsity of the raw data from the raw response value, optimizes a noise value for the raw data based on the sparsity and optimization conditions, and generates a query response value by applying the optimized noise value corresponding to the sparsity to the raw data, and wherein the optimization conditions are such that the security level of the query response value satisfies a preset security level and the difference between the statistical distribution of the query response value and the statistical distribution of the raw response value is minimized. Claim 2 delete Claim 3 A statistical data providing device according to claim 1, wherein the data providing program generates the query response value by matching a noise value optimized for the raw data to satisfy the optimization condition based on the sparsity and noise distribution of the raw data, and the noise distribution is generated based on the sensitivity of the query and the preset security level. Claim 4 In paragraph 3, the data providing program generates a first temporary response value by applying a first noise value to the raw data based on the scarcity of the raw data and the probability of a noise value occurring in the noise distribution, and calculates an optimization score of the first temporary response value for the optimization condition, a statistical data providing device. Claim 5 In paragraph 4, the data providing program generates a second temporary response value by applying either the first noise value or a second noise value greater than or smaller than the first noise value to the raw data, and calculates an optimization score of the second temporary response value for the optimization condition, a statistical data providing device. Claim 6 In paragraph 5, the data providing program is a statistical data providing device that sets a temporary response value corresponding to the largest optimization score as the query response value. Claim 7 A method for providing statistical data, comprising: receiving a query; generating a raw response value corresponding to the query; calculating sparsity for raw data in the raw response value; matching an optimized noise value for the sparsity to satisfy optimization conditions; and applying the optimized noise value corresponding to the sparsity to the raw data to generate a query response value, wherein the optimization conditions are such that the security level of the query response value satisfies a preset security level, and the difference between the statistical distribution of the query response value and the statistical distribution of the raw response value is minimized. Claim 8 delete Claim 9 A method for providing statistical data according to claim 7, wherein the step of matching the optimized noise value matches the noise value optimized for the raw data to satisfy the optimization condition based on the sparsity and noise distribution of the raw data, and the noise distribution is generated based on the sensitivity of the query and the preset security level. Claim 10 In claim 9, the step of matching the optimized noise value comprises applying a first noise value to the raw data based on the scarcity of the raw data and the probability of a noise value occurring in the noise distribution to generate a first temporary response value, and calculating an optimization score of the first temporary response value for the optimization condition, a statistical data provision method. Claim 11 In claim 10, the step of matching the optimized noise value comprises applying either the first noise value or a second noise value greater than or smaller than the first noise value to the raw data to generate a second temporary response value, and calculating an optimization score of the second temporary response value for the optimization condition, a statistical data provision method. Claim 12 In claim 11, the step of generating the query response value is a statistical data provision method in which a temporary response value corresponding to the largest optimization score is set as the query response value.