A classification method for large-scale datasets based on privacy computing
By applying privacy-based computing methods in large-scale data set classification, a tag library and tag identification model are established, and the target data is initially classified and merged, which solves the problems of insufficient processing efficiency and accuracy and lack of flexibility in the existing technology, and realizes efficient and flexible data set classification.
Patent Information
- Application Number
- CN202411035778.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-07-31
AI Technical Summary
The existing privacy computing technology has difficulty meeting the processing efficiency and accuracy in large-scale data set classification, and lacks flexibility to adapt to the classification needs in different fields and scenarios.
A large-scale data set classification method based on privacy computing is proposed. By establishing a tag library and a tag identification model, the target data is initially classified, the user's classification needs are identified, the target category is determined based on the needs, and the unit classification is merged to form an initial classification.
The initial processing of target data is achieved, and the major categories and initial classifications of targets are obtained, classification efficiency is improved, data analysis is reduced, and classification needs are adapted to different fields and scenarios.
Smart Images

Figure CN119004288B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large-scale data set classification, and specifically is a large-scale data set classification method based on privacy computing. Background Art
[0002] With the rapid development of information technology, big data has become an important resource in modern society. In many fields, such as finance, medical care, and e-commerce, massive data sets are generated every day. These data sets are not only large in scale, but also contain rich information, which is of great value for applications such as machine learning and data mining. However, when using these data for classification and analysis, there is often a risk of data privacy leakage. Traditional data set classification systems usually need to store data in a centralized server for unified processing and analysis. However, this approach has obvious security risks. Once the server is hacked or the data is leaked by internal personnel, it will lead to the leakage and abuse of user privacy. In addition, due to the centralized storage of data, it is also easy to cause disputes over data ownership and usage rights.
[0003] In order to solve these problems, privacy computing technology has received widespread attention in recent years. Privacy computing aims to analyze and calculate data without leaking the original data to the data provider. By adopting technical means such as encryption, secure multi-party computing (MPC), and federated learning, privacy computing can maximize the value of data while protecting data privacy. However, existing privacy computing technology still has some limitations in the classification of large-scale data sets. First, since large-scale data sets usually contain massive amounts of data and complex features, traditional privacy computing methods often fail to meet the requirements in terms of processing efficiency and accuracy. Second, existing privacy computing systems usually lack flexibility and cannot adapt to classification needs in different fields and scenarios.
[0004] Based on this, the present invention proposes a large-scale data set classification method based on privacy computing. Summary of the invention
[0005] In order to solve the problems existing in the above solutions, the present invention provides a large-scale data set classification method based on privacy computing.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] A classification method for large-scale data sets based on privacy computing, the method comprising:
[0008] Step 1: Establish a label library, which is used to store various classification labels, determine target data, and the target data includes each unit data; perform initial classification on the target data to obtain each unit classification;
[0009] Furthermore, the method for initially classifying the target data includes:
[0010] Establishing a label recognition model based on the label library, wherein the label recognition model is used to identify a unit label group corresponding to the unit data, wherein the unit label group is composed of various classification labels;
[0011] Analyze each unit data in the target data through the label recognition model to obtain the unit label group corresponding to each unit data;
[0012] The unit data with the same unit label group are classified into one category to obtain several unit classifications.
[0013] Step 2: Identify the user's classification needs, and determine each target category according to the classification needs; sort out each unit classification according to each target category, and obtain the target overall data corresponding to each target category and the corresponding irrelevant classification data; the target overall data is composed of each unit classification belonging to the corresponding target category;
[0014] Step 3: Merge each unit classification in the overall target data corresponding to each target category to obtain each initial classification;
[0015] Furthermore, the method for merging the unit classifications in the target overall data includes:
[0016] Step SA1: Identify the main label group of each unit classification;
[0017] Step SA2: Establish a label similarity statistical table, which is used to count the similarities between various classification labels in various label types; perform a monopolar trend adjustment on each unit classification based on the label similarity statistical table to obtain each regional class;
[0018] Step SA3: identifying each main label group within the region class, and calculating the difference value between each main label group within the region class;
[0019] When the difference value is greater than the threshold X1, the corresponding main label groups are not merged;
[0020] When the difference value is not greater than the threshold X1, the corresponding main label groups are merged to form several merged classes;
[0021] Determine whether the merging requirements are met between the merged classes and between the merged classes and the main label group;
[0022] When it is determined that the merging requirements are met, the corresponding merging is performed to obtain a new merging class;
[0023] When it is judged that the merger requirements are not met, the corresponding merger will not be carried out;
[0024] This process is repeated until no more merging is possible, and the merged classes corresponding to each regional class are obtained;
[0025] Step SA4: Determine whether the merging requirements are met between the merged classes within different regional classes, between the merged classes and the main label groups, and between the main label groups, and perform corresponding merging according to the judgment results to obtain new merged classes; mark each merged class as an initial classification.
[0026] Furthermore, in step SA2, the method for adjusting the monopole trend of each unit classification based on the label similarity statistical table includes:
[0027] Identify the main label group corresponding to each unit classification in the target overall data, and determine the label types; set the benchmark label corresponding to each label type; determine the similarity between the benchmark labels according to the label similarity statistics table, set the polar axis corresponding to each benchmark label according to the similarity between the benchmark labels, and generate the corresponding polar axis diagram;
[0028] According to the label similarity statistics table, the similarity between each classification label in each main label group and each benchmark label is matched and marked as the matching similarity, and the polar axis corresponding to the benchmark label with the highest matching similarity is marked as the main polar axis; the corresponding repulsive axis and proximal axis are determined according to the main polar axis;
[0029] Identify the matching similarity between the reference labels of the repulsive axis and the proximal axis and the corresponding classification labels in the main label group; mark the label points corresponding to the main label group according to the matching similarity corresponding to the repulsive axis and the proximal axis and the corresponding main polar axis;
[0030] Determine each polar region according to the distribution of each label point in the polar axis diagram; and divide each main label group in the polar region into a regional class.
[0031] Furthermore, in step SA3, the method for calculating the combined value between the main tag groups includes:
[0032] Determine the single value corresponding to each label type in the two main label groups according to the label similarity statistics table; mark the single value as DYi; i represents the corresponding label type, i = 1, 2, ..., n, n is a positive integer;
[0033] According to the formula Calculate the corresponding difference value;
[0034] Where: CY is the difference value; e is the natural constant.
[0035] Step 4: Determine each candidate optimization classification scheme based on each initial classification;
[0036] Step Five: Set corresponding candidate privacy solutions according to each candidate optimization classification solution, and combine each candidate privacy solution and candidate optimization classification solution into each test combination;
[0037] Step Six: Establish a test terminal, test the corresponding test combinations through the test terminal, determine the corresponding target classification combination, and set the corresponding application classification solution according to the candidate privacy solution and candidate optimization classification solution corresponding to the target classification combination;
[0038] Further, the method for establishing a test terminal includes:
[0039] Set each simulation combination, determine each to-be-established channel according to each simulation combination; establish corresponding test channels according to each to-be-established channel;
[0040] Establish a corresponding test terminal according to each test channel.
[0041] Further, another method for establishing a test terminal includes:
[0042] Determine each reserve channel, and set the test scope of each reserve channel;
[0043] Set various test combinations according to the test scope of each reserve channel;
[0044] Obtain the number of reserve channels corresponding to each test combination, and set the protection safety value of each test combination;
[0045] Evaluate the implementation cost of each test combination;
[0046] Remove the dimension and take its numerical value for calculation, and calculate the corresponding test priority value according to the formula QKY = b1×AF - b2×CBD;
[0047] In the formula: QKY is the test priority; b1 and b2 are both proportionality coefficients, and the value range is 0 < b1 ≤ 1, 0 < b2 ≤ 1; AF is the protection safety value; CBD is the implementation cost;
[0048] Establish a test terminal according to the test combination with the highest test priority value.
[0049] Step Seven: Classify data according to the application classification solution.
[0050] Compared with the prior art, the beneficial effects of the present invention are:
[0051] Through the settings of steps one to three, the target data is initially processed to obtain each target category and the initial classification corresponding to each target category. The classification can be performed according to the existing preset classification method, so as to achieve flexible adjustment according to the classification needs of the user. By classifying based on the target category, a large amount of data that does not need to be classified will be quickly screened out, thereby improving the classification efficiency. By first dividing each main label group into each regional category, it is avoided to analyze the relationship between the two main label groups one by one, otherwise it will lead to the analysis of a large amount of data. By first dividing into each regional category and analyzing the main label groups in each regional category, the amount of data analysis will be greatly reduced and the merging efficiency will be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0053] Figure 1 is a flow chart of the method of the present invention;
[0054] Figure 2 This is an example diagram of the monopolar trend of the present invention. DETAILED DESCRIPTION
[0055] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0056] like Figure 1 to Figure 2 As shown, a classification method for large-scale data sets based on privacy computing includes:
[0057] Step 1: Establish a label library, which is used to store various data labels related to data classification, and mark the corresponding data labels as classification labels; establish a corresponding label recognition model according to the various classification labels stored in the label library, and the label recognition model is used to identify various data and determine the various classification labels corresponding to them. Specifically, the label recognition model can be established based on neural networks, etc., and the classification labels corresponding to various data can be marked in advance and integrated into a training set for training; the existing label recognition model can also be directly applied;
[0058] Obtain a large-scale data set that needs to be classified, and mark the corresponding data set as target data; the target data has been pre-processed, such as data cleaning, data standardization or normalization, etc.; each single data in the target data is marked as unit data; that is, the target data is composed of each unit data;
[0059] Each unit data in the target data is analyzed through the label recognition model to obtain the classification labels corresponding to the unit data, forming a label group, which is marked as a unit label group; the unit data with the same unit label group are classified into one category to obtain several unit classifications.
[0060] Step 2: Identify the user's classification needs, that is, the user can adjust the corresponding classification needs in subsequent applications according to actual needs, to avoid fixed classification methods, which lead to lack of flexibility and cannot adapt to dynamic needs; determine the target categories according to the classification needs; that is, according to the user's classification needs, determine which data the user wants to classify from the target data, or which data the user wants to apply; determine the corresponding target categories according to each goal; such as machine learning data for a certain purpose, user behavior analysis data and other related classification needs; sort out the unit classifications according to the target categories, obtain the unit classifications corresponding to the target categories, integrate them into the target overall data of the target category, and classify the unit label groups corresponding to the unit classifications; integrate the remaining unit classifications that are not in the target categories into irrelevant classification data; that is, there is no need to classify and analyze the irrelevant classification data in the future.
[0061] Step 3: Merge each unit classification in the overall target data corresponding to each target category to obtain each initial classification.
[0062] Through the settings of steps one to three, the target data can be initially processed to obtain each target category and the initial classification corresponding to each target category. Subsequently, the data can be classified according to the existing preset classification method, so as to realize flexible adjustment according to the classification needs of the user. By classifying based on the target category, a large amount of data that does not need to be classified can be quickly screened out, thereby improving the classification efficiency.
[0063] Among them, the method of merging each unit classification in the target overall data includes:
[0064] Step SA1: Mark the main label group in the unit label group corresponding to each unit classification according to the target category, that is, the combination of classification labels that makes the unit classification belong to the target category, and the classification label combination is the main label group;
[0065] Step SA2: Identify various classification tags in the tag library, identify the classification meanings corresponding to each classification tag, and classify the classification tags into different tag types according to the classification meanings, such as film and television tag types, animal tag types, etc.; divide them specifically according to needs; then set the similarity between different classification tags in each tag type, analyze the corresponding similarity based on the corresponding classification meanings; perform corresponding sorting to obtain the corresponding tag similarity statistics table;
[0066] Based on the label similarity statistics table, the classification of each unit is adjusted to a single pole trend to obtain each polar region; each main label group in the polar region is divided into a regional class.
[0067] For example, Figure 2 As shown, according to the identification of each main label group, 8 types of classification labels are determined, namely Figure 2 For label types 1-8, the corresponding polar axes are equally divided according to the number of label types, and the reference labels corresponding to each label type, i.e., the classification labels within the type of label, are set by optional or manual setting; the similarity between the reference labels is identified, and the polar axes are distributed according to the similarity, such as the polar axis with the lowest similarity corresponding to a polar axis is the farthest from it; the two polar axes corresponding to the lowest similarity are marked as mutually exclusive polar axes; the two polar axes with the highest similarity are marked as close polar axes;
[0068] Calculate the similarity between each classification label and each benchmark label in each main label group, and mark the polar axis corresponding to the benchmark label with the highest similarity in the main label group as the main polar axis, that is, Figure 2 The polar axis corresponding to the label type 1 in the polar axis; the mutually exclusive polar axis corresponding to the main polar axis is identified, and marked as the repulsive axis; the adjacent polar axis corresponding to the main polar axis is identified, and marked as the paraxial axis; the similarity corresponding to the reference label corresponding to the repulsive axis in the main label group is identified, and A1=100-the similarity of the repulsive axis; the similarity corresponding to the reference label corresponding to the paraxial axis in the main label group is identified, and A2=the similarity of the paraxial axis; the main label group is regarded as a label point and marked in the polar axis diagram according to A1 and A2; each polar region is determined according to the distribution of each label point in the polar axis diagram; in one embodiment, each label point in the polar axis diagram can be clustered to obtain a number of clustering regions, and the clustering regions are marked as polar regions; for the clustering process, various clustering algorithms such as K-means clustering algorithm, DBSCAN, hierarchical clustering, OPTICS, etc. can be used for clustering, and corresponding restrictions can be preset, and a large clustering allowable error is obtained, which does not affect the subsequent classification accuracy; in other embodiments, other existing technologies can also be applied to form each polar region according to the distribution of label points.
[0069] By first dividing each main label group into each regional class, we avoid analyzing the pairwise relationships between numerous main label groups one by one, which would otherwise lead to analyzing a large amount of data. Instead, by first dividing into each regional class and analyzing the main label groups within each regional class, the amount of data analysis will be greatly reduced and the merging efficiency will be improved.
[0070] Step SA3: identifying each main label group within the region class, and calculating the difference value between each main label group within the region class;
[0071] When the difference value is greater than the threshold X1, the corresponding main label groups are not merged;
[0072] When the difference value is not greater than the threshold value X1, the corresponding main label groups are merged to form several merged classes; that is, the difference values between the main label groups in a merged class are not greater than the threshold value X1;
[0073] Determine whether the merged classes and the merged classes and the main label group meet the merge requirements; for the merged classes, if the difference values between the main label groups in the merged class are not greater than the threshold X1, it is considered that the merge requirements are met; the same applies to the merged class and the main label group;
[0074] When it is determined that the merging requirements are met, the corresponding merging is performed to obtain a new merging class;
[0075] When it is judged that the merging requirements are not met, the corresponding merging is not performed; at this time, the merging within the regional category is completed, and the merging outside the regional category is performed later;
[0076] And so on, until it cannot be merged anymore; obtain the merged class corresponding to each regional class;
[0077] The calculation method of the combined value between the main label groups includes:
[0078] Determine the similarity of the classification labels belonging to the same label type in the two main label groups according to the label similarity statistics table; mark it as a single value; mark the single value as DYi; i represents the corresponding label type, i = 1, 2, ..., n, n is a positive integer;
[0079] According to the formula Calculate the corresponding difference value;
[0080] Where: CY is the difference value; e is the natural constant.
[0081] Step SA4: Determine whether the merging requirements are met between the merged classes within different regional classes, between the merged classes and the main label groups, and between the main label groups, and perform corresponding merging according to the judgment results; refer to step SA3; until merging is impossible; obtain each new merged class; mark each merged class as the initial classification.
[0082] Step 4: Determine the candidate optimization classification schemes based on the initial classifications; that is, after determining the initial classifications, combine the various existing classification data processing methods to determine which schemes are available for subsequent classification processing of the initial classifications. Specifically, the platform will count the various schemes that exist under the premise of having the initial classifications, make corresponding adjustments to them, form alternative schemes, mark the initial classifications applicable to each alternative scheme, establish a corresponding scheme library, and then perform corresponding matching; or directly set the corresponding candidate optimization classification schemes based on the initial classifications manually. Although each candidate optimization classification scheme can meet the subsequent classification processing, the adaptability of privacy computing technology also needs to be considered. Therefore, each corresponding scheme is used as a candidate optimization classification scheme.
[0083] Step 5: Determine various privacy computing technologies suitable for the application based on each candidate optimization classification scheme, and mark them as candidate privacy schemes. The determination is mainly based on various factors that have an impact on the application of privacy computing technology, such as the data processing method and data type corresponding to the corresponding candidate optimization classification scheme; combine each candidate optimization classification scheme and the corresponding candidate privacy scheme to form several feasible combinations, which are marked as test combinations.
[0084] Step 6: Establish a test terminal, test the corresponding test combination through the test terminal, determine the corresponding target classification combination, and set the corresponding application classification scheme according to the candidate privacy scheme and the candidate optimization classification scheme corresponding to the target classification combination; the test terminal is used to test various test combinations to understand the privacy protection effect of the test combination in the data classification process;
[0085] In one embodiment, the method for establishing a test terminal includes:
[0086] Based on various historical test data of privacy computing technology, determine the various possible test combinations, and then determine whether each test combination can be tested in the same way based on historical test data, and integrate those that can be tested together to form several simulation combinations, that is, one simulation combination corresponds to a test combination within a common range; finally, determine the test channels that need to be established based on each simulation combination, that is, one test channel corresponds to a range of test combinations within a range; mark the test channels that need to be established as channels to be established;
[0087] Establish corresponding test channels according to the channels to be built;
[0088] The platform will establish corresponding test terminals based on each test channel.
[0089] In another embodiment, another method of establishing a test end includes:
[0090] The platform determines various directly accessible test channels, which do not need to be established from scratch by the platform. Mark the corresponding test channels as reserve channels and determine the test scope of each reserve channel;
[0091] Obtain various test combinations, that is, combine each reserve channel according to each test scope to form each candidate channel combination. Specifically, combine according to each test scope. For example, a candidate channel combination can test all possible test scopes;
[0092] Obtain the number of reserve channels corresponding to each test combination, conduct test simulations on each test combination, and obtain the protection safety value corresponding to each test combination. That is, directly conduct protection simulations manually to determine the overall security protection ability of the test combination. The value range of the protection safety value is [0, 100]. Specifically, first determine various possible protection capabilities, sort the protection capabilities in descending order, mark the protection safety value ranked first as 100, the lowest as 0, and set corresponding protection safety values for others according to the protection ability differences; form a matching table corresponding to the protection safety value; match the corresponding protection safety value for each reserve channel according to the matching table; calculate the average value of the protection safety values of each reserve channel in the test combination, and use the corresponding average value as the protection safety value of the test combination;
[0093] Evaluate the implementation cost of each test combination, that is, the cost of establishing a test end according to the test combination;
[0094] Remove the dimension and take its numerical value for calculation. Calculate the corresponding test priority value according to the formula QKY = b1×AF - b2×CBD; where: QKY is the test priority; b1 and b2 are both proportionality coefficients, and the value range is 0 < b1 ≤ 1, 0 < b2 ≤ 1; AF is the protection safety value; CBD is the implementation cost.
[0095] Establish a test end according to the test combination with the highest test priority value.
[0096] Step Seven: Classify data according to the application classification scheme.
[0097] The above formulas are all calculated by removing the dimension and taking its numerical value. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through a large amount of data simulation.
[0098] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A classification method for large-scale data sets based on privacy computing, characterized in that: Methods include: Step 1: Establish a label library, which is used to store various classification labels and determine target data, which includes each unit data; Perform initial classification on the target data to obtain the classification of each unit; Step 2: Identify the user's classification needs, and determine each target category according to the classification needs; sort out each unit classification according to each target category, and obtain the target overall data corresponding to each target category and the corresponding irrelevant classification data; the target overall data is composed of each unit classification belonging to the corresponding target category; Step 3: Merge each unit classification in the overall target data corresponding to each target category to obtain each initial classification; Step 4: Determine each candidate optimization classification scheme based on each initial classification; Step 5: setting corresponding privacy schemes to be selected according to the optimization classification schemes to be selected, and combining the privacy schemes to be selected and the optimization classification schemes to be selected into test combinations; Step 6: Establish a test terminal, test the corresponding test combination through the test terminal, determine the corresponding target classification combination, and set the corresponding application classification scheme according to the candidate privacy scheme and the candidate optimization classification scheme corresponding to the target classification combination; Step 7: Classify data according to the application classification scheme.
2. According to claim 1, a privacy computing-based large-scale data set classification method is characterized in that: Methods for initial classification of target data include: Establishing a label recognition model based on the label library, wherein the label recognition model is used to recognize a unit label group corresponding to the unit data, wherein the unit label group is composed of various classification labels; Analyze each unit data in the target data through the label recognition model to obtain the unit label group corresponding to each unit data; The unit data with the same unit label group are classified into one category to obtain several unit classifications.
3. According to claim 1, a privacy computing-based large-scale data set classification method is characterized in that: Methods for merging the classification of each unit in the target overall data include: Step SA1: Identify the main label group of each unit classification; Step SA2: Establish a label similarity statistical table, which is used to count the similarities between various classification labels in various label types; perform a monopolar trend adjustment on each unit classification based on the label similarity statistical table to obtain each regional class; Step SA3: identifying each main label group within the region class, and calculating the difference value between each main label group within the region class; When the difference value is greater than the threshold X1, the corresponding main label groups are not merged; When the difference value is not greater than the threshold X1, the corresponding main label groups are merged to form several merged classes; Determine whether the merging requirements are met between the merged classes and between the merged classes and the main label group; When it is determined that the merging requirements are met, the corresponding merging is performed to obtain a new merging class; When it is judged that the merger requirements are not met, the corresponding merger will not be carried out; This process is repeated until no more merging is possible, and the merged classes corresponding to each regional class are obtained; Step SA4: Determine whether the merging requirements are met between the merged classes within different regional classes, between the merged classes and the main label groups, and between the main label groups, and perform corresponding merging according to the judgment results to obtain new merged classes; mark each merged class as an initial classification.
4. According to claim 3, a privacy computing-based large-scale data set classification method is characterized in that: In step SA2, the method for adjusting the monopole trend of each unit classification based on the label similarity statistical table includes: Identify the main label groups corresponding to each unit classification in the target overall data, and determine the respective label types; set the reference labels corresponding to each label type; determine the similarity between each reference label according to the label similarity statistical table, set the polar axis corresponding to each reference label according to the similarity level between each reference label, and generate the corresponding polar axis diagram; Match the similarity between each classification label in each main label group and each reference label according to the label similarity statistical table, mark it as the matching similarity, and mark the polar axis corresponding to the reference label with the highest matching similarity as the main polar axis; determine the corresponding repulsion axis and proximity axis according to the main polar axis; Identify the matching similarity between the reference labels of the repulsion axis and the proximity axis and the corresponding classification labels in the main label group; mark the label points corresponding to the main label group according to the matching similarity corresponding to the repulsion axis and the proximity axis and the corresponding main polar axis; Determine each polar region according to the distribution of each label point in the polar axis diagram; divide each main label group within the polar region into a region class.
5. According to claim 3, a privacy-oriented large-scale data set classification method is characterized in that: In step SA3, the calculation method of the merging value between main label groups includes: Determine the single value corresponding to each label type in two main label groups according to the label similarity statistical table; mark the single value as DYi; i represents the corresponding label type, i = 1, 2,..., n, and n is a positive integer; According to the formula Calculate the corresponding difference value; In the formula: CY is the difference value; e is the natural constant.
6. The privacy computing-based classification method for large-scale data sets according to claim 1, characterized in that: The method for establishing a test end includes: Set each simulation combination, determine each to-be-built channel according to each simulation combination; establish the corresponding test channel according to each to-be-built channel; Establish the corresponding test end according to each test channel.
7. The privacy computing-based classification method for large-scale data sets according to claim 1, characterized in that: The method for establishing a test end includes: Determine each reserve channel, and set the test range of each reserve channel; Set various test combinations according to the test range of each reserve channel; Obtain the number of reserve channels corresponding to each test combination, and set the protection safety value of each test combination; Evaluate the implementation cost of each test combination; Remove the dimension and take its numerical value for calculation, and calculate the corresponding test priority value according to the formula QKY = b1 × AF - b2 × CBD; In the formula: QKY is the test priority; b1 and b2 are both proportionality coefficients, and the value range is 0 < b1 ≤ 1, 0 < b2 ≤ 1; AF is the protection safety value; CBD is the implementation cost; Establish a test end according to the test combination with the highest test priority value.
Citation Information
Patent Citations
Federal learning privacy protection method and system based on adversarial training
CN113609521A
Data processing method and system based on privacy computing
CN115168045A