CTGAN network structure, equipment and medium

By introducing a business rule database and a rule-vector mapping engine into the CTGAN network, and combining a generator and a discriminator, data that conforms to business rules is generated. Privacy enhancement is achieved through association graphs, thus solving the problem of logical errors in the data generated by the CTGAN network and realizing both accuracy and privacy protection.

CN120996095APending Publication Date: 2025-11-21广州数据集团有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511046726.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

When generating data that is similar to statistical data, the existing CTGAN network does not explicitly model the business logic relationships between entities, resulting in logical errors in the generated data and failure to conform to real business rules.

Method used

By introducing a business rule database and a rule-vector mapping engine, business rules are combined with the labels of the generated data. The data generated by the generator is judged by a discriminator and the business rule database to filter out compliant data that conforms to the business rules, and privacy enhancement processing is performed through the association graph.

Benefits of technology

The generated data conforms to business logic and is accurate, with effective privacy protection, avoiding the break in data relationships and enhancing the practical application value of the generated data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996095A_ABST
    Figure CN120996095A_ABST
Patent Text Reader

Abstract

The invention discloses a CTGAN network structure, equipment and a medium, and the network structure comprises a rule-vector mapping engine, one end of which is used for obtaining a label of data needing to be generated and is connected with a business rule database, and the other end of which is connected with a generator and is used for obtaining a rule coding vector according to the label of the data needing to be generated and the business rule database, outputting the regular coding vector to a generator; the other end of the generator is connected with the discriminator, and the generator is used for obtaining constraint compliance data according to the rule coding vector and outputting the constraint compliance data to the discriminator; one end of the discriminator is connected with the generator and the business rule database, the other end of the discriminator outputs data, the discriminator is used for determining compliance data in the constraint compliance data according to the business rule database and the constraint compliance data and outputting the compliance data, and the compliance data is the generated data. Accurate data meeting service logic can be generated in combination with actual service rules.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of data simulation, in particular to a CTGAN network structure, a device and a medium. BACKGROUND

[0002] In current synthetic data technology based on a generative adversarial network (GAN), CTGAN (Conditional TabularGANs) is a network for tabular data, which performs excellently in preserving data statistical distribution. It learns the data edge distribution through adversarial training, and can generate simulation data similar to the statistical characteristics of the original data, and has a wide application basis in the fields of public data opening and enterprise data simulation.

[0003] However, in the existing technology for generating data similar to statistical data, CTGAN only learns the data edge distribution through adversarial training, without explicitly modeling the business logic association between entities, resulting in logical errors and inaccuracy of the generated data, which does not conform to the real business rules. SUMMARY

[0004] The present application aims to at least solve one of the technical problems existing in the prior art. To this end, the present application proposes a CTGAN network structure which can generate data conforming to business logic and accurate in combination with actual business rules.

[0005] The present application also proposes a device and a medium for implementing the above-mentioned CTGAN network structure.

[0006] According to the CTGAN network structure of the first aspect of the present application, it comprises:

[0007] a business rule database connected to a rule-vector mapping engine and a discriminator respectively;

[0008] The rule-vector mapping engine is used to obtain the label of the data to be generated and is connected to the business rule database at one end, and is connected to the generator at the other end, and is used to obtain a rule encoding vector according to the label of the data to be generated and the business rule database, and output the rule encoding vector to the generator;

[0009] The generator is connected to the rule-vector mapping engine at one end and is connected to the discriminator at the other end, and is used to receive the rule encoding vector output by the rule-vector mapping engine, obtain constraint compliant data according to the rule encoding vector, and output the constraint compliant data to the discriminator;

[0010] The discriminator, one end is connected with the generator and the business rule database, and the other end outputs data, is used for receiving the constraint compliance data sent by the generator, determining the compliance data in the constraint compliance data according to the business rule database and the constraint compliance data, and outputting the compliance data, and the compliance data is the generated data.

[0011] According to the CTGAN network structure provided by the embodiment of the application, the following beneficial effects are achieved: the CTGAN network structure combines the business rule database and the rule-vector mapping engine, the rule-vector mapping engine is used to combine the business rule and the label of the required generated data, so that the rule encoding vector required by the generator is obtained, the data generated by the generator is discriminated by the discriminator and the business rule database, and the data meeting the business rule, that is, the compliance data, is screened out; the CTGAN network structure is improved, and the business rule database is combined from the generator to the discriminator, so that the generated data is combined with the actual business rule, meets the business logic, and is accurate.

[0012] According to some embodiments of the application, the method further comprises:

[0013] The association graph, one end is connected with the discriminator, and the other end outputs data, is used for performing association and reservation type privacy enhancement processing on the compliance data sent by the discriminator, and specifically comprises:

[0014] Detecting the sensitive field in the compliance data;

[0015] According to the sensitive field associated with the corresponding sensitive entity;

[0016] Adding privacy noise to the detected sensitive field and the sensitive field associated with the sensitive entity; wherein the intensity of the privacy noise is adjusted according to the sensitivity level of the corresponding sensitive field.

[0017] Output the processed data, and the output data is the generated data.

[0018] According to some embodiments of the application, the intensity of the privacy noise is adjusted according to the sensitivity level of the corresponding sensitive field, and specifically comprises:

[0019] The calculation formula of the intensity of the privacy noise is as follows:

[0020] ∈ i = alpha * log (sensitivityLevel i + 1), wherein, ∈ i is the intensity of the privacy noise corresponding to the i th sensitive field, alpha is a preset parameter, sensitivityLevel i is the sensitivity level of the i th sensitive field.

[0021] According to some embodiments of the present invention, the discriminator includes an association compliance discrimination branch and an authenticity discrimination branch;

[0022] The discriminator outputs compliance data based on the business rule base and the constraint compliance data, including:

[0023] The discriminator inputs the constraint compliance data into the associated compliance discrimination branch;

[0024] In the related compliance determination branch, the business rule base is searched according to the tags of the constraint compliance data to obtain the search results;

[0025] Based on the search results, determine whether the data value of the constraint compliance data conforms to the business rules. If yes, it is considered compliant data; otherwise, it is considered non-compliant data.

[0026] The subset of the data that conforms to the rules and the data output by the authenticity discrimination branch is used as the compliant data output.

[0027] According to some embodiments of the present invention, the CTGAN network structure further includes: a loss calculation module, one end of which is connected to the discriminator and the other end of which is connected to the generator, for calculating the loss function of the compliant data output by the discriminator and adjusting the parameters of the generator according to the loss function;

[0028] The loss function includes a constraint compliance loss term to drive the generator to output constraint compliance data that conforms to business rules; the formula for the constraint compliance loss term is as follows:

[0029] L constraint = 1 - JointComplianceRate, where L constraint The constraint compliance loss term is defined as JointComplianceRate, which is the joint compliance rate output by the associated compliance discrimination branch.

[0030] According to some embodiments of the present invention, the rule-vector mapping engine obtains the rule encoding vector based on the labels of the data to be generated and the business rule base, specifically including:

[0031] Based on the tags corresponding to the required number of data to be generated, the business rule base is searched, and the business rules corresponding to the tags of the required data to be generated are retrieved to obtain the tag-business rule correspondence.

[0032] Based on the correspondence between the tag and the business rule, the corresponding business rule is mapped to the corresponding rule encoding vector, and a correspondence between the tag and the rule encoding vector is formed.

[0033] According to some embodiments of the present invention, the method for constructing the rule-vector mapping engine includes:

[0034] Obtain the real dataset;

[0035] Based on the business rule base, determine the business rules that the real data in the real dataset must conform to;

[0036] Based on the categories of business rules that the real data needs to conform to, the corresponding encoding vectors for each category are obtained;

[0037] The real data, the rule encoding vector corresponding to the real data, and the correspondence between the real data and the rule encoding vector corresponding to the real data are stored in the rule-vector mapping engine.

[0038] Based on the categories of business rules that the real data needs to conform to, the corresponding encoding vectors for each category are obtained, including: Boolean scalar, sparse vector, and probability vector;

[0039] The Boolean scalar, the sparse vector, and the probability vector are concatenated to obtain the rule encoding vector corresponding to the real data.

[0040] According to some embodiments of the present invention, obtaining the encoding vector corresponding to each category based on the categories of business rules that the real data needs to conform to specifically includes: extracting fields from the real data for the categories of business rules that the real data needs to conform to, including inequality rules; and performing comparison operations based on the extracted fields to obtain the Boolean scalar.

[0041] The categories of business rules that the real data needs to conform to include category mapping rules are used to encode the real data using one-hot encoding to obtain the sparse vector;

[0042] The categories of business rules that the real data needs to comply with include numerical range rules. The real data is then input into the interval calculator and the compliance calculator in sequence to obtain the probability vector.

[0043] An electronic device according to a second aspect of the present invention includes:

[0044] Memory, used to store programs;

[0045] A processor for executing a program stored in the memory, wherein when the processor executes the program stored in the memory, the processor is configured to deploy a network structure as described in any one of the first aspects.

[0046] According to a third aspect of the present invention, a storage medium stores computer-executable instructions for deploying a network structure as described in any one of the first aspects.

[0047] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the description, claims, and drawings. Attached Figure Description

[0048] The accompanying drawings are provided to further understand the technical solutions of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the technical solutions of the present invention, and do not constitute a limitation on the technical solutions of the present invention.

[0049] Figure 1 This is a schematic diagram of a CTGAN network structure provided in an embodiment of the present invention;

[0050] Figure 2 This is a schematic diagram of a CTGAN network structure provided in another embodiment of the present invention;

[0051] Figure 3 This is a schematic diagram illustrating the construction of a rule-vector mapping engine in a CTGAN network structure according to another embodiment of the present invention;

[0052] Figure 4 This is a schematic diagram of the correlation graph in a CTGAN network structure provided in another embodiment of the present invention. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0054] It should be understood that in the description of the embodiments of the present invention, "multiple" (or "amounts") means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. If "first," "second," etc., are used in the description, they are only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.

[0055] like Figure 1 As shown, an embodiment of the present invention provides a CTGAN network structure, including:

[0056] The business rules database is connected to the rule-vector mapping engine and the discriminator, respectively.

[0057] The rule-vector mapping engine has two ends: one end is used to obtain the labels of the data to be generated and connect to the business rule database, and the other end is connected to the generator. The generator obtains the rule encoding vector based on the labels of the data to be generated and the business rule database, and outputs the rule encoding vector to the generator.

[0058] The generator connects to the rule-vector mapping engine on one end and the discriminator on the other. It receives the rule encoding vectors output by the rule-vector mapping engine, obtains the constraint compliance data based on the rule encoding vectors, and outputs the constraint compliance data to the discriminator.

[0059] The discriminator connects to the generator and business rule database on one end and to the loss calculation module on the other. It receives constraint compliance data sent by the generator, determines the compliant data in the constraint compliance data based on the business rule database and constraint compliance data, and outputs the compliant data, which is the generated data.

[0060] The CTGAN network structure combines a business rule database and a rule-vector mapping engine. First, the rule-vector mapping engine combines business rules with the labels of the data to be generated, obtaining the rule encoding vector required for the generator input. The data generated by the generator is then judged by the discriminator and the business rule database, and data that conforms to the business rules, i.e., compliant data, is selected. This invention improves the CTGAN network structure by combining the business rule database from the generator to the discriminator, so that the generated data combines actual business rules, conforms to business logic, and is accurate.

[0061] It should be noted that the data includes labels and data values. For example, the label is "approval time" and the data value is "3 days". The generator's input also includes external features and noise.

[0062] like Figure 2 As shown, in one embodiment, the CTGAN network structure further includes:

[0063] The association graph, connected to the discriminator at one end and outputting data at the other, is used to perform association-preserving privacy enhancement processing on the compliant data sent by the discriminator, specifically including:

[0064] Detect sensitive fields in compliant data;

[0065] Associate the corresponding sensitive entity based on the sensitive field;

[0066] Privacy noise is added to the detected sensitive fields and sensitive fields associated with sensitive entities; the intensity of the privacy noise is adjusted according to the sensitivity level of the corresponding sensitive field.

[0067] The processed data is then output, and the output data is the generated data.

[0068] Because when a CTGAN network directly generates sensitive data, the superimposed privacy protection technology (such as differential privacy) can easily damage the data association relationship (e.g., the association break rate reaches 40% after enterprise data anonymization), resulting in the generated data not conforming to business rules and losing business value. Therefore, this embodiment detects sensitive fields through association graphs and processes all sensitive fields included in the entities corresponding to the sensitive fields to varying degrees to avoid the breakage of business association relationships.

[0069] It's easy to understand that sensitive fields and sensitive entities are essentially tags for data;

[0070] like Figure 4 As shown, the sensitive field "Citizen ID Number" was detected, which is associated with the sensitive entity "Personal Basic File". Other sensitive fields associated with "Personal Basic File" include "Social Security Payment Record" and "Individual Income Tax Declaration Record". Differential privacy technology (i.e., ...) is employed. Figure 4 The differential privacy engine adds noise (sensitive noise) to both the "social security payment record" and "individual income tax declaration record" and outputs the confidential data of the "social security payment record" and "individual income tax declaration record" respectively.

[0071] In one embodiment, adjusting the intensity of privacy noise according to the sensitivity level of the corresponding sensitive field specifically includes:

[0072] The formula for calculating the intensity of privacy noise is as follows:

[0073] ∈ i =α·log(sensitivityLevel) i +1), where, ∈ i Let α be the intensity of the privacy noise corresponding to the i-th sensitive field, and α be a preset parameter, sensitivityLevel. i Let be the sensitivity level of the i-th sensitive field.

[0074] It should be noted that ∈ i It can usually be understood as a certain metric related to the i-th object (a sensitive field in this example), and associated with privacy budgets and perturbation parameters in scenarios such as differential privacy, used to control the degree of protection of sensitive information or the intensity of interference during data processing;

[0075] The role of α is to adjust the "amplitude" of the entire calculation. It can be set and optimized according to actual needs (such as privacy protection requirements, the tolerance of the business scenario for the result, etc.) to scale the results of subsequent logarithmic operations and determine ∈ i With sensitivityLeveli The rate of change;

[0076] sensitivityLevel i The "sensitivity level" representing the i-th object is a quantitative indicator of the sensitive attributes of data or entities. The higher the value, the higher the sensitivity. For example, in a medical data scenario, data samples containing core patient privacy (such as genes and serious diseases) have a sensitivityLevel of 1. i It will be higher than ordinary basic information data;

[0077] The above formula first applies to sensitivityLevel i Performing the "(+1)" operation is to avoid sensitivityLevel i When the value is 0, the logarithm is meaningless (ensuring the input to the logarithmic function is always greater than 0). Then, calculate the natural logarithm log(sensitivityLevel). i +1), unless otherwise specified, generally refers to the natural logarithm ln(sensitivityLevel). i +1), of course it could also be other base logarithms (needs to be confirmed in the specific context, but it is essentially a logarithmic transformation), then multiply by α to get ∈ i The overall process involves logarithmic transformation to adjust the sensitivityLevel. i The linear change is transformed into a smoother nonlinear change, allowing ∈ i As the sensitivity level increases, the growth is relatively controllable and not too dramatic. At the same time, the strength of this correlation can be flexibly adjusted by using α. This is often used in scenarios where data utilization and sensitive information protection need to be balanced (such as privacy computing, data desensitization, etc.), and the parameters for privacy protection or data processing are dynamically adjusted according to the data sensitivity level.

[0078] In one embodiment, the CTGAN network structure further includes a preprocessing module connected to the business rule base, used to preprocess real data and parse associated rules;

[0079] The construction methods for the business rule base include:

[0080] Extract features from real data to obtain the distribution of numerical data and the features of categorical data;

[0081] Based on the extracted features (i.e., the distribution of numerical data and the features of categorical data), a set of rule constraints is constructed, which includes numerical constraints (such as "material acceptance time ≥ material submission time") and categorical constraints (such as "prescription code includes diagnosis code primary code"). The set of rule constraints is stored in the business rule base, thereby constructing the business rule base.

[0082] In one embodiment, the construction method of the business rule base further includes:

[0083] Receive custom extreme scenario rules sent by the administrator; whereby the custom extreme scenario rules are the distribution and characteristics of data under extreme scenarios;

[0084] Based on the custom extreme scenario rules, the constraint data under the extreme scenario is obtained and stored in the business rule library.

[0085] like Figure 1 , Figure 2 As shown, in one embodiment, the discriminator includes an association compliance discrimination branch and an authenticity discrimination branch;

[0086] The discriminator outputs compliance data based on the business rule base and constraint compliance data, including:

[0087] The discriminator inputs constraint compliance data into the associated compliance discrimination branch;

[0088] In the related compliance judgment branch, the business rule base is searched based on the tags of the constraint compliance data to obtain the search results;

[0089] Based on the search results, determine whether the data value of the constraint compliance data conforms to the business rules. If it does, it is considered compliant data; otherwise, it is considered non-compliant data.

[0090] The subset of the data that conforms to the rules and the data output by the authenticity discrimination branch is output as compliant data.

[0091] It should be noted that for the authenticity discrimination branch, the discriminator inputs the constraint compliance data and the real dataset into the authenticity discrimination branch. The authenticity discrimination branch is used to determine the degree of fit between the constraint compliance data and the real dataset, and outputs the constraint compliance data that meets the fit requirements. The processing of the authenticity discrimination branch is a common method in CTGAN network structures. The data output by the discriminator is the intersection of the data output by the authenticity discrimination branch and the data output by the associated compliance discrimination branch.

[0092] like Figure 1 , Figure 2 As shown, in one embodiment, the CTGAN network structure further includes: a loss calculation module, one end of which is connected to the discriminator and the other end of which is connected to the generator, used to calculate the loss function of the compliant data output by the discriminator and adjust the parameters of the generator according to the loss function;

[0093] Specifically, a constraint compliance loss term is added to the loss function to drive the generator to output constraint compliance data that conforms to business rules; the formula for the constraint compliance loss term is as follows:

[0094] L constraint= 1 - JointComplianceRate, where L constraint To constrain compliance loss items, JointComplianceRate is the joint compliance rate output by the associated compliance discrimination branch.

[0095] It should be noted that the loss calculation module is actually feedback, used to drive the generator to generate more reasonable data; the larger the value of the constraint compliance loss item, the lower the compliance level of the generated data. By allowing the generator to optimize to reduce this loss, more compliant data is generated.

[0096] In CTGAN scenarios, this constraint is used to drive the generator to prioritize the generation of compliant data. The CTGAN network structure is often used to generate tabular data (such as data with business rules in fields like finance and healthcare). Adding this compliance loss term ensures that the generated data not only closely approximates real data in terms of statistical characteristics like distribution, but also meets industry and business compliance standards and attribute constraints (such as format specifications and data privacy requirements for financial data, and coding and confidentiality rules for medical data). This improves the usability and legality of the generated data in actual business applications, preventing the generated data from being unusable or causing risks due to non-compliance.

[0097] In one embodiment, the loss calculation module adjusts the generator parameters according to the loss function, including:

[0098] The generator's data generation is optimized with the goal of minimizing the Wasserstein distance (i.e., bulldozer distance) between the generated data and the real data, thereby improving the accuracy of statistical distribution restoration.

[0099] The methods by which the generator generates data include:

[0100] Based on the parameters adjusted by the loss calculation module, the preset generation function, and the real dataset, generate the pre-output data;

[0101] Determine the bulldozer distance between the pre-output data and the actual data, and output the pre-output data whose bulldozer distance is less than the Wasserstein distance threshold.

[0102] In one embodiment, the Wasserstein distance threshold is 0.15.

[0103] In one embodiment, the rule-vector mapping engine obtains the rule encoding vector based on the labels of the data to be generated and the business rule base, specifically including:

[0104] Based on the tags corresponding to the required number of data to be generated, the business rule base is searched, and the business rules corresponding to the tags of the required data to be generated are retrieved to obtain the tag-business rule correspondence.

[0105] Based on the correspondence between tags and business rules, and the corresponding business rules, the corresponding business rules are mapped to the corresponding rule encoding vectors, thus forming a correspondence between tags and rule encoding vectors.

[0106] like Figure 3 As shown, in one embodiment, the method for constructing a rule-vector mapping engine includes:

[0107] Obtain the real dataset;

[0108] Based on the business rule base, determine the business rules that the real data in the real dataset must comply with;

[0109] Based on the categories of business rules that the real data needs to conform to, the corresponding encoding vectors for each category are obtained;

[0110] Store the real data, the rule encoding vectors corresponding to the real data, and the correspondence between the real data and the rule encoding vectors corresponding to the real data into the rule-vector mapping engine;

[0111] Based on the categories of business rules that the real data needs to conform to, the corresponding encoding vectors for each category are obtained, including Boolean scalars, sparse vectors, and probability vectors. The Boolean scalars, sparse vectors, and probability vectors are concatenated to obtain the rule encoding vectors corresponding to the real data.

[0112] It should be noted that a single piece of data may satisfy multiple business rules.

[0113] like Figure 3 As shown, in one embodiment, according to the categories of business rules that the real data needs to conform to, the corresponding encoding vectors are obtained respectively. Specifically, this includes: for the categories of business rules that the real data needs to conform to, including inequality rules, extracting fields from the real data; performing comparison operations based on the extracted fields to obtain Boolean scalars; for example, for fields A and B, if field A ≥ field B, the Boolean scalar of A is 1, and the Boolean scalar of B is 0.

[0114] The categories of business rules that real data needs to conform to include category mapping rules. One-hot encoding is performed on the real data to obtain sparse vectors. The sparse vectors are essentially encoded using the one-hot encoding rule.

[0115] The categories of business rules that real data must comply with include numerical range rules. Real data is then input into the interval calculator and compliance calculator in sequence to obtain a probability vector.

[0116] It should be noted that when you input the range calculator and compliance calculator in sequence, the actual process is to first extract the numerical range from the real data, and then take a fixed value to indicate that the input data is within the numerical range.

[0117] In one embodiment, the method for constructing a rule-vector mapping engine further includes:

[0118] Inject boundary value constraints (such as "Daily inbound calls to the 12345 complaint hotline = historical peak value × 1.35") to make the generator prioritize sampling boundary data;

[0119] It should be noted that by inputting boundary data such as peak values ​​and boundary value constraints, the generator can simulate and generate a large number of boundary values ​​for specific tests. Because the CTGAN algorithm samples according to a normal distribution by default, boundary values ​​are inherently scarce in the data. Without specific control, the generated data will also have relatively few boundary values, making it impossible to simulate special test scenarios with a large number of boundary values.

[0120] In one embodiment, based on the categories of business rules that the real data needs to conform to, the encoding vectors corresponding to each category are obtained, and the method further includes:

[0121] The categories of business rules that real data needs to comply with include multi-condition rules (AND / OR). The conditions corresponding to the real data are encoded into binary vectors layer by layer to obtain multi-layer binary vectors.

[0122] This invention also provides an electronic device, which includes, but is not limited to:

[0123] Memory, used to store programs;

[0124] The processor is used to execute programs stored in memory. When the processor executes the programs stored in memory, the processor is used to deploy one of the CTGAN network structures described above.

[0125] The processor and memory can be connected via a bus or other means.

[0126] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs, such as the network structure described in the embodiments of this invention. The processor deploys the aforementioned network structure by running the non-transitory software programs and instructions stored in the memory.

[0127] The memory may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store the network architecture deployed as described above. Furthermore, the memory may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0128] The non-transitory software programs and instructions required to implement the above-described terminal selection method are stored in memory and, when executed by one or more processors, deploy the above-described network structure.

[0129] This invention also provides a storage medium storing computer-executable instructions for deploying the aforementioned network structure.

[0130] In one embodiment, the storage medium stores computer-executable instructions that are executed by one or more control processors.

[0131] The embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0132] Those skilled in the art will understand that all or some of the network structures disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically include computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0133] This document describes embodiments of the invention, including preferred embodiments known to the inventors for carrying out the invention. Variations of these embodiments will become apparent to those skilled in the art upon reading the foregoing description. The inventors encourage those skilled in the art to adopt such variations as appropriate, and the inventors intend to practice embodiments of the invention in ways other than those specifically described herein. Therefore, the scope of the invention includes all modifications and equivalents of the subject matter set forth in the appended claims, as permitted by applicable law. Furthermore, the scope of the invention covers any combination of the foregoing elements in all possible variations thereof, unless otherwise indicated herein or otherwise clearly contradicted by the context.

Claims

1. A CTGAN network structure, characterized in that, include: The business rules database is connected to the rule-vector mapping engine and the discriminator, respectively. The rule-vector mapping engine has one end for obtaining the labels of the data to be generated and connecting to the business rule database, and the other end for connecting to the generator. It is used to obtain the rule encoding vector based on the labels of the data to be generated and the business rule database, and output the rule encoding vector to the generator. The generator is connected to the rule-vector mapping engine at one end and to the discriminator at the other end. It is used to receive the rule encoding vector output by the rule-vector mapping engine, obtain constraint compliance data according to the rule encoding vector, and output the constraint compliance data to the discriminator. The discriminator is connected to the generator and the business rule database at one end, and outputs data at the other end. It is used to receive the constraint compliance data sent by the generator, determine the compliance data in the constraint compliance data according to the business rule database and the constraint compliance data, and output the compliance data, which is the generated data.

2. The CTGAN network structure according to claim 1, characterized in that, Also includes: The association graph, connected at one end to the discriminator and outputting data at the other end, is used to perform association-preserving privacy enhancement processing on the compliant data sent by the discriminator, specifically including: Detect sensitive fields in the compliance data; The sensitive fields are associated with the corresponding sensitive entities; Privacy noise is added to the detected sensitive fields and the sensitive fields associated with the sensitive entities; wherein the intensity of the privacy noise is adjusted according to the sensitivity level of the corresponding sensitive field. The processed data is then output, and the output data is the generated data.

3. The CTGAN network structure according to claim 2, characterized in that, The intensity of the privacy noise is adjusted according to the sensitivity level of the corresponding sensitive field, specifically including: The formula for calculating the intensity of the privacy noise is as follows: ∈ i =α·log(sensitivityLevel) i +1), where, ∈ i Let α be the intensity of the privacy noise corresponding to the i-th sensitive field, and α be a preset parameter, sensitivityLevel. i Let be the sensitivity level of the i-th sensitive field.

4. The CTGAN network structure according to claim 1, characterized in that, The discriminator includes a compliance discrimination branch and an authenticity discrimination branch; The discriminator outputs compliance data based on the business rule base and the constraint compliance data, including: The discriminator inputs the constraint compliance data into the associated compliance discrimination branch; In the related compliance determination branch, the business rule base is searched according to the tags of the constraint compliance data to obtain the search results; Based on the search results, determine whether the data value of the constraint compliance data conforms to the business rules. If yes, it is considered compliant data; otherwise, it is considered non-compliant data. The subset of the data that conforms to the rules and the data output by the authenticity discrimination branch is used as the compliant data output.

5. A CTGAN network structure according to claim 4, characterized in that, The CTGAN network structure further includes: a loss calculation module, one end of which is connected to the discriminator and the other end of which is connected to the generator, used to calculate the loss function of the compliant data output by the discriminator and adjust the parameters of the generator according to the loss function; The loss function includes a constraint compliance loss term to drive the generator to output constraint compliance data that conforms to business rules; the formula for the constraint compliance loss term is as follows: L constraint = 1 - JointComplianceRate, where L constraint The constraint compliance loss term is defined as JointComplianceRate, which is the joint compliance rate output by the associated compliance discrimination branch.

6. The CTGAN network structure according to claim 1, characterized in that, The rule-vector mapping engine obtains rule encoding vectors based on the labels of the data to be generated and the business rule base, specifically including: Based on the tags corresponding to the required number of data to be generated, the business rule base is searched, and the business rules corresponding to the tags of the required data to be generated are retrieved to obtain the tag-business rule correspondence. Based on the correspondence between the tag and the business rule, the corresponding business rule is mapped to the corresponding rule encoding vector, and a correspondence between the tag and the rule encoding vector is formed.

7. A CTGAN network structure according to claim 1, characterized in that, The construction method of the rule-vector mapping engine includes: Obtain the real dataset; Based on the business rule base, determine the business rules that the real data in the real dataset must conform to; Based on the categories of business rules that the real data needs to conform to, the corresponding encoding vectors for each category are obtained; The real data, the rule encoding vector corresponding to the real data, and the correspondence between the real data and the rule encoding vector corresponding to the real data are stored in the rule-vector mapping engine. Based on the categories of business rules that the real data needs to conform to, the corresponding encoding vectors for each category are obtained, including: Boolean scalar, sparse vector, and probability vector; The Boolean scalar, the sparse vector, and the probability vector are concatenated to obtain the rule encoding vector corresponding to the real data.

8. A CTGAN network structure according to claim 7, characterized in that, The step of obtaining the corresponding encoding vector for each category of business rules that the real data needs to conform to specifically includes: extracting fields from the real data for categories of business rules that the real data needs to conform to, including inequality rules; and performing comparison operations based on the extracted fields to obtain the Boolean scalar. The categories of business rules that the real data needs to conform to include category mapping rules are used to encode the real data using one-hot encoding to obtain the sparse vector; The categories of business rules that the real data needs to comply with include numerical range rules. The real data is then input into the interval calculator and the compliance calculator in sequence to obtain the probability vector.

9. An electronic device, characterized in that, include: Memory, used to store programs; A processor for executing a program stored in the memory, wherein when the processor executes the program stored in the memory, the processor is configured to deploy a network structure as described in any one of claims 1 to 8.

10. A storage medium, characterized in that, It stores computer-executable instructions for deploying a network structure as described in any one of claims 1 to 8.