Data synthesis method, system, device, storage medium and program product

CN122795960APending Publication Date: 2026-09-22ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610896210.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

该类方法通常需要对高敏感度数据进行大规模传输与存储,带来较高的运维成本与合规风险;同时,由于生成过程往往依赖于在某一时间点形成的数据快照,合成数据集难以保持时效性,从而影响下游分析或模型训练效果

Benefits of technology

[0012]由上述实施例可知,本说明书提供的数据合成系统包括数据合成层、中间代理层以及多个数据源。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122795960A_ABST
    Figure CN122795960A_ABST
Patent Text Reader

Abstract

The specification provides a data synthesis method, system, device, storage medium and program product. The data synthesis system comprises a data synthesis layer, an intermediate agent layer and a plurality of data sources. The intermediate agent layer establishes a logical connection with each data source to obtain metadata, generates a unified logical view through semantic alignment and provides the unified logical view to the data synthesis layer. The data synthesis layer sends a data synthesis request according to the view, and indicates a target logical field set to be synthesized and a required statistical feature type. The intermediate agent layer determines a target data source based on the target logical field, generates a corresponding statistical request and delivers the statistical request; each target data source returns statistical feature data after executing the statistical request; the intermediate agent layer forwards the statistical feature data to the data synthesis layer, and the data synthesis layer generates a synthesized data set according to the statistical feature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of data synthesis technology, and more particularly to a data synthesis method, a data synthesis system, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] With the continuous advancement of data elementization and AI-driven innovation, the value of data often needs to be fully realized through cross-system flow and in-depth utilization, such as in scenarios like AI model training, business analysis, and system testing. Such utilization typically requires data to have sufficient scale and coverage, and to reflect the latest state of the business.

[0003] However, the flow and reuse of data are significantly constrained in reality: raw data often contains personal privacy information and trade secrets, and due to compliance requirements and inter-organizational authority boundaries, its secure sharing and cross-departmental / cross-organizational flow are difficult, easily leading to data silos. In this situation, data elements cannot achieve efficient flow while meeting governance compliance requirements, thus affecting the sustainability and stability of data supply.

[0004] To alleviate these issues, one common industry practice is to migrate or replicate real data to an isolated environment, and then generate or process the data offline. This approach typically requires large-scale transmission and storage of highly sensitive data, leading to high operational costs and compliance risks. Furthermore, because the generation process often relies on data snapshots taken at a specific point in time, the synthesized datasets are difficult to keep up with, thus affecting the effectiveness of downstream analysis or model training.

[0005] Another common practice is to anonymize the real data. However, anonymized data may still pose privacy risks under conditions of multi-source fusion, statistical correlation, or background knowledge supplementation. At the same time, anonymization often destroys the true distribution of some fields in the original data and the statistical relationships between fields, reducing the usability of the data for real business scenarios and affecting the effectiveness of data use. Summary of the Invention

[0006] In view of the above, one or more embodiments of this specification provide the following technical solutions: According to a first aspect of one or more embodiments of this specification, a data synthesis system is proposed, including a data synthesis layer, an intermediate proxy layer, and multiple data sources; The intermediate proxy layer is used to establish logical connections with the multiple data sources respectively, obtain the metadata of each data source, the metadata being used to describe the field attributes and inter-field relationships of the fields contained in the corresponding data source; based on the metadata of each of the multiple data sources, semantically align the fields in the multiple data sources, determine multiple field sets respectively composed of semantically equivalent fields, generate corresponding logical fields for each field set, and determine the logical relationships between different logical fields based on the inter-field relationships described by the metadata; and generate a unified logical view based on the logical fields and the logical relationships and provide it to the data synthesis layer. The data synthesis layer is used to send a data synthesis request generated according to the unified logical view to the intermediate proxy layer; the data synthesis request is used to indicate the target logical field set to be synthesized and the statistical feature type. The intermediate proxy layer is used to determine the target data source among the multiple data sources based on the target logical field set, and to generate a statistical request for the target data source corresponding to the statistical feature type; The target data source is used to execute the statistical request, obtain and return statistical feature data to the intermediate proxy layer; The intermediate proxy layer is used to send the statistical feature data to the data synthesis layer; The data synthesis layer is used to generate a synthetic dataset that conforms to the statistical characteristics.

[0007] According to a second aspect of one or more embodiments of this specification, a data synthesis method is proposed, applied to an intermediate proxy layer in the data synthesis system described in the first aspect, the method comprising: Logical connections are established with multiple data sources respectively to obtain metadata of each data source. The metadata is used to describe the field attributes of the fields contained in the corresponding data source and the relationship between the fields. Based on the metadata of the multiple data sources, the fields in the multiple data sources are semantically aligned to determine multiple field sets composed of semantically equivalent fields, and corresponding logical fields are generated for each field set. Based on the field associations described by the metadata, the logical relationships between different logical fields are determined, and a unified logical view is generated based on the logical fields and the logical relationships and provided to the data synthesis layer. Receive the data synthesis request generated by the data synthesis layer according to the unified logical view, wherein the data synthesis request is used to indicate the target logical field set to be synthesized and the statistical feature type; Based on the target logical field set, a target data source among the multiple data sources is determined, and a statistical request corresponding to the statistical feature type is generated for the target data source and sent to the target data source; The system receives statistical feature data returned by the target data source after executing the statistical request, and forwards it to the data synthesis layer so that the data synthesis layer can generate a synthetic dataset that conforms to the statistical feature data.

[0008] According to a second aspect of one or more embodiments of this specification, a data synthesis method is proposed, applied to the data synthesis layer of the data synthesis system described in the first aspect, the method comprising: Obtain a unified logical view provided by the intermediate proxy layer. The unified logical view consists of logical fields generated by the intermediate proxy layer based on metadata from multiple data sources and the logical relationships between different logical fields. Receive a data synthesis request generated according to the unified logical view, the data synthesis request being used to indicate the target logical field set to be synthesized and the statistical feature type; The data synthesis request is sent to the intermediate proxy layer; Receive statistical feature data forwarded by the intermediate proxy layer, wherein the statistical feature data is returned by the target data source after executing a statistical request corresponding to the data synthesis request; Generate a synthetic dataset that conforms to the statistical characteristics.

[0009] According to a fourth aspect of the embodiments of this specification, an electronic device is provided, comprising: processor; Memory used to store processor-executable instructions; Wherein, when the processor executes the executable instructions, it is used to implement the method described in the second aspect or the third aspect.

[0010] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in the second or third aspect.

[0011] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the method described in the second or third aspect.

[0012] As can be seen from the above embodiments, the data synthesis system provided in this specification includes a data synthesis layer, an intermediate proxy layer, and multiple data sources.

[0013] First, by generating a unified logical view based on the metadata of each data source through the intermediate proxy layer, the data synthesis layer can initiate a data synthesis request through a unified logical field without being aware of the differences in the physical structure and field naming conflicts of the multiple underlying data sources, thus reducing the complexity and integration cost of using multi-source heterogeneous data.

[0014] Second, since the intermediate proxy layer only sends statistical requests for statistical feature data to the target data source, the granularity of the request does not include the original data records. This means that the original data records do not need to leave their data source during the exploration phase, effectively avoiding the risk of data leakage and compliance costs caused by physical migration or copying of data.

[0015] Third, the target data source returns real-time or near-real-time statistical feature data after executing the statistical request. The synthetic dataset generated by the data synthesis layer can reflect the data distribution characteristics at the current moment, avoiding the problem of the synthesis result being out of sync with the actual data distribution due to reliance on static data snapshots.

[0016] Fourth, the synthetic dataset generated by the data synthesis layer does not contain the original records. This ensures that the synthetic dataset retains statistical usability while eliminating the risk of directly exposing real individual information at the record level. As a result, it can provide data with a higher security boundary for downstream analysis, testing, and model training.

[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description

[0018] Figure 1 This is a structural diagram of a data synthesis system provided in an exemplary embodiment.

[0019] Figure 2 This is an exemplary embodiment of the data synthesis layer, intermediate proxy layer, and interaction diagram between multiple data sources.

[0020] Figure 3 This is another structural diagram of a data synthesis system provided in an exemplary embodiment.

[0021] Figure 4 This is a flowchart of a data synthesis method provided in an exemplary embodiment.

[0022] Figure 5 This is a flowchart of another data synthesis method provided in an exemplary embodiment.

[0023] Figure 6 This is a schematic diagram of the structure of a device provided in an exemplary embodiment. Detailed Implementation

[0024] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0025] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with relevant laws, regulations and standards, and corresponding operation entry points shall be provided for users to choose to authorize or refuse.

[0026] In multi-source heterogeneous data environments, relevant technical solutions generally follow the technical paradigm of "acquiring or modifying raw data for use." The core assumption is that statistical characteristic data acquisition or subsequent data generation tasks can only be completed with direct access to or possession of the original data records. Under this paradigm constraint, the improvement ideas of those skilled in the art typically focus on: optimizing data replication and migration efficiency, enhancing the irreversibility of de-identification algorithms, or improving the performance of centralized modeling, etc., but it is difficult to escape the direct dependence on the original data ontology.

[0027] However, the aforementioned paradigm faces inherent contradictions in scenarios with multi-source heterogeneity and strong compliance constraints: the deeper the physical access to the original data, the greater the difficulty in managing security and compliance boundaries and the higher the risk of leakage; while the anonymization or truncation processes applied to reduce risks often come at the cost of sacrificing the integrity of the statistical distribution of data and the fidelity of the relationships between fields. This makes it difficult for related technologies to dynamically and continuously provide a supply of high-quality synthetic data that truly reflects the current data distribution characteristics while strictly adhering to data security requirements.

[0028] Based on this, this specification proposes a data synthesis system. By introducing an intermediate proxy layer, it completes field semantic alignment and logical modeling based solely on the metadata of each data source without obtaining the original data content, constructing a unified logical view. Furthermore, based on this unified logical view, data requirements are transformed into statistical requests for each data source, which then returns statistical features. This generates a synthetic dataset conforming to these statistical features at the data synthesis layer. This eliminates the reliance on centralized processing of the original data, achieving data feature utilization and data generation while ensuring data security and compliance.

[0029] In some embodiments, please refer to Figure 1 This paper provides a schematic diagram of the structure of a data synthesis system. The data synthesis system includes a data synthesis layer 30, an intermediate proxy layer 20, and multiple data sources 10.

[0030] (1) Regarding the intermediate proxy layer 20. In the system architecture provided in the embodiments of this specification, the intermediate proxy layer 20 logically assumes the role of a trusted third party. For example, the intermediate proxy layer 20 establishes secure connections with multiple data sources 10 and with the data synthesis layer 30 based on preset trust relationships.

[0031] For the data source 10, the intermediate proxy layer 20 is an authorized and legitimate access point. The statistical query requests issued by it are limited to the aggregation and statistical level and are subject to security policies. Based on this, the data source 10 can be sure that the original data records will not be leaked to the outside through the intermediate proxy layer 20.

[0032] For the data synthesis layer 30, the intermediate proxy layer 20 is the only channel for it to obtain multi-source statistical feature data. Moreover, this channel has embedded desensitization processing and permission verification mechanisms during execution. The data synthesis layer 30 does not need to, and cannot, directly access the physical storage of any original data source 10.

[0033] By positioning the intermediate proxy layer 20 as a trusted coordination layer independent of the data source 10 and the data users, this embodiment achieves unified scheduling of statistical feature exploration and synthesis requests for multi-source heterogeneous data without requiring direct sharing of raw data between data sources 10 or between data source 10 and data synthesis layer 30. The establishment of this trust model makes "non-intrusive exploration" and "dynamic synthesis" practically feasible: data source 10 only needs to open its statistical query capabilities to the trusted intermediate proxy layer 20, while data synthesis layer 30 only needs to initiate a declarative request to the unified logical view provided by the intermediate proxy layer 20 to obtain a synthesized dataset reflecting the latest data distribution characteristics.

[0034] (2) Regarding data source 10.

[0035] In the embodiments of this specification, data source 10 refers to various systems or platforms that store or manage raw data, such as relational databases, non-relational databases, data warehouses, data lakes, and data interfaces of business systems. Each data source 10 is independently managed by its respective organization or department and is subject to corresponding data security policies and access control.

[0036] In traditional data utilization models, data source 10 often passively cooperates with data extraction, copying, or export operations to support downstream analysis and modeling needs. However, in the system architecture provided in the embodiments of this specification, the role of data source 10 has undergone a fundamental change: data source 10 no longer needs to provide raw data records externally, but only needs to respond to statistical query requests issued by the intermediate proxy layer 20 and return aggregated statistical feature data to it. The raw data records always reside locally on data source 10 and do not undergo cross-domain transfer.

[0037] Based on the pre-established trust relationship with the intermediate proxy layer 20, the data source 10 can regard the intermediate proxy layer 20 as an authorized secure access point, trusting that the statistical query requests issued by it have been verified by security policies and will not exceed the constraints of field-level permissions and probing granularity. This mechanism enables the data source 10 to provide necessary statistical feature support for cross-source data fusion and synthetic data generation while ensuring the security and compliance of the original data.

[0038] (3) Regarding the data synthesis layer 30.

[0039] In the embodiments described in this specification, the data synthesis layer 30 is the initiator of the request for synthesized data. The data synthesis layer 30 can be an interactive platform for data analysts, AI engineers, or business personnel, or it can be a data synthesis task module in an automated scheduling system. The core requirement of the data synthesis layer 30 is to acquire a high-fidelity synthesized dataset that reflects the current business data distribution characteristics and does not contain real individual information, for use in downstream applications such as model training, algorithm verification, system stress testing, or business simulation.

[0040] In the traditional model, the data synthesis layer 30 typically needs to first obtain a copy of the original data and then execute the synthesis algorithm based on the copy; or it can directly process and use the de-identified real data. Regardless of the method, the data synthesis layer 30 needs to directly or indirectly access the original data records.

[0041] In the system architecture provided in the embodiments of this specification, the role of the data synthesis layer 30 is redefined: the data synthesis layer 30 does not need to be aware of the physical location, storage structure, or access method of the underlying data source 10, nor does it need to directly process any raw data records. The data synthesis layer 30 only needs to construct data synthesis requests in a declarative manner, using logical fields as the granularity, based on the unified logical view provided by the intermediate proxy layer 20. After receiving real-time statistical feature data forwarded from the intermediate proxy layer 20, the data synthesis layer 30 uses it as conditional input to drive the synthetic data generation model, thereby generating a synthetic dataset that meets the requirements.

[0042] Since the data synthesis layer 30 only deals with logical views and statistical features throughout the process, without involving the original data records, its responsibility boundaries in terms of data security and compliance are clearly defined, thus enabling it to focus more on realizing the value of the synthesized data in downstream applications.

[0043] Please see Figure 1 as well as Figure 2 The interaction process between the data synthesis layer 30, the intermediate proxy layer 20, and the data source 10 will be illustrated by an example below: During the generation of the unified logical view, the intermediate proxy layer 20 performs the following operations: First, the intermediate proxy layer 20 establishes logical connections with multiple data sources 10 respectively (S20). Here, logical connection means that the intermediate proxy layer 20 establishes a communication link through the standard data access interface provided by each data source 10, such as JDBC / ODBC interface, RESTful API interface, GraphQL interface or message queue subscription interface, which enables it to obtain the metadata of the data source 10 and execute statistical queries, without physically migrating or copying the original data records in the data source 10 to the local intermediate proxy layer 20.

[0044] After establishing a logical connection, the intermediate proxy layer 20 obtains metadata from the system directory table, schema information table, or dedicated metadata service of each data source 10 (S22). The metadata of each data source 10 is used to describe the field attributes and relationships between the fields contained in the corresponding data source 10. The metadata includes, but is not limited to: field name, field data type, field length or precision, whether it can be nullable, field comment description, primary key constraint, foreign key constraint, unique constraint, and the relationships between views or tables defined in the system of the data source 10.

[0045] Subsequently, the intermediate proxy layer 20 performs semantic alignment processing based on the acquired metadata. The purpose of semantic alignment is to resolve the inconsistency in the semantic representation of fields among multiple data sources 10 due to differences in naming conventions and design standards, thereby identifying and merging fields that actually express the same or highly similar business meanings.

[0046] In one alternative implementation, semantic alignment can be achieved through a combination of one or more of the following methods: First, text similarity is based on field names. For example, the intermediate proxy layer 20 extracts field name strings from each data source 10, removes predefined prefixes, suffixes, or separators, and calculates edit distance, Jaccard similarity, or semantic similarity based on word vectors. For instance, it uses a pre-trained word embedding model to convert field names into vectors and then calculates cosine similarity. When the similarity between two field names exceeds a preset threshold, they are identified as candidate semantically equivalent fields.

[0047] Secondly, based on the overlap of field value range distributions. For example, without obtaining the original data record details, the intermediate proxy layer 20 uses lightweight statistical probing, such as requesting the data source 10 to return the deduplication count, maximum and minimum values ​​of the field, or a list of high-frequency values ​​for enumerated type fields, to calculate the intersection and union ratio of the two field value ranges. When the overlap exceeds a preset threshold, the two are identified as candidate semantically equivalent fields.

[0048] Third, the similarity is based on the metadata annotation information of the fields in their respective data source 10. For example, the intermediate proxy layer 20 extracts the field annotation text and column description information, uses natural language processing technology to extract keywords or perform vectorization, and calculates the similarity between the annotation texts. When the similarity exceeds a preset threshold, the two are identified as candidate semantically equivalent fields.

[0049] In practical implementation, the intermediate proxy layer 20 can integrate the above-mentioned multiple identification results, such as by using a weighted voting mechanism or a rule engine, to cluster fields identified as semantically equivalent, and finally determine multiple field sets, where each field set is semantically equivalent to the others. The above semantic alignment process can be executed fully automatically, or it can be manually confirmed or adjusted by the administrator through the management interface.

[0050] For fields belonging to the same field set, the intermediate proxy layer 20 generates corresponding logical fields. The generation process of logical fields includes: extracting common field attributes of each field in the field set and forming a unified field representation. Field attributes may include: ① logical field name, such as the name that appears most frequently in the field set, or a standard name generated based on the descriptive information in the metadata annotation; ② data type, such as selecting a compatible data type that can accommodate the value range of all physical fields without loss when there are differences in the data types of the fields in the set, such as unifying multiple numeric types into a double-precision floating-point type; ③ value range, such as taking the union of the value ranges of each physical field as the value range of the logical field; ④ field description information, such as merging the annotation information of each physical field or selecting representative descriptions. The logical fields generated in this way serve as a unified abstraction of multi-source heterogeneous physical fields, shielding the upper-layer data synthesis layer 30 from the structural differences of the underlying physical data source 10.

[0051] Furthermore, the intermediate proxy layer 20 determines the logical relationships between different logical fields based on the field associations described by the metadata of each data source 10. These logical relationships are used to express structural constraints or statistical dependencies between logical fields in a unified logical view.

[0052] In one alternative implementation, the logical relationships can be determined in one or more of the following ways: In the first approach, the explicit primary and foreign key relationships declared in the metadata are converted into logical relationships between logical fields. When two physical fields belong to different data sources 10 and each has a foreign key association with a certain common dimension table, the intermediate proxy layer 20 can deduce that there is an indirect association between the two logical fields and mark it in the logical view.

[0053] In the second approach, logical relationships are inferred based on the statistical correlation between field value ranges. For example, the intermediate proxy layer 20 can send a statistical query request to the relevant data source 10 to obtain the Pearson correlation coefficient, mutual information value, or chi-square test statistic of the two fields. When the correlation measure exceeds a preset threshold, it is determined that there is a statistical dependency between the two logical fields.

[0054] In the third approach, functional dependencies or business constraints between fields are extracted from metadata. For example, if the metadata of a data source 10 defines a calculation expression or validation rule between field A and field B, the intermediate proxy layer 20 can map this functional dependency into a constraint relationship between the corresponding logical fields.

[0055] In the fourth approach, the temporal relationship between fields is extracted from the metadata relationship of the timestamp field. For example, if a data source 10 defines the order constraint between the order creation time field and the order completion time field, the intermediate proxy layer 20 can record this temporal relationship as the temporal logical relationship between the corresponding logical fields.

[0056] After determining the logical fields and logical relationships, the intermediate proxy layer 20 generates a unified logical view representing multiple data sources 10 (S24). The unified logical view can be physically represented using a relational schema, such as a virtual table where each logical field corresponds to a column, and logical relationships are described by foreign key connections or associated metadata; or it can be represented using a graph structure, such as nodes representing logical fields and edges representing logical relationships. The unified logical view is provided to the data synthesis layer 30 (S26), which can obtain the view definition by calling the standard query interface provided by the intermediate proxy layer 20.

[0057] At the device implementation level, the intermediate proxy layer 20 can be deployed as an independent server process or a containerized service cluster, connected to the network domain where each data source 10 resides via a security gateway or dedicated line. Internally, the intermediate proxy layer 20 can be divided into multiple functional modules, such as a metadata collection module, a semantic alignment engine, a logical view builder, and a statistical query scheduler. These modules collaborate via an internal message bus or remote procedure calls. Furthermore, the intermediate proxy layer 20 can maintain a persistent storage area to save the generated unified logical view definitions, metadata snapshots, and semantic alignment results, supporting retrospective queries of historical view versions.

[0058] Through the above process, the intermediate proxy layer 20 constructs a unified logical view across multiple heterogeneous data sources 10 without physically aggregating the original data records, providing a structured semantic foundation for subsequent non-intrusive statistical feature exploration and dynamic synthesis.

[0059] The following section provides a detailed explanation of the process by which the intermediate proxy layer 20 generates a unified logical view, using a specific example.

[0060] Suppose there are three data sources 10, namely data source A, data source B and data source C.

[0061] Data source A is a customer relationship management system in a relational database, containing a table customer_info with the following fields: cust_id (integer, primary key), cust_name (string), birth_date (date), mobile_phone (string), and reg_time (timestamp).

[0062] Data source B is the order analysis system in the data warehouse, which contains the table order_record, whose fields include: order_id (integer, primary key), user_id (integer, foreign key related to the user table), order_amt (floating-point), and order_time (timestamp).

[0063] Data source C is a user profile storage in a NoSQL database, containing a collection called user_profile, whose fields include: uid (string), full_name (string), birthday (string, format YYYY-MM-DD), phone_number (string), and create_timestamp (integer timestamp).

[0064] The fields in the three data sources 10 mentioned above differ in name, type, and format, but some fields are semantically equivalent or related.

[0065] Step 1: The intermediate proxy layer 20 establishes logical connections through the standard interfaces provided by each data source 10 to obtain metadata.

[0066] For data source A, connect via JDBC interface and retrieve metadata of the customer_info table from the database system directory, including field names, data types, primary key constraints, field comments, etc.

[0067] For data source B, the metadata and foreign key constraint information of the order_record table are also obtained through the JDBC interface.

[0068] For data source C, obtain the schema definition information of the user_profile collection through the RESTful API.

[0069] The obtained metadata, after parsing, is recorded as shown in Table 1: Table 1 Step 2: The intermediate proxy layer 20 starts the semantic alignment engine, which combines the similarity of field names, the overlap of value range distribution, and the similarity of metadata annotations to perform semantic clustering on the above fields and determine the field set.

[0070] Taking the "customer unique identifier" related field as an example: data source A's cust_id (annotated "customer unique identifier"), data source B's user_id (annotated "order user ID"), and data source C's uid (annotated "user unique identifier") all contain "id", "user_id", and "uid" in their names respectively. Through word form restoration and synonym mapping, such as "cust", "user", and "uid" all pointing to user entity identifiers, the calculated name similarity exceeds a preset threshold.

[0071] The intermediate proxy layer initiates lightweight probing to three data sources: it retrieves the deduplication count and value range for A's cust_id, B's user_id, and C's uid. For example, A returns a value range of [10001, 99999], B returns [10005, 98800], and C returns a string prefixed with "UID_". After removing the prefix, the numerical portion highly overlaps with A and B. The value range overlap calculation results show that the overlap rate of the three exceeds a preset threshold.

[0072] After comprehensive judgment, cust_id, user_id, and uid are grouped into the same field set, which is denoted as field set 1.

[0073] Similarly, semantic alignment is performed on other fields to obtain the field set shown in Table 2: Table 2 Step 3: For each set of fields, the intermediate proxy layer 20 extracts the common attributes of each physical field to form the corresponding logical fields and their field attributes.

[0074] Taking field set 1 (user identifier) ​​as an example: Logical field name: Based on the most frequent description in the comment information, it is determined to be user_id.

[0075] Data type: The physical field types within the collection are INT (A, B) and STRING (C). To ensure lossless storage, a general type compatible with both integer and string representations is selected, such as VARCHAR(64).

[0076] Value range: Take the union of the value ranges of each physical field and record the format mapping rules, such as automatically removing the prefix "UID_" from the value of C to align with other data sources.

[0077] Field description information: Merging all comments yields the "user's unique identifier".

[0078] The resulting logical field is denoted as LV_user_id, and its field attributes are recorded in Table 3: Table 3 Similarly, other logical fields are generated, as shown in Table 4: Table 4 Step 4: The intermediate proxy layer 20 infers the logical relationships between different logical fields based on the associations and statistical correlations recorded in the metadata.

[0079] (1) Primary and foreign key relationships based on metadata: In data source B, order_record.user_id is a foreign key, which is associated with the primary key of the user table. Combining the semantic alignment results, user_id corresponds to the logical field LV_user_id, while the order table itself corresponds to the logical fields LV_order_amt and LV_order_time. Therefore, a logical relationship is established: there is a one-to-many relationship between LV_user_id and LV_order_amt and LV_order_time (one user corresponds to multiple orders).

[0080] In data source A, cust_id is the primary key and has a one-to-one relationship with cust_name, birth_date, etc. After mapping, it is confirmed that there is a one-to-one relationship between LV_user_id and LV_user_name, LV_birth_date, LV_phone, and LV_create_time.

[0081] (2) Logical relationship inferred from statistical correlation: The intermediate proxy layer 20 initiates a statistical query to data source B to calculate the Pearson correlation coefficient between order_amt and order_time (e.g., if it returns 0.12, which is below the threshold), and determines that there is no strong correlation; however, it explores the time sequence constraints between birth_date and reg_time / create_timestamp in data sources A and C respectively. For example, through metadata annotations and value range comparison, it is determined that birth_date should be earlier than create_time. Therefore, a logical relationship is established: there is a time sequence constraint relationship between LV_birth_date and LV_create_time (the former is not later than the latter).

[0082] The logical relationships of the final records are shown in Table 5: Table 5 Step 5: The intermediate proxy layer 20 integrates the above logical fields and logical relationships to generate a unified logical view. In an optional implementation, this view is represented by the relational schema shown in Table 6: View name: unified_user_order_view Table 6 Logical relationships within the view: user_id is the main entity identifier, and together with user_name, birth_date, phone, and create_time, they form a one-to-one attribute group.

[0083] user_id, order_amt, and order_time form a one-to-many parent-child relationship, and can be used to perform join queries through user_id.

[0084] There is a time-series constraint between birth_date and create_time, and this constraint must be satisfied when generating composite data.

[0085] This unified logical view is provided to the data composition layer 30 in the form of a metadata description file (such as a JSON Schema or SQL view definition). In subsequent operations, the data composition layer 30 only needs to initiate a data composition request based on the logical field names in this view (such as user_id, order_amt), without needing to concern itself with the physical distribution and field differences of the underlying data sources A, B, and C.

[0086] As can be seen from the above example, the intermediate proxy layer 20 integrates the original fields from the three heterogeneous data sources 10 into a unified logical view with a clear structure and consistent semantics through semantic alignment, logical field generation, and logical relationship determination, providing a standardized data description foundation for subsequent non-intrusive statistical exploration and dynamic synthesis.

[0087] After obtaining the unified logical view provided by the intermediate proxy layer 20, the data synthesis layer 30 presents the unified logical view to the requester 40 with data synthesis needs in an interactive form (S26). In an optional implementation, the data synthesis layer 30 provides a graphical interface or application programming interface to display the logical field names, data types, field descriptions, and logical relationships in the unified logical view in the form of a directory tree or relationship diagram. The requester 40 can construct a data synthesis request and send it to the data synthesis layer 30 by selecting logical fields, selecting the required statistical feature types, and setting the generation quantity parameters (S28). After receiving the request, the data synthesis layer 30 encapsulates it into a message body with a preset format and sends it to the intermediate proxy layer 20 (S30).

[0088] Data synthesis requests include, but are not limited to, the following information: (1) Target logical field set. The target logical field set is used to indicate the logical fields that should be included in this synthesized dataset. Each field in the target logical field set is selected from the logical fields in the unified logical view. Before sending the request, the data synthesis layer 30 can verify whether the selected fields exist in the currently valid version of the unified logical view.

[0089] (2) Required statistical feature types. The statistical feature types are used to indicate which categories of statistical features the intermediate proxy layer 20 should obtain from the target data source 10 to support subsequent data synthesis. Optionally, the statistical feature types include, but are not limited to: single-field value range, null value ratio, unique value count, frequency distribution histogram; multi-field correlation coefficient matrix, conditional probability distribution; time interval distribution of time series fields; and enumeration constraints or format constraints that field values ​​should satisfy.

[0090] (3) Target Generation Quantity Parameter. The target generation quantity parameter indicates the number of synthetic data records expected to be generated in this synthesis request. This quantity can be less than, equal to or greater than the total number of records in the original data source 10, and does not depend on the physical inventory of the original data records.

[0091] The following example illustrates the process of constructing and sending a data synthesis request.

[0092] For example, suppose a unified logical view contains logical fields user_id (user identifier), user_level (user level), order_amt (order amount), and order_time (order time), and the view has declared a one-to-many relationship between user_id and order_amt and order_time. A party with data synthesis needs wants to obtain a synthetic dataset for training an AI model, which needs to reflect the statistical distribution characteristics of user ordering behavior in the current business.

[0093] In this scenario, the party performs the following operations through the interactive interface provided by the data synthesis layer 30: (1) In the field list of the unified logical view, select the target logical fields: user_id, user_level, order_amt, and order_time to form a set of target logical fields; (2) In the statistical feature type options, check: single field frequency distribution histogram (for user_level, order_amt), value range (for order_amt, order_time), and correlation coefficient between fields (for user_level and order_amt). (3) Enter the value “100000” in the quantity parameter to indicate that you expect to generate 100,000 composite records.

[0094] Based on the above operations, the data synthesis layer 30 generates a corresponding data synthesis request. This data synthesis request can be internally represented as a structured data object, for example, described using JSON format.

[0095] The data synthesis layer 30 sends a data synthesis request generated according to the unified logical view to the intermediate proxy layer 20. After receiving the data synthesis request, the intermediate proxy layer 20 first parses the request, extracting the target logical field set, the required statistical feature type, and optional generation quantity parameters contained therein. After parsing, based on the target logical field set, it determines the target data source 10 among multiple data sources 10, and generates a statistical request corresponding to the statistical feature type for the target data source 10 (S32).

[0096] In determining the target data source 10, the intermediate proxy layer 20, based on each target logical field in the target logical field set, traces back the mapping relationship established during the unified logical view generation phase to obtain the field set corresponding to each target logical field. As previously stated, each field set contains multiple semantically equivalent physical fields from different data sources 10. Therefore, for each target logical field, the intermediate proxy layer 20 can obtain a list of data sources 10 associated with that target logical field, which records which data sources 10 contain physical fields semantically equivalent to that target logical field.

[0097] Secondly, the intermediate proxy layer 20, based on the list of data sources 10 associated with each target logical field, determines at least one subset of candidate data sources 10 that can cover all target logical fields in the target logical field set. Here, coverage means that the data sources 10 in the subset of candidate data sources, combined, can provide at least one corresponding physical field for each target logical field in the target logical field set. In one implementation, the intermediate proxy layer 20 can model this problem as a set coverage problem, where each data source 10 corresponds to a subset of the target logical field set it can provide, and the goal is to find the combination of data sources 10 that can completely cover all target logical fields.

[0098] When multiple subsets of candidate data sources 10 can meet the coverage conditions, the intermediate proxy layer 20 determines the final target data source 10 subset from the multiple candidate data source 10 subsets according to a preset selection strategy. The selection strategy prioritizes minimizing the number of target data sources 10, that is, it prioritizes the candidate data source 10 subset containing the fewest data sources 10. The purpose of this is to reduce the number of data sources 10 that need to issue statistical query requests, reduce network communication overhead, shorten the overall response time of statistical probing, and reduce the coordination complexity in the process of collaborative probing of multiple data sources 10.

[0099] When multiple candidate data source subsets 10 still exist that meet the primary objective, the selection strategy further performs a secondary selection based on the priority of each data source 10 within the subset, selecting the subset of data sources 10 with the better overall priority evaluation as the target data source subset 10. In one implementation, the priority of the data source subset 10 can be determined based on a weighted sum or minimum value of the priorities of each data source 10 within the subset. The priority of each data source 10 is pre-evaluated based on at least one of the following dimensions: (1) Data quality score of data source 10, such as a quantitative score calculated based on the quality indicators of data source 10 such as completeness, accuracy and consistency.

[0100] (2) The update timeliness level of data source 10, such as the level divided according to indicators such as data refresh frequency and data delay time of data source 10.

[0101] (3) Access response latency level of data source 10, such as the level divided according to the percentile of response time of historical statistical queries.

[0102] (4) The trust level of the security domain where data source 10 is located, such as the degree of trust based on the network area and security protection level of data source 10.

[0103] The aforementioned priority information can be obtained by the intermediate proxy layer 20 from the metadata provided by the data source 10, or pre-configured by the administrator through the management interface. Through the above-mentioned hierarchical selection strategy, the intermediate proxy layer 20 can prioritize selecting the combination of data sources 10 with the fewest number and the best overall quality, timeliness, response performance, and security and reliability as the target data source 10, while ensuring the integrity of the exploration coverage.

[0104] After identifying the target data source 10, the intermediate proxy layer 20 generates a corresponding statistical request for each target data source 10. The statistical request instructs the target data source 10 to perform local statistical calculations and return statistical feature data corresponding to the target logical field set and the required statistical feature types. The generation process of the statistical request includes the following operations: (1) Mapping and restoration from logical fields to physical fields: For the target data source 10 being processed, the intermediate proxy layer 20 filters out a subset of target logical fields from the target logical field set that can provide physical fields corresponding to the data source 10. For each target logical field in this subset, the intermediate proxy layer 20 determines its corresponding physical field name, physical field data type, and the data table or data set in the target data source 10. This mapping relationship has been established and persisted during the semantic alignment process in the unified logical view generation stage, and can be directly called here.

[0105] (2) Conversion of statistical feature type to query operation: The intermediate proxy layer 20 converts the required statistical feature type indicated in the data synthesis request into a statistical query operation that the target data source 10 can execute. Different types of statistical features correspond to different query primitives.

[0106] For example, for the statistical feature of "value range", the intermediate proxy layer 20 generates query statements to obtain the maximum and minimum values ​​of the physical field, such as the MAX() and MIN() functions in SQL, or the range aggregation command in NoSQL.

[0107] For example, for the statistical feature of "null value ratio", the intermediate proxy layer 20 generates a query statement to calculate the ratio of the number of null value records to the total number of records, such as a combination of COUNT(*) and COUNT(column).

[0108] For example, for the "unique value count" statistical feature, the intermediate proxy layer 20 generates a query statement to calculate the number of duplicate values, such as COUNT(DISTINCT column).

[0109] For example, for the statistical feature of "frequency distribution histogram", the intermediate proxy layer 20 generates a group counting query statement, such as GROUP BY column with COUNT(*), and can specify the upper limit of the number of groups or the group boundary.

[0110] For example, for the statistical feature of "quantiles", the intermediate proxy layer 20 generates a query statement to calculate the specified quantile value, such as the PERCENTILE_CONT() function in SQL or equivalent quantile calculation logic.

[0111] For example, for the statistical feature of "correlation coefficient matrix", when multiple physical fields are involved and the target data source 10 supports covariance or correlation coefficient calculation, the intermediate proxy layer 20 generates a query statement to calculate the Pearson correlation coefficient between fields, such as CORR(column_a, column_b); when the target data source 10 does not support direct calculation of correlation coefficient, the intermediate proxy layer 20 can request intermediate statistics such as the mean, standard deviation and joint moments of each field, and the intermediate proxy layer 20 will aggregate and calculate them on its own.

[0112] (3) Assembly and encapsulation of query statements: The intermediate proxy layer 20 assembles the statistical query operations obtained above into complete query statements or query request message bodies that conform to the query syntax of the target data source 10, based on the access interface type of the target data source 10. For relational data sources 10, it can be assembled into SQL statements; for data sources 10 that provide RESTful APIs, it can be encapsulated into a JSON-formatted query parameter dictionary; for NoSQL data sources 10 that support native aggregation frameworks, it can be assembled into the corresponding aggregation pipeline description. After assembly, the intermediate proxy layer 20 assigns a request identifier to the statistical request and associates the request identifier with the data synthesis request identifier to facilitate subsequent response matching and link tracing.

[0113] Through the above process, the intermediate proxy layer 20 converts the declarative composition request for the logical view into specific statistical query instructions for each target data source 10.

[0114] In some embodiments, please refer to Figure 3The intermediate proxy layer 20 internally maintains or associates a data security policy storage module 21, which contains pre-configured data security policies. The data security policies include at least field sensitivity levels and whitelists of statistical feature types corresponding to each sensitivity level.

[0115] Field sensitivity level is a label indicating the degree of sensitivity of the information contained in a logical field. For example, it can be divided into high sensitivity level, medium sensitivity level and low sensitivity level.

[0116] The whitelist of statistical feature types corresponding to each sensitivity level is used to limit: for logical fields of that sensitivity level, the range of statistical feature types allowed to be requested from data source 10. For example, for logical fields with a high sensitivity level, the corresponding statistical feature type whitelist can be configured to only allow requests for "value range" and "percentage of null values," while prohibiting requests for "frequency distribution histogram" or "unique value list"; for logical fields with a low sensitivity level, the whitelist can contain a more complete set of statistical feature types. The above policies can be configured and updated by the security administrator through the management interface provided by the intermediate proxy layer 20.

[0117] The first operation performed based on the data security policy includes: an admission check on the data synthesis request.

[0118] In one implementation, the intermediate proxy layer 20 performs an admission check on the data synthesis request after parsing it and before determining the target data source 10.

[0119] First, the intermediate proxy layer 20 extracts the set of target logical fields indicated in the data synthesis request. For each target logical field in this set, it queries the data security policy storage module for the sensitivity level corresponding to that logical field.

[0120] Secondly, it is determined whether the sensitivity level allows the request to be made. Under certain configurations, logical fields of a specific sensitivity level can be set to prohibit any form of synthetic request referencing. If any field in the target logical field set is marked as prohibited from requesting, the admission check fails, and the intermediate proxy layer 20 generates and returns a prompt message, which may include the identifier of the rejected field and the reason for the rejection.

[0121] Next, for each requested target logical field, the intermediate proxy layer 20 further obtains a whitelist of statistical feature types corresponding to the sensitivity level of that field, and compares the required statistical feature type indicated in the data synthesis request with this whitelist. If any item in the required statistical feature type is outside the whitelist range, the access check fails, and the intermediate proxy layer 20 returns a corresponding prompt message.

[0122] Only after all fields in the target logical field set have passed the above verification can the intermediate proxy layer 20 proceed to the subsequent target data source 10 determination and statistical request generation stage. Through access control, synthetic intents that do not conform to the security policy can be intercepted at the forefront of request processing, preventing invalid or unauthorized statistical probing requests from being sent to data source 10.

[0123] The second operation performed based on the data security policy includes: constraint processing of statistical requests.

[0124] In another alternative implementation, the intermediate proxy layer 20 performs a second operation during the process of generating statistical requests to be sent to each target data source 10, namely, constraining the statistical requests based on data security policies.

[0125] After the intermediate proxy layer 20 determines the target data source 10 based on the target logical field set, it needs to generate a statistical request corresponding to its physical fields for each target data source 10. This statistical request instructs the target data source 10 to perform local statistical calculations and return the results. During the construction of the statistical request, the intermediate proxy layer 20 imposes at least one of the following constraints based on the data security policy: (1) Limit the fields to be explored to not exceed the range of fields allowed by the statistical feature type whitelist. For the physical fields in the target data source 10 that correspond to the target logical fields, the intermediate proxy layer 20 generates only the statistical query items allowed in the whitelist based on the sensitivity level of the logical field. If the whitelist does not contain a certain statistical feature type, even if the data synthesis request contains that type, the intermediate proxy layer 20 will not generate the corresponding query statement or query parameters in the statistical request sent to the data source 10, thereby narrowing the exploration scope at the exploration source.

[0126] (2) Limiting the granularity of probing to exclude detailed data. When constructing statistical requests, the intermediate proxy layer 20 ensures that the query operations in the request only involve aggregation calculations, such as using aggregation functions like COUNT, SUM, AVG, STDEV, PERCENTILE_CONT, or requesting the return of pre-calculated grouped statistical results, and must not contain SELECT * or any query format that may return details of a single data record. This constraint ensures from the probing mechanism level that the content returned by the data source 10 to the intermediate proxy layer 20 is always aggregated statistical information, and the original data records do not leave the data source 10.

[0127] (3) Restrict the returned results to aggregate statistics only. The intermediate proxy layer 20 may attach return format constraints to the statistical request, requiring the target data source 10 to return only the numerical results after aggregate calculation or the grouped statistical summary, and not to include any auxiliary information that can be used to reconstruct individual records (such as data block identifiers, physical row addresses, etc.).

[0128] Through the above constraint processing, the intermediate proxy layer 20 embeds the execution of security policies during the generation of statistical requests, ensuring that the statistical probing behavior sent to each target data source 10 is consistent with the predefined security rules, and ensuring that the probing process itself does not constitute a channel for the leakage of sensitive information.

[0129] The first and second operations described above can be executed at different stages of the process of determining the target data source 10. The first operation is usually executed before the target data source 10 is determined, and is used to filter synthetic requests that do not comply with the security policy; the second operation is executed after the target data source 10 is determined, during the process of generating statistical requests, and is used to impose fine-grained constraints on specific query content. Either operation can be performed or a combination thereof, and the intermediate proxy layer 20 can be flexibly enabled according to the system configuration.

[0130] At the device implementation level, a security policy execution engine can be set up inside the intermediate proxy layer 20. This engine is connected to the data security policy storage module and is registered in the request processing pipeline as a pre-interceptor (corresponding to the first operation) and a statistics request construction interceptor (corresponding to the second operation). The security policy execution engine performs policy matching and decision-making for each data synthesis request and outputs the allowed, denied, or modified statistics request. This modular design allows security policy updates without modifying the core scheduling logic of the intermediate proxy layer 20, improving the maintainability of the system and the real-time performance of policy response.

[0131] Through the above mechanism, the intermediate proxy layer 20 moves the execution point of the data security policy to the stage of receiving and generating exploration requests, realizing automated security compliance assurance under a non-intrusive exploration architecture, and providing a reliable security foundation for the subsequent acquisition of statistical feature data and the generation of synthetic data.

[0132] In some embodiments, the intermediate proxy layer 20 sends the statistical requests generated for each target data source 10 to that target data source 10 through the established logical connection (S34). In one implementation, the intermediate proxy layer 20 can send statistical requests to multiple target data sources 10 in parallel to reduce the overall probing wait time; in another implementation, when there are dependencies between statistical requests (e.g., it is necessary to first obtain the value range of a certain field to determine the grouping boundary of the histogram), the intermediate proxy layer 20 can send them serially in the order of dependency.

[0133] The target data source 10 is used to execute the received statistical request (S36) and obtain and return statistical feature data to the intermediate proxy layer 20 (S38).

[0134] The target data source 10 performs aggregation calculations on the data records stored locally. For example, when a statistical request indicates the retrieval of a frequency distribution histogram for a certain field, the target data source 10 groups and counts all values ​​for that field, generating the record frequencies corresponding to each value interval or enumerated value; when a statistical request indicates the retrieval of the correlation coefficient between fields, the target data source 10 calculates the covariance and standard deviation of the two fields, thereby obtaining the Pearson correlation coefficient value. The target data source 10 encapsulates the aggregation calculation results into statistical feature data in a preset format and returns it to the intermediate proxy layer 20 through the established logical connection. Throughout this process, the original data records in the target data source 10 remain in the local storage of the target data source 10, without cross-domain transfer or physical copying.

[0135] The intermediate proxy layer 20 is also used to send statistical feature data to the data synthesis layer 30 (S40).

[0136] In another implementation, the data security strategy also includes de-identification rules. After receiving the statistical feature data returned by each target data source 10, the intermediate proxy layer 20 can perform a third operation based on the data security strategy: perform a security check on the statistical feature data based on the data security strategy. The security check includes: identifying whether the statistical feature data contains sensitive field information that exceeds the allowed range or detailed data that violates the exploration granularity limit; thereby establishing a final security defense before the statistical feature data enters the data synthesis layer 30.

[0137] First, the system identifies whether the statistical feature data contains sensitive field information that exceeds the permitted scope. The intermediate proxy layer 20 compares the physical field identifiers corresponding to each statistical item in the returned statistical feature data with the list of fields that the target data source 10 is allowed to explore in the data security policy. If the returned data is found to contain statistical features of unauthorized physical fields, it is determined that the security check has failed.

[0138] Secondly, the intermediate proxy layer 20 identifies whether the statistical feature data contains detailed data that violates the granularity limit of the investigation. The intermediate proxy layer 20 examines the data structure of the returned statistical feature data to determine if there is any content that may implicitly contain individual record information. For example, if the count value of a certain group in the frequency distribution histogram is 1, and the value range corresponding to that group is extremely narrow, then this statistical item may indirectly expose the existence of a single individual; if the number of returned groups is exactly equal to the count of duplicate values ​​in that field of the target data source 10, then it is essentially equivalent to returning a complete list of values. In the above cases, the intermediate proxy layer 20 can mark this part of the statistical features as a violation according to preset risk judgment rules.

[0139] After performing the security check, the intermediate proxy layer 20 can send the statistical feature data that passed the security check to the data synthesis layer 30; and, using the de-identification processing rules, perform de-identification processing on the illegal parts identified in the statistical feature data that failed the security check, and send the processed statistical feature data to the data synthesis layer 30.

[0140] Desensitization rules should include at least one or more of the following: (1) Masking of identifier fields. When the statistical feature data contains statistical descriptions of identifier fields (such as name, ID number, telephone number, etc.), the intermediate proxy layer 20 masks and replaces the identifiers of the returned statistical items. For example, for the returned list of high-frequency values, the specific name value is replaced with a general placeholder such as "VALUE_MASKED", and only the frequency statistics are retained; for the value range, it can be narrowed to a wider range that does not expose the individual boundary values.

[0141] (2) Perform interval or bucketing processing on sensitive numerical fields. When the statistical feature data contains detailed statistics on sensitive numerical fields (such as amount, health indicators, etc.), the intermediate proxy layer 20 converts the original continuous values ​​or fine statistical results into interval representations. For example, the original returned list of quantile values ​​is truncated to retain only the median and interquartile range; the fine grouping boundaries of the frequency distribution histogram are replaced with wider grouping intervals, or groups with counts below a preset threshold are merged into adjacent groups.

[0142] (3) Truncate the granularity of the output statistics. When the precision of the statistics in the statistical feature data is too high, which may imply the risk of reverse inference of the original records, the intermediate proxy layer 20 truncates or rounds the statistical results. For example, the floating-point mean or standard deviation is retained to a finite number of decimal places; the total number of records is rounded up or down by a preset order of magnitude.

[0143] In one implementation, the intermediate proxy layer 20 can reorganize statistical feature data from different target data sources 10 according to the target logical fields, forming a unified set of statistical features indexed by the logical fields, so that the data synthesis layer 30 can use it directly. For example, for the logical field user_id, the intermediate proxy layer 20 can merge the statistical features from data source A and data source C to form a complete description of the value range and frequency distribution information of the logical field.

[0144] The data synthesis layer 30 can generate a synthetic dataset that conforms to the received statistical feature data (S42) based on the received statistical feature data, and provide the synthetic dataset to the demand side 40 (S44).

[0145] For example, please refer to Figure 3In the process of generating the synthetic dataset, the data synthesis layer 30 is used to: obtain the synthesis quantity parameter carried in the data synthesis request; this parameter specifies the expected number of synthetic data records to be generated. The synthesis quantity parameter, along with the statistical feature data corresponding to the target logical field set, are input as conditions into the pre-built data synthesis model 31 to drive the data synthesis model to generate synthetic data one by one or in batches according to the distribution pattern and correlation constraints represented by the statistical feature data, until the number of synthetic data records reaches the value specified by the synthesis quantity parameter. Each synthetic data record is constructed and generated by the model through random sampling and constraint mapping, and does not contain any direct copying or simple transformation of any original data records from multiple data sources 10.

[0146] The data synthesis model 31 can be built into the data synthesis layer 30, or the data synthesis layer 30 can be associated with at least one external data synthesis model 31. This embodiment does not impose any restrictions on this.

[0147] The data synthesis layer 30 can perform quality verification on the generated synthetic dataset. For the synthetic dataset, the data synthesis layer 30 calculates statistics corresponding to the statistical feature type indicated in the data synthesis request, and calculates a distribution difference measure based on the difference between the statistics under the same statistical feature type and the received statistical feature data. For example, for numerical fields, the KS test statistic or Wasserstein distance between the generated sample and the target distribution can be calculated; for categorical fields, the total variational distance or Jensen-Shannon divergence between the generated sample and the target frequency distribution can be calculated.

[0148] If the calculated distribution difference metric exceeds a preset threshold, it indicates that the generated synthetic dataset deviates significantly from the real data distribution in terms of statistical characteristics. The data synthesis layer 30 then triggers a regeneration of the synthetic dataset. The regeneration process can be performed using an adjusted random seed or fine-tuned model parameters until the generated quality meets the preset threshold requirements. If multiple regenerations still fail to meet the quality requirements, the data synthesis layer 30 can return a prompt message to the party requiring data synthesis, suggesting adjustments to the statistical feature type or target logical field set in the data synthesis request.

[0149] Through the aforementioned generation and verification closed loop, the data synthesis layer 30 can continuously produce high-fidelity synthetic datasets that meet statistical constraints without the participation of original data records, providing usable and secure data supply for downstream AI model training, system testing, and data analysis.

[0150] For example, the synthetic data generation model is configured to receive at least one of the following statistical features as conditional inputs: ① marginal distribution features of a single field, including but not limited to: the range of field values, the proportion of null values, frequency distribution histograms or probability quality functions, and quantile sets; ② correlation features of multiple fields, including but not limited to: covariance matrices or correlation coefficient matrices between fields, conditional probability distribution tables, and mutual information values; ③ time series features, including but not limited to: the time interval distribution of event occurrences and state transition probability matrices; ④ business constraint features, including but not limited to: enumerated dictionaries that field values ​​must satisfy, format regular expressions, numerical range constraints, or functional dependencies between fields.

[0151] Synthetic data generation models include, but are not limited to, one or more of the following: conditional generative adversarial networks, variational autoencoders, Bayesian networks, Gaussian mixture models, sampling models based on kernel density estimation, differential privacy generation models, autoregressive models, streaming models, and hybrid rule-based and statistical generation models.

[0152] For example, in a conditional generative adversarial network, statistical feature data is encoded as a conditional vector; in a variational autoencoder, statistical feature data is used as prior distribution parameters; and in a Bayesian network, statistical feature data is directly instantiated into a conditional probability table. The specific transformation method is set according to the model architecture used, and this embodiment does not impose any restrictions on it.

[0153] The training or configuration process for synthetic data generation models varies depending on the model type. For example, for generative adversarial networks or variational autoencoder models, surrogate datasets are used for training, a distribution alignment loss term is introduced to constrain the statistical properties of the generated samples to be consistent with the target distribution, and differential privacy noise can be optionally applied to protect parameter privacy. Another example is for Bayesian networks or probabilistic graphical models, where the network topology is constructed based on a unified logical view, and the conditional probability distribution parameters of each node are directly instantiated using statistical feature data, without iterative training.

[0154] In summary, the data synthesis system provided in this embodiment can achieve the following technical effects: First, a non-intrusive probing mechanism based on the intermediate proxy layer 20. This solution abandons the traditional approach of copying and migrating real data, innovatively utilizing the intermediate proxy layer 20 as an intermediary between the data synthesis layer 30 and multiple data sources 10. This intermediate proxy layer 20 establishes logical connections with each data source 10, directly probing the real data distributed across different data sources 10 in real time, acquiring only the metadata and statistical feature data necessary for generating the synthesized data, while the original data records always reside locally on each data source 10. This mechanism fundamentally avoids touching original sensitive data during data probing, eliminates the physical path to data leakage, and achieves non-intrusive data utilization.

[0155] Second, a closed loop of dynamic synthesis and real-time feature extraction. This scheme deeply couples the synthetic data generation capability of the data synthesis layer 30 with the real-time feature exploration capability of the intermediate proxy layer 20, forming a closed-loop process of "on-demand request - real-time exploration - dynamic generation". Data synthesis requests are dynamically triggered. Each time a request is processed, the intermediate proxy layer 20 explores the latest statistical features of each target data source 10 in real time through logical connections. Based on this, the data synthesis layer 30 drives the synthetic data generation model to generate a synthetic dataset. This ensures that the synthetic dataset always reflects the current distribution of real data, solving the problem of insufficient timeliness of synthetic data.

[0156] Third, the collaborative architecture of the intermediate proxy layer 20 and the data synthesis layer 30. The system architecture of this solution uses the intermediate proxy layer 20 and the data synthesis layer 30 as the two core functional layers. The intermediate proxy layer 20 is responsible for providing secure and compliant cross-source data insights, including generating a unified logical view, determining the target data source 10, issuing statistical requests, and performing security checks and anonymization processing on the returned statistical feature data. The data synthesis layer 30 is responsible for generating high-fidelity synthetic datasets that conform to statistical feature constraints on demand. The two layers collaborate through standardized interfaces, achieving separation of responsibilities and efficient integration between data exploration and data generation. This architecture has high flexibility and scalability: the intermediate proxy layer 20 can interface with various heterogeneous data sources 10; the data synthesis layer 30 can integrate various synthetic data generation models to adapt to the differentiated needs of data synthesis algorithms in different application scenarios.

[0157] Therefore, the data synthesis system provided in this embodiment can realize real-time exploration of statistical features of multi-source heterogeneous data and on-demand generation of high-fidelity synthetic datasets while adhering to data security and privacy protection requirements, providing effective technical support for the secure circulation and in-depth utilization of data elements.

[0158] In some embodiments, at the specific engineering implementation level, the intermediate proxy layer 20 may be constructed using data virtualization technology or data weaving technology.

[0159] Data virtualization technology is a technical means to establish a logical abstraction layer between data source 10 and data consumer. Its core feature is that it provides a unified data access interface and logical view to the outside world, while internally it converts query requests for the logical view into native query instructions for each physical data source 10 through query rewriting and push-down mechanisms, thereby achieving the effect of "data does not move, logic is accessible".

[0160] Data weaving is a technical framework that further integrates proactive metadata management, knowledge graph construction, and automated data orchestration capabilities on the basis of data virtualization. Compared to traditional data virtualization, data weaving emphasizes semantic understanding of data assets, discovery of relationships, and intelligent execution of data supply strategies.

[0161] In practical deployments, data virtualization technology or data weaving technology can be flexibly selected as the underlying implementation framework of the intermediate proxy layer 20, depending on the complexity of the data source 10 environment and the requirements for intelligent functions. Neither implementation method changes the logical functions and security boundaries undertaken by the intermediate proxy layer 20 in the system.

[0162] In some implementation scenarios, multiple data sources 10 belong to independent management domains and are constrained by laws, regulations, or internal governance requirements, making cross-domain statistical exploration impossible. In such environments, the intermediate proxy layer 20 can adopt a federated learning mechanism as an alternative interaction method: the statistical feature exploration requirement is decomposed into computational tasks that each data source 10 can execute locally. After each data source 10 completes the statistical calculation locally, it applies privacy protection processing to the results and then returns them to the intermediate proxy layer 20 for aggregation and reconstruction to obtain statistical feature data equivalent to the global data distribution.

[0163] The aforementioned federated learning mechanism serves as an optional extension of the probing capabilities of the intermediate proxy layer 20. Its activation or deactivation does not affect the core function of the system in issuing statistical query requests based on logical connections under normal conditions. When the data source 10 environment meets the centralized probing conditions, the aforementioned statistical request issuance method can be used to reduce system complexity; when the data source 10 environment has strong isolation requirements, the federated learning mechanism can be enabled through configuration to adapt to higher levels of data governance constraints.

[0164] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they are not described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments is also within the scope of this specification.

[0165] In some embodiments, please refer to Figure 4 This specification also provides a data synthesis method for use in the intermediate proxy layer of the aforementioned data synthesis system. The method includes: In S400, logical connections are established with multiple data sources to obtain metadata from each data source. The metadata is used to describe the field attributes of the fields contained in the corresponding data source and the relationships between the fields.

[0166] In S402, based on the metadata of multiple data sources, the fields in the multiple data sources are semantically aligned to determine multiple field sets composed of semantically equivalent fields, and corresponding logical fields are generated for each field set. Based on the relationship between fields described by each metadata, the logical relationship between different logical fields is determined, and a unified logical view is generated based on the logical fields and logical relationships and provided to the data composition layer.

[0167] In S404, a data synthesis request generated by the data synthesis layer according to the unified logical view is received. The data synthesis request is used to indicate the target set of logical fields to be synthesized and the statistical feature type.

[0168] In S406, based on the target logical field set, the target data source among multiple data sources is determined, and a statistical request corresponding to the statistical feature type is generated for the target data source and sent to the target data source.

[0169] In S408, the statistical feature data returned by the target data source after executing the statistical request is received and forwarded to the data synthesis layer so that the data synthesis layer can generate a synthetic dataset that conforms to the statistical feature data.

[0170] In one implementation, the intermediate proxy layer has a pre-defined data security policy, which includes at least: field sensitivity levels, a whitelist of statistical feature types corresponding to each sensitivity level, and de-identification rules. The de-identification rules include at least one or more of the following: masking of identifier fields, range-based or bucket-based processing of sensitive numerical fields, and truncation of the granularity of output statistics.

[0171] The method also includes: Perform at least one of the following actions based on the data security policy: In the first operation, after receiving the data synthesis request, an access check is performed on the data synthesis request based on the data security policy. The access check includes: verifying whether the sensitivity level of each target logical field indicated by the data synthesis request is allowed to be requested, and whether the required statistical feature type is in the whitelist of statistical feature types corresponding to the sensitivity level; if the verification fails, a prompt message indicating that the access check has failed is returned. In the second operation, during the process of generating statistical requests to be sent to each target data source, the statistical requests are constrained based on the data security policy. The constraint processing includes at least one of the following: restricting the fields to be explored to not exceed the field range allowed by the statistical feature type whitelist, restricting the exploration granularity to not include detailed data, and restricting the returned results to only aggregate statistics. In the third operation, after receiving the statistical feature data returned by each target data source, the intermediate proxy layer performs a security check on the statistical feature data based on the data security policy. The security check includes: identifying whether the statistical feature data contains sensitive field information that exceeds the allowed range or detailed data that violates the exploration granularity limit; sending the statistical feature data that passes the security check to the data synthesis layer; and using the de-identification processing rules, performing de-identification processing on the illegal parts identified in the statistical feature data that fails the security check, and sending the processed statistical feature data to the data synthesis layer.

[0172] In one implementation, a target data source is determined from multiple data sources based on a set of target logical fields, including: Based on the set of fields corresponding to each target logical field in the target logical field set, obtain the list of data sources associated with each target logical field; Based on the list of data sources associated with each target logical field, determine at least one subset of data sources from multiple data sources that can cover all target logical fields in the target logical field set; When multiple subsets of data sources exist, the target subset of data sources is determined from the multiple subsets of data sources according to a preset selection strategy; The data source subset selection strategy includes: taking the minimum number of data sources in the selected target data source subset as the primary objective; when there are multiple data source subsets that satisfy the primary objective, a secondary selection is performed based on the priority of each data source in the data source subset, and the data source subset with the highest priority is selected as the target data source subset; The priority of each data source is determined based on at least one of the following dimensions: data quality score of the data source, data update timeliness level, access response latency level, and trust level of the security domain in which the data source is located.

[0173] In one implementation, semantic alignment includes: for fields in multiple data sources, identifying semantically equivalent fields and grouping them into the same set of fields based on at least one of textual similarity of field names, overlap of field value range distribution, and similarity of metadata annotation information of the fields in their respective data sources. Logical fields are generated in the following way: For each field belonging to the same field set, the common field attributes of each field in the field set are extracted to form logical fields and field attributes of logical fields. The field attributes of logical fields include: logical field name, data type, value range, and field description information. The logical relationships between different logical fields are determined based on the primary and foreign key relationships, functional dependencies between fields, business constraints between fields, temporal relationships between fields, and / or statistical correlations between field value ranges recorded in the metadata of each data source.

[0174] In some embodiments, please refer to Figure 5 This specification also provides a data synthesis method for use in the data synthesis layer of the aforementioned data synthesis system. The method includes: In S500, a unified logical view is obtained from the intermediate proxy layer. The unified logical view consists of logical fields generated by the intermediate proxy layer based on metadata from multiple data sources and the logical relationships between different logical fields.

[0175] In S502, a data synthesis request generated according to a unified logical view is received. The data synthesis request is used to indicate the target set of logical fields to be synthesized and the statistical feature type.

[0176] In S504, the data synthesis request is sent to the intermediate proxy layer.

[0177] In S506, statistical feature data forwarded by the intermediate proxy layer is received. The statistical feature data is returned by the target data source after executing the statistical request corresponding to the data synthesis request.

[0178] In S508, a synthetic dataset that conforms to statistical characteristics is generated.

[0179] In one implementation, a synthetic dataset that conforms to statistical characteristics is generated, including: Obtain the number of data to be synthesized carried in the data synthesis request; The synthesis quantity parameter and the statistical feature data corresponding to the target logical field set are input as conditions into the pre-built data synthesis model to drive the data synthesis model to generate a synthetic dataset that matches the synthesis quantity parameter based on the statistical feature data. For synthetic datasets, calculate the statistics corresponding to the statistical feature types, and calculate the distribution difference measure based on the differences between the statistics under the same statistical feature type and the statistical feature data; If the distribution difference metric exceeds a preset threshold, the synthetic dataset will be regenerated.

[0180] For details on the implementation of the above method, please refer to the relevant sections of the above system; they will not be repeated here.

[0181] Figure 6 This is a schematic structural diagram of a device provided in an exemplary embodiment. For example... Figure 6As shown, device 600 mainly consists of a communication interface 602, a user interface 604, a processor 606, and a data storage 608. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 610. The communication interface 602 enables device 600 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 602 may include an antenna and related processing devices for wireless communication with a radio access network or access point. Furthermore, the communication interface 602 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 602 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 602 may also include multiple physical communication interfaces, such as Wi-Fi interfaces, Bluetooth interfaces, and wide-area wireless interfaces.

[0182] User interface 604 includes receiving user input and providing output to the user. Therefore, user interface 604 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 604 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 604 may include software, circuitry, or other forms of logic capable of transmitting and receiving data from external user input / output devices. Additionally or alternatively, device 600 may support remote access from other devices via communication interface 602 or another physical interface (not shown). User interface 604 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 604 may also be configured as a display device for rendering or displaying text fragments.

[0183] Processor 606 may contain one or more general-purpose processors and / or special-purpose processors.

[0184] Data storage 608 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 606. Data storage 608 may include removable and non-removable components.

[0185] Processor 606 is capable of executing program instructions 618 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 608 to perform the various functions described herein. Data storage 608 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 600, enable device 600 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 618 by processor 606 may result in processor 606 using data 612.

[0186] For example, program instructions 618 may include an operating system 622 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 600 and one or more applications 620 (e.g., a browser, social application, or game application). Similarly, data 612 may include operating system data 616 and application data 614. Operating system data 616 is primarily accessible to the operating system 622, while application data 614 is primarily accessible to one or more applications 620. Application data 614 may reside in a file system visible or hidden from the user of device 600.

[0187] Application 620 can communicate with operating system 622 through one or more application programming interfaces (APIs). These APIs help application 620 read and / or write application data 614, transmit or receive information via communication interface 602, receive or display information on user interface 604, etc.

[0188] In some terminology, application 620 may be simply referred to as "app". Furthermore, application 620 can be downloaded to device 600 through one or more online app stores or app markets. However, applications can also be installed on device 600 in other ways, such as through a web browser or a physical interface on device 600 (e.g., a USB port).

[0189] In some embodiments, the data synthesis apparatus corresponding to the intermediate processing layer can be applied to, for example... Figure 6 The device shown implements the technical solution described in this specification. For details on the specific implementation of this data synthesis apparatus, please refer to the relevant sections mentioned above; further details will not be repeated here.

[0190] In some embodiments, the data composition apparatus corresponding to the data composition layer can be applied to, for example... Figure 6 The device shown implements the technical solution described in this specification. For details on the specific implementation of this data synthesis apparatus, please refer to the relevant sections mentioned above; further details will not be repeated here.

[0191] Based on the same concept as the methods described above, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor performs the steps of the method as described in any of the above embodiments by executing the executable instructions.

[0192] Based on the same concept as the methods described above, this specification also provides a computer-readable storage medium having computer instructions stored thereon that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.

[0193] Based on the same concept as the methods described above, this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.

[0194] What those skilled in the art will understand is: In this specification, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in a process, method, product, or apparatus that includes said elements is not excluded.

[0195] In this specification, “a,” “an,” and “the” do not specifically refer to the singular, but may also include the plural.

[0196] In this specification, ordinal numbers such as "first," "second," etc., do not necessarily indicate order; they are often used to distinguish between objects. For example, "first server" and "second server" usually refer to two servers. To differentiate between these two servers, they are described as "first server" and "second server." Of course, sometimes these two servers may be the same server.

[0197] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving the data sent by B, or it can be understood as A indirectly receiving the data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending the data directly to A, or it can be understood as B indirectly sending the data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.

[0198] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.

[0199] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.

[0200] Although one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is only one of many possible execution orders and does not represent the only execution order. Therefore, when the claims involve method steps, any changes or adjustments to the order of such steps, or the parallelism between steps, are also within the scope of protection of the claims.

Claims

1. A data synthesis system, comprising a data synthesis layer, an intermediate proxy layer, and multiple data sources; The intermediate proxy layer is used to establish logical connections with the multiple data sources respectively, obtain the metadata of each data source, the metadata being used to describe the field attributes and inter-field relationships of the fields contained in the corresponding data source; based on the metadata of each of the multiple data sources, semantic alignment is performed on the fields in the multiple data sources to determine multiple field sets composed of semantically equivalent fields, and corresponding logical fields are generated for each field set; and based on the inter-field relationships described by the metadata, the logical relationships between different logical fields are determined. A unified logical view is generated based on the logical fields and the logical relationships and provided to the data synthesis layer; The data synthesis layer is used to send a data synthesis request generated according to the unified logical view to the intermediate proxy layer; the data synthesis request is used to indicate the target logical field set to be synthesized and the statistical feature type. The intermediate proxy layer is used to determine the target data source among the multiple data sources based on the target logical field set, and to generate a statistical request for the target data source corresponding to the statistical feature type; The target data source is used to execute the statistical request, obtain and return statistical feature data to the intermediate proxy layer; The intermediate proxy layer is used to send the statistical feature data to the data synthesis layer; The data synthesis layer is used to generate a synthetic dataset that conforms to the statistical characteristics.

2. The system according to claim 1, wherein the intermediate proxy layer is pre-configured with a data security policy, the data security policy including at least: Field sensitivity levels, and whitelists of statistical feature types corresponding to each sensitivity level; The intermediate proxy layer is also used to perform at least one of the following operations based on the data security policy: In the first operation, after receiving the data synthesis request, an access check is performed on the data synthesis request based on the data security policy. The access check includes: verifying whether the sensitivity level of each target logical field indicated by the data synthesis request is allowed to be requested, and whether the required statistical feature type is in the whitelist of statistical feature types corresponding to the sensitivity level; if the verification fails, a prompt message indicating that the access check has failed is returned. In the second operation, during the process of generating statistical requests to be sent to each target data source, the statistical requests are constrained based on the data security policy. The constraint processing includes at least one of the following: restricting the fields to be explored to not exceed the field range allowed by the statistical feature type whitelist, restricting the exploration granularity to not include detailed data, and restricting the returned results to only aggregate statistics.

3. The system according to claim 2, wherein the data security strategy further includes de-identification processing rules; The desensitization processing rules include at least the following: One or more of the following can be used: masking of identifier fields, interval or bucketing of sensitive numerical fields, and truncation of the granularity of output statistics. The at least one operation also includes: In the third operation, after receiving the statistical feature data returned by each target data source, the intermediate proxy layer performs a security check on the statistical feature data based on the data security policy. The security check includes: identifying whether the statistical feature data contains sensitive field information that exceeds the allowed range or detailed data that violates the exploration granularity limit. The statistical feature data that passes the security check is sent to the data synthesis layer; and, using the desensitization processing rules, the non-compliant parts identified in the statistical feature data that fail the security check are desensitized, and the processed statistical feature data is sent to the data synthesis layer.

4. In the system according to claim 1, during the process of determining the target data source, the intermediate proxy layer is used for: Based on the set of fields corresponding to each target logical field in the target logical field set, obtain the list of data sources associated with each target logical field; Based on the list of data sources associated with each target logical field, determine at least one subset of data sources from the plurality of data sources that can cover all target logical fields in the set of target logical fields; In the case of multiple data source subsets, the target data source subset is determined from the multiple data source subsets according to a preset selection strategy; The data source subset selection strategy includes: taking the minimum number of data sources in the selected target data source subset as the first objective; when there are multiple data source subsets that satisfy the first objective, a secondary selection is performed based on the priority of each data source in the data source subset, and the data source subset with the highest priority is selected as the target data source subset; The priority of each data source is determined based on at least one of the following dimensions: data quality score of the data source, data update timeliness level, access response latency level, and trust level of the security domain in which the data source is located.

5. The system according to claim 1, wherein the semantic alignment comprises: For fields in the multiple data sources, based on at least one of the following: text similarity of field names, overlap of field value range distribution, and similarity of field metadata annotation information in their respective data sources, semantically equivalent fields are identified and grouped into the same field set. The logical fields are generated in the following way: For each field belonging to the same field set, the common field attributes of each field in the field set are extracted to form logical fields and field attributes of logical fields. The field attributes of logical fields include: logical field name, data type, value range and field description information. The logical relationships between different logical fields are determined based on the primary and foreign key relationships, functional dependencies between fields, business constraints between fields, temporal relationships between fields, and / or statistical correlations between field value ranges recorded in the metadata of each data source.

6. The system according to claim 1, wherein during the generation of the synthetic dataset, the data synthesis layer is used to: Obtain the synthesis quantity parameter carried in the data synthesis request; The synthesis quantity parameter and the statistical feature data corresponding to the target logical field set are input together as conditions into the pre-built data synthesis model to drive the data synthesis model to generate a synthetic dataset that matches the synthesis quantity parameter based on the statistical feature data. For the synthetic dataset, calculate the statistic corresponding to the statistical feature type, and calculate the distribution difference measure based on the difference between the statistic and the statistical feature data under the same statistical feature type; If the distribution difference metric exceeds a preset threshold, the synthetic dataset is regenerated.

7. A data synthesis method, applied to an intermediate proxy layer in a data synthesis system as described in any one of claims 1 to 6, the method comprising: Logical connections are established with multiple data sources respectively to obtain metadata of each data source. The metadata is used to describe the field attributes of the fields contained in the corresponding data source and the relationship between the fields. Based on the metadata of the multiple data sources, the fields in the multiple data sources are semantically aligned to determine multiple field sets composed of semantically equivalent fields, and corresponding logical fields are generated for each field set. Based on the field associations described by the metadata, the logical relationships between different logical fields are determined, and a unified logical view is generated based on the logical fields and the logical relationships and provided to the data synthesis layer. Receive the data synthesis request generated by the data synthesis layer according to the unified logical view, wherein the data synthesis request is used to indicate the target logical field set to be synthesized and the statistical feature type; Based on the target logical field set, a target data source among the multiple data sources is determined, and a statistical request corresponding to the statistical feature type is generated for the target data source and sent to the target data source; The system receives statistical feature data returned by the target data source after executing the statistical request, and forwards it to the data synthesis layer so that the data synthesis layer can generate a synthetic dataset that conforms to the statistical feature data.

8. A data synthesis method, applied to the data synthesis layer of a data synthesis system as described in any one of claims 1 to 6, the method comprising: Obtain a unified logical view provided by the intermediate proxy layer. The unified logical view consists of logical fields generated by the intermediate proxy layer based on metadata from multiple data sources and the logical relationships between different logical fields. Receive a data synthesis request generated according to the unified logical view, the data synthesis request being used to indicate the target logical field set to be synthesized and the statistical feature type; The data synthesis request is sent to the intermediate proxy layer; Receive statistical feature data forwarded by the intermediate proxy layer, wherein the statistical feature data is returned by the target data source after executing a statistical request corresponding to the data synthesis request; Generate a synthetic dataset that conforms to the statistical characteristics.

9. An electronic device, comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method as claimed in any one of claims 7 or 8 by executing the executable instructions.

10. A computer-readable storage medium having stored thereon computer instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 7 or 8.

11. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 7 or 8.