Data delivery method and device, equipment and storage medium
By generating synthetic data with similar characteristics to business data through generative adversarial networks, the problem of insufficient data in user organization systems is solved, the effectiveness of quality optimization operations is improved, and data security is guaranteed.
Patent Information
- Application Number
- CN202311120295.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-31
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-08-31
AI Technical Summary
The user organization's system does not obtain enough business data, resulting in poor quality optimization operations. Furthermore, the anonymized data is insufficient to meet the requirements of quality optimization operations, affecting the accuracy and security of data features.
By initializing a generative adversarial network and training it through mutual adversarial interaction between the generator and the discriminator, synthetic data with the same characteristics as business data but without real information is generated, thus meeting the data needs of user organization systems.
The generated synthetic data is correlated with business data, meets the data quality and quantity requirements of quality optimization operations, improves the effectiveness of quality optimization operations in user organizations' systems, and ensures data security without leaking sensitive information.
Smart Images

Figure CN117216557B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing, and in particular relates to a data distribution method, apparatus, device and storage medium. Background Technology
[0002] With the development of electronic information technology, more and more business operations can be realized through electronic business systems. In order to provide users with more tailored business services, business data can be acquired for quality optimization operations such as user profiling, data analysis, and business monitoring. However, to ensure user data security, user organizations can only access data within their own authorized scope, such as business data within their own jurisdiction. This results in the quantity of business data acquired being insufficient to meet the needs of quality optimization operations, leading to poor effectiveness of these operations. Summary of the Invention
[0003] This application provides a data distribution method, apparatus, device, and storage medium that can meet the quality and quantity requirements of user organizations for data during quality optimization operations.
[0004] In a first aspect, embodiments of this application provide a data distribution method, comprising: initializing a generative adversarial network (GAN) based on business data from multiple data sources, wherein the GAN includes a generator and a discriminator, and the data sources are obtained based on data requirement information of a user organization system; extracting a correlation vector from the business data based on acquired input requirement parameters, wherein the correlation vector represents the distribution of effective fields in the business data; loading the correlation vector into the generator, iteratively training the GAN until the GAN meets the iteration requirements, and determining the generator in the GAN that meets the iteration requirements as the business data generation model, wherein the synthetic data generated by the generator and the business data serve as inputs to the discriminator, and the output information of the discriminator is used to train the generator; and generating synthetic data using the business data generation model and the acquired data output strategy and distributing it to the user organization system through a subscribed data distribution system.
[0005] Secondly, embodiments of this application provide a data distribution apparatus, comprising: a model training module, configured to initialize a generative adversarial network (GAN) based on business data from multiple data sources, the GAN including a generator and a discriminator, the data sources being obtained based on data requirement information from a user organization system; and a module for extracting correlation vectors from the business data based on acquired input requirement parameters, the correlation vectors representing the distribution of effective fields in the business data; and a module for loading the correlation vectors into the generator, iteratively training the GAN until it meets the iteration requirements, and determining the generator in the GAN that meets the iteration requirements as the business data generation model, wherein the synthetic data generated by the generator and the business data serve as inputs to the discriminator, and the output information of the discriminator is used to train the generator; and a data distribution module, configured to generate synthetic data using the business data generation model and the acquired data output strategy, and distribute it to the user organization system through a subscribed data distribution system.
[0006] Thirdly, embodiments of this application provide an electronic device, including: a processor and a memory storing computer program instructions; the processor executes the computer program instructions to implement the data delivery method of the first aspect.
[0007] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the data delivery method of the first aspect.
[0008] This application provides a data distribution method, apparatus, device, and storage medium that initializes a generative adversarial network (GAN) based on business data from a data source matching the data requirements of a user organization's system. Based on input requirements parameters, a correlation vector is extracted to represent the distribution of valid fields in the business data from multiple data sources. By loading this correlation vector into the generator within the GAN, the synthesized data output by the generator conforms to the distribution of valid fields in the business data. The generator and discriminator in the GAN train against each other through multiple iterations, ensuring the GAN meets iteration requirements. The generator in the adversarial GAN that meets these requirements is used as the business data generation model, and synthesized data is output according to a data output strategy. This synthesized data is correlated with the business data from the data source matching the user organization's system's data requirements and conforms to the distribution of valid fields in the introduced data source's business data, reflecting the characteristics of the business data required by the user organization's system. The business data generation model can continuously generate synthesized data that reflects the characteristics of the business data required by the user organization's system, meeting the quality and quantity requirements of the user organization's system for quality optimization operations, thereby improving the effectiveness of the user organization's quality optimization operations. Attached Figure Description
[0009] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 A flowchart illustrating a data distribution method provided in an embodiment of this application;
[0011] Figure 2 A flowchart illustrating a data delivery method provided in another embodiment of this application;
[0012] Figure 3 This is a schematic diagram illustrating an example of the process of loading business data from a data source into a relational vector and then onto a generator, as provided in an embodiment of this application.
[0013] Figure 4 A flowchart illustrating an example of the process for obtaining a business data generation model provided in an embodiment of this application;
[0014] Figure 5 A flowchart illustrating a data distribution method provided in yet another embodiment of this application;
[0015] Figure 6 A schematic diagram illustrating an example of a data delivery process provided in an embodiment of this application;
[0016] Figure 7 A flowchart illustrating an example of a real-time data delivery process provided in an embodiment of this application;
[0017] Figure 8 This is a schematic diagram of the structure of a data transmission device provided in an embodiment of this application;
[0018] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0019] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples. It should be noted that the acquisition, storage, use, and processing of information and data in the embodiments of this application are all authorized by users or relevant organizations and comply with the relevant provisions of national laws and regulations.
[0020] With the development of electronic information technology, more and more business operations can be realized through electronic business systems. To provide users with more tailored business services, business data can be acquired for quality optimization operations such as user profiling, data analysis, and business monitoring. Data providers can register as publishers on the data distribution server to push business data. The data distribution system registers as subscribers on the data distribution server and subscribes to publishers to obtain business data in real time. User organization systems can obtain business data distributed by the data distribution system. Because business data may involve sensitive user information, to ensure user data security, on the one hand, user organizations can only access data within their own authorized scope, such as business data within their own jurisdiction, forming data silos; on the other hand, business data is anonymized before being distributed to user organization systems. The amount of data that user organization systems can access within their own authorized scope cannot meet the needs of quality optimization operations, resulting in poor quality optimization effectiveness. Furthermore, due to the lack of key information in the anonymized business data, the data content is insufficient to meet the needs of quality optimization operations, leading to poor quality optimization effectiveness. In some cases, the amount of business data for a specific business scenario or business cycle targeted by quality optimization operations is too small to meet the needs of quality optimization operations, resulting in poor quality optimization results. However, if business data from other business scenarios or business cycles is obtained for quality optimization operations, in addition to the business data from the specific business scenario or business cycle, the business data from other business scenarios and business cycles cannot reflect the characteristics of the business data from the specific business scenario or business cycle, which may have a negative impact on the quality optimization operations and reduce their effectiveness.
[0021] This application provides a data distribution method, apparatus, device, and storage medium. It can iteratively train a generative adversarial network (GAN) based on business data from multiple data sources and the input requirements of the user organization's business data needs. The generator in the iteratively trained GAN can continuously output synthetic data. This synthetic data is not real business data and does not contain real information, but it has the same or similar characteristics as real business data. It is not subject to the user organization's system permissions, and the user organization's system can customize cross-regional and cross-type synthetic data for quality optimization operations. This synthetic data does not involve sensitive user information, and because it has the same or similar characteristics as real business data, it will not damage the data characteristics, ensuring the quality and quantity of the synthetic data. This satisfies the user organization's system's data requirements for quality optimization operations and improves the effectiveness of quality optimization operations.
[0022] The data distribution method, apparatus, equipment, and storage medium provided in this application will be described below.
[0023] The first aspect of this application provides a data distribution method that can be applied to scenarios where data is provided to user organization systems. This data distribution method can be executed by a data distribution device, equipment, system, platform, etc., and is not limited thereto. Figure 1 A flowchart of a data distribution method provided in an embodiment of this application is shown below. Figure 1 As shown, the data distribution method may include steps S101 to S104.
[0024] In step S101, an adversarial network is initialized and generated based on business data from multiple data sources.
[0025] Data sources can be categorized based on factors such as scenario, time, and permissions. The multiple data sources introduced are the data sources participating in the data distribution method in this application embodiment. The introduced data sources can be obtained based on the data requirement information of the user organization system. The data requirement information of the user organization system can characterize the user organization system's data needs. The data requirement information of the user organization system can be analyzed to determine the user organization system's required business scenarios. Based on these scenarios, the introduced data sources can be determined, including which data sources to introduce and the amount of business data obtained from each data source. For example, if the data sources are categorized by province, and the user organization system's data requirement information characterizes its needs including business data in region A1, then the introduced data sources can include data sources from each province within region A1. As another example, if the data sources are categorized by time (by week), and the user organization system's data requirement information characterizes its needs including business data in November, then the introduced data sources can include data sources from each week in November. The types of business data can be determined based on specific business areas. For example, in the transaction payment field, business data can include transaction flow data.
[0026] A generative adversarial network (GAN) can be initialized using business data from multiple data sources. The GAN consists of a generator and a discriminator. Initialization of the GAN includes initializing the generator and the discriminator. The generator produces synthetic data, which is not real data like the business data and does not contain any real information. The discriminator distinguishes the synthetic data from other data.
[0027] In step S102, the correlation vector of the business data is extracted based on the obtained input requirement parameters.
[0028] Input requirement parameters can be used to reflect the requirements for business data. In some examples, input requirement parameters may include, but are not limited to, one or more of the following: data size, data format, data constraints, and data relationships. Based on the input requirement parameters, a relationship vector can be extracted from the business data. The relationship vector represents the distribution of valid fields in the business data. Valid fields are the fields in the data domains that need to be focused on. For example, if the business data includes user gender, user date of birth, and user place of origin, and the data domains of focus include user gender and user date of birth, then the fields corresponding to user gender and user date of birth in the business data are the valid fields.
[0029] In step S103, the correlation vector is loaded into the generator, and the generative adversarial network is iteratively trained until the generative adversarial network meets the iteration requirements. The generator in the generative adversarial network that meets the iteration requirements is determined as the business data generation model.
[0030] The correlation vector is loaded into the generator, and the synthetic data output by the generator will be constrained by the correlation vector. The synthetic data generated by the generator and the business data can be used as input to the discriminator. The output information of the discriminator can be used to train the generator. The generator's goal is to generate synthetic data that is difficult for the discriminator to distinguish, and the discriminator's goal is to distinguish the synthetic data generated by the generator. The generator and the discriminator train against each other, thereby improving the correlation between the synthetic data generated by the generator and the business data.
[0031] The iteration requirement can be a condition for the end of iterative training. In some examples, the iteration requirement may include the adversarial generative network reaching a preset number of iterations, and / or the loss value of the adversarial generative network after iterative training being within a preset loss value range.
[0032] In step S104, synthetic data is generated using the business data generation model and the acquired data output strategy, and then distributed to the user organization system through the subscribed data distribution system.
[0033] Data output strategies can be pre-defined according to the needs of the user organization's system. In some examples, the data output strategy may include, but is not limited to, one or more of the following: data distribution frequency, data distribution scale, and data distribution format. The business data generation model can output synthetic data according to this data output strategy. Synthetic data conforming to the data output strategy can be offline sent to the database, and then the database can publish the synthetic data to the data distribution server. The data distribution system can subscribe to obtain synthetic data from the data distribution server in real time and distribute it to the user organization's system, enabling the user organization's system to obtain synthetic data in real time.
[0034] In this embodiment, a generative adversarial network (GAN) can be initialized based on business data from a data source that matches the data requirements of the user organization system. Based on input requirements parameters, a correlation vector is extracted to represent the distribution of valid fields in the business data from multiple data sources. By loading this correlation vector into the generator within the GAN, the synthesized data output by the generator conforms to the distribution of valid fields in the business data. The generator and discriminator in the GAN train against each other. After multiple iterations, the GAN satisfies the iteration requirements. The generator in the adversarial GAN that meets these requirements is used as the business data generation model, and synthesized data is output according to a data output strategy. This synthesized data is correlated with the business data from the data source matching the user organization system's data requirements and conforms to the distribution of valid fields in the business data from the introduced data source, thus reflecting the characteristics of the business data required by the user organization system. The business data generation model can continuously generate synthesized data that reflects the characteristics of the business data required by the user organization system, meeting the quality and quantity requirements of the user organization system for quality optimization operations, thereby improving the effectiveness of the user organization system's quality optimization operations. Furthermore, the synthetic data output by the business data generation model is not real data and does not contain real information, thus preventing the leakage of users' sensitive information and ensuring data security.
[0035] In some embodiments, data features of business data can be extracted first, and then correlation vectors can be selectively generated based on the data features to perform multivariate pre-training of business data before iterative training of the generative adversarial network. Figure 2 A flowchart illustrating a data distribution method provided in another embodiment of this application. Figure 2 and Figure 1 The difference is that, Figure 1 Step S102 can be further refined as follows: Figure 2 Steps S1021 to S1023 in the process, Figure 1 Step S103 can be further refined as follows: Figure 2 Steps S1031 to S1034 in the process.
[0036] In step S1021, the data features of the business data are extracted.
[0037] Data features are extracted from business data from the introduced data source, and these data features can reflect the characteristics of the business data.
[0038] In step S1022, the data features are integrated and filtered according to the input requirements parameters to obtain effective data features.
[0039] Based on the input parameters, the system can determine the data features that need to be integrated, forgotten, and remembered. Effective data features can include those after integration and filtering. In some examples, a Long Short-Term Memory (LSTM) network can be used to integrate and filter data features to obtain effective data features. The LSTM network selects data features for reinforcement and those for forgetting, thus yielding effective data features.
[0040] In step S1023, a correlation vector is generated based on the characteristics of the valid data.
[0041] Based on the distribution of valid fields in the business data as reflected by the valid data characteristics, a correlation vector is generated.
[0042] Figure 3 This is a schematic diagram illustrating an example of the process of loading business data from a data source into a relational vector and then into a generator, as provided in an embodiment of this application. Figure 3 As shown, assuming four data sources are introduced, namely data source 1, data source 2, data source 3 and data source 4, after receiving business data from the four data sources, data features can be extracted and divided into data feature set 1, data feature set 2, data feature set 3, data feature set 4 and data feature set 5. Each data feature set may include at least one data feature. The above five data feature sets are integrated and selected to generate a correlation vector, which can be output and loaded into the generator.
[0043] In step S1031, the correlation vector is loaded into the generator, and the generator generates the synthetic data.
[0044] By loading correlation vectors into the generator and iteratively training the generator in subsequent processes, compared with traditional training methods, the addition of correlation vectors can solve the problem of generating a large amount of synthetic data with the same or similar features as the business data. This solves the problem of gradient vanishing or gradient exploding caused by insufficient data for iterative training, and can also solve the problem of data feature dilution when multiple data sources are introduced.
[0045] In step S1032, the synthesized data and business data are mixed and input into the discriminator. The discriminator samples the data domains of the synthesized data and business data, and randomly swaps the fields in the data domains of different data to generate combined data.
[0046] The discriminator samples the data domains of the synthetic data and the business data to obtain the fields in each data domain of the synthetic data and the fields in each data domain of the business data. It can randomly swap the fields in the same data domain of different data to obtain new data, which is the combined data.
[0047] In step S1033, the discriminator outputs result information based on the combined data.
[0048] This result information can be used to train the generator. In some examples, a discriminator can evaluate the relevance of the combined data to obtain data relevance parameters; the discriminator then generates result information based on the combined data with the highest relevance indicated by the data relevance parameters. Relevance evaluation can be achieved using a search-number algorithm combined with the data requirements of the user organization's system. Data relevance parameters characterize the relevance between combined data and business data. Data relevance parameters can be positively or negatively correlated with relevance. For ease of explanation, this application embodiment uses the example of data relevance parameters being positively correlated with relevance. Data relevance parameters may include a relevance score; the higher the relevance score, the higher the relevance. The combined data with the highest relevance indicated by the data relevance parameters is the combined data with the highest relevance to the business data.
[0049] In step S1034, the result information is used as training parameters to train the generator until the generative adversarial network meets the iteration requirements.
[0050] The information carried by the output of this combined data can better help the generator to be trained, making the features of the synthesized data generated by the generator closer to the business data.
[0051] Each time the generator outputs synthesized data, it can be determined whether the generative adversarial network meets the iteration requirements. If the iteration requirements are met, the generator is used as the business data generation model; if the iteration requirements are not met, the process can return to step S1032 to continue training the discriminator and generator.
[0052] To facilitate understanding, an example is provided here to illustrate the process of obtaining the business data generation model described above. Figure 4 A flowchart illustrating an example of the process for obtaining a business data generation model provided in an embodiment of this application is shown below. Figure 4 As shown, the process of obtaining the business data generation model may include steps a1 to a17.
[0053] In step a1, the required business scenarios are analyzed based on the user organization's requirements information.
[0054] In step a2, the data source and data scale to be introduced are determined based on the required business scenario.
[0055] In step a3, the adversarial network is initialized and generated based on the business data from the introduced data source.
[0056] In step a4, the multi-source pre-training module of the generative adversarial network is initialized.
[0057] The multi-source pre-training module of the generative adversarial network can be initialized using business data from the introduced data source and the received input requirement parameters. The multi-source pre-training module is used to execute the process of generating correlation vectors as described in the above embodiments.
[0058] In step a5, data features are extracted from the business data of the introduced data source.
[0059] In step a6, the data features are integrated and selected.
[0060] In step a7, a correlation vector is generated based on the integrated and selected data features.
[0061] In step a8, the correlation vector is loaded into the generator, and iterative training begins.
[0062] In step a9, the generator generates the synthetic data.
[0063] In step a10, it is determined whether the Generative Adversarial Network (GAN) meets the iteration requirements. If the GAN meets the iteration requirements, proceed to step a17; if the GAN does not meet the iteration requirements, proceed to step a11.
[0064] In step a11, the synthesized data and business data are input into the discriminator.
[0065] In step a12, the discriminator performs data domain sampling on the synthesized data and the business data.
[0066] In step a13, the discriminator randomly swaps the fields of the data to obtain combined data. Specifically, the discriminator randomly swaps the fields of the data field of the combined data and the fields of the data domain of the business data.
[0067] In step a14, the discriminator scores the correlation of the combined data.
[0068] In step a15, determine whether the correlation score of the combined data is the highest score. If the correlation score of the combined data is the highest score, proceed to step a16; if the correlation score of the combined data is not the highest score, return to step a13.
[0069] In step a16, the discriminator outputs the result information of the highest correlation score, and returns to step a8, using the result information as training parameters to train the generator.
[0070] In step a17, the generator is determined as the output of the business data generation model.
[0071] The specific details of steps a1 to a17 above can be found in the relevant descriptions in the above embodiments, and will not be repeated here.
[0072] In some embodiments, the synthetic data output by the business data generation model may also be subject to constraints such as data output strategy, data format, data distribution rules, and data group leader rules, and will ultimately be distributed to the user organization system. Figure 5 A flowchart illustrating a data distribution method provided in another embodiment of this application. Figure 5 and Figure 1 The difference is that, Figure 1 Step S104 can be further refined as follows: Figure 5 Steps S1041 to S1044 in the process.
[0073] In step S1041, the business data generation model and data output strategy are used to output synthetic data that conforms to the data output strategy.
[0074] The data output strategy includes the output rules for the synthetic data generated by the business data generation model, which can be determined according to the needs of the user organization's system. The data output strategy can be loaded into the business data generation model so that the business data generation model outputs synthetic data that conforms to the data output strategy.
[0075] In step S1042, the synthesized data that conforms to the data output strategy is formatted into a synthesized data format supported by the data distribution system.
[0076] The data distribution system has format requirements for the synthesized data it distributes. Therefore, the synthesized data obtained in step S1041 needs to be formatted to ensure its format matches the data formats supported by the data distribution system. The formatted synthesized data can be input into a preset database and then distributed from that database to the data distribution system.
[0077] In step S1043, if the synthesized data conforms to the data distribution rules configured by the data distribution system, the data distribution system assembles the synthesized data according to the acquired data assembly rules and distributes the assembled synthesized data.
[0078] Data distribution rules can be set by the user organization's system and configured in the data distribution system. Data distribution rules can be used to specify the scope and content of the synthetic data received by the user organization's system. In some examples, data distribution rules may include data acquisition rules, data filtering rules, etc. For example, if the data distribution rule states that field 'a' in data field B1 is equal to 'a' and the first four bytes of a field in data field B2 are within the range of C1, then the synthetic data conforming to this data distribution rule is the synthetic data required by the user organization's system.
[0079] Data assembly rules can be set by the user organization's system and configured in the data distribution system. Data assembly rules can be used to specify the assembly format of data received by the user organization's system. In some examples, data assembly rules may include, but are not limited to, data deformation rules, data packetization rules, etc.
[0080] User organization systems can receive the synthesized data in real time and use it for quality optimization.
[0081] In step S1044, if the synthesized data does not conform to the data distribution rules, new synthesized data is obtained again from the business data generation model and distributed to the user organization system through the subscribed data distribution system.
[0082] If the synthesized data does not conform to the data distribution rules, the synthesized data will not be distributed to the user organization system. Instead, new synthesized data will be obtained from the business data generation model, and then the process will proceed through step S1042 and the judgment of whether the synthesized data conforms to the data distribution rules until the obtained synthesized data conforms to the data distribution rules, at which point step S1043 will be executed.
[0083] Figure 6 A schematic diagram illustrating an example of the data delivery process provided in this application embodiment, such as... Figure 6 As shown, business data from the data source, i.e., real data, can be input into the model training module. The model training module is used to introduce business data from multiple different data sources according to the actual needs of the user organization, and iteratively train the business data generation model using the model training method described in this embodiment to trigger the data generation and distribution module. The data generation and distribution module can be used to set customized data distribution rules, data assembly rules, etc., according to the actual needs of the user organization system, integrating the synthesized data into a standard distribution format of the data distribution system. This data is then connected to the pipeline distribution system through the data publishing terminal. The pipeline distribution system distributes the synthesized data to the corresponding user organization system according to the customizations of different user organization systems. For example, it distributes the synthesized data that meets the actual needs of user organization system 1 to... Figure 6 User organization system 1 will distribute the synthesized data that meets the actual needs of user organization system 2 to... Figure 6 User organization system 2 will distribute the synthesized data that meets the actual needs of user organization system 3 to Figure 6 User organization system 3 will distribute the synthesized data that meets the actual needs of user organization system 4 to Figure 6 User organization system 4 will distribute the synthesized data that meets the actual needs of user organization system 5 to Figure 6 The user organization system 5. Business data from the data source can also be distributed to the data distribution system through different types of data publishing terminals, such as Figure 6 As shown, business data may include transaction flow data. Business data can be distributed to the data distribution system through data distribution terminals such as transfer data distribution terminals, omnichannel data distribution terminals, and QR code data distribution terminals. The data distribution system can distribute business data within the user organization system's permissions to the corresponding user organization system.
[0084] To make it easier to understand, an example is given here to illustrate the process of real-time data distribution. Figure 7 A flowchart illustrating an example of the real-time data delivery process provided in this application embodiment is shown below. Figure 7 As shown, the real-time data delivery process may include steps b1 to b12.
[0085] In step b1, a data output strategy is loaded, and synthetic data is generated from the model using business data. The data output strategy can represent the actual needs of the user.
[0086] In step b2, the synthesized data is formatted so that the formatted synthesized data conforms to the data format supported by the data distribution system.
[0087] In step b3, the formatted composite data is input into a preset database.
[0088] In step b4, the synthesized data in the database is published to the data distribution server.
[0089] In step b5, the data distribution system subscribes to the synthesized data.
[0090] In step b6, the data distribution system acquires the synthesized data in real time.
[0091] In step b7, the data distribution rules customized by the user organization system are obtained, and the data distribution rules are configured in the data distribution system.
[0092] In step b8, it is determined whether the synthesized data conforms to the data distribution rules. If it does, proceed to step b9; otherwise, return to step b6.
[0093] In step b9, the data assembly rules of the user organization system are obtained, and the data assembly rules are configured in the data distribution system.
[0094] In step b10, the assembled synthetic data is issued according to the data assembly rules.
[0095] In step b11, the availability of the communication line of the user organization system is checked. If available, proceed to step b12; if unavailable, return to step b6.
[0096] In step b12, the synthesized data is sent to the user organization system.
[0097] The specific details of steps b1 to b12 above can be found in the relevant descriptions in the above embodiments, and will not be repeated here.
[0098] The data delivery method provided in this application embodiment can connect the generated synthetic data to the existing real-time delivery channel for real data, i.e., business data, thereby reducing the transformation cost of user organization systems.
[0099] The second aspect of this application provides a data distribution device. Figure 8 This is a schematic diagram of the structure of a data transmission device provided in an embodiment of this application, as shown below. Figure 8 As shown, the data distribution device 200 includes a model training module 201 and a data distribution module 202.
[0100] The model training module 201 can be used to initialize a generative adversarial network (GAN) based on business data from multiple data sources. The GAN includes a generator and a discriminator. The data sources are obtained based on the data requirements of the user organization system. It can also be used to extract correlation vectors from the business data based on the acquired input requirement parameters. These correlation vectors represent the distribution of effective fields in the business data. Furthermore, it can load these correlation vectors into the generator and iteratively train the GAN until it meets the iteration requirements. The generator in the GAN that meets the iteration requirements is then identified as the business data generation model. The synthesized data generated by the generator and the business data serve as inputs to the discriminator, and the output information of the discriminator is used to train the generator.
[0101] The data distribution module 202 can be used to generate synthetic data by utilizing business data to create models and acquire data output strategies, and then distribute the synthetic data to user organization systems through the subscribed data distribution system.
[0102] In this embodiment, a generative adversarial network (GAN) can be initialized based on business data from a data source that matches the data requirements of the user organization system. Based on input requirements parameters, a correlation vector is extracted to represent the distribution of valid fields in the business data from multiple data sources. By loading this correlation vector into the generator within the GAN, the synthesized data output by the generator conforms to the distribution of valid fields in the business data. The generator and discriminator in the GAN train against each other. After multiple iterations, the GAN satisfies the iteration requirements. The generator in the adversarial GAN that meets these requirements is used as the business data generation model, and synthesized data is output according to a data output strategy. This synthesized data is correlated with the business data from the data source matching the user organization system's data requirements and conforms to the distribution of valid fields in the business data from the introduced data source, thus reflecting the characteristics of the business data required by the user organization system. The business data generation model can continuously generate synthesized data that reflects the characteristics of the business data required by the user organization system, meeting the quality and quantity requirements of the user organization system for quality optimization operations, thereby improving the effectiveness of the user organization system's quality optimization operations. Furthermore, the synthetic data output by the business data generation model is not real data and does not contain real information, thus preventing the leakage of users' sensitive information and ensuring data security.
[0103] In some embodiments, the model training module 201 may be specifically used to: extract data features from business data; integrate and filter the data features according to input requirement parameters to obtain effective data features; and generate correlation vectors based on the effective data features.
[0104] In some examples, the model training module 201 can be specifically used to: integrate and filter data features using a long short-term memory network to obtain effective data features.
[0105] In some embodiments, the model training module 201 may be specifically used to: load the correlation vector into the generator, and generate synthetic data by the generator; mix the synthetic data and business data and input them into the discriminator, and the discriminator samples the data domain of the synthetic data and business data, randomly swaps the fields in the data domain of different data, and generates combined data; the discriminator outputs result information based on the combined data; and uses the result information as training parameters to train the generator until the generative adversarial network meets the iteration requirements.
[0106] In some examples, the model training module 201 may be specifically used to: evaluate the correlation of the combined data by the discriminator to obtain the data correlation parameters; and generate result information by the discriminator based on the combined data with the highest correlation indicated by the data correlation parameters.
[0107] In some embodiments, the data distribution module 202 may be specifically used to: use a business data generation model and a data output strategy to output synthetic data that conforms to the data output strategy; format the synthetic data that conforms to the data output strategy into synthetic data in a data format supported by the data distribution system; and, if the synthetic data conforms to the data distribution rules configured by the data distribution system, assemble the synthetic data according to the acquired data assembly rules through the data distribution system, and distribute the assembled synthetic data.
[0108] In some examples, the data delivery module 202 can be specifically used to: obtain new synthetic data from the business data generation model and deliver it to the user organization system through the subscribed data distribution system when the synthetic data does not conform to the data delivery rules.
[0109] A third aspect of this application also provides an electronic device. Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 9 As shown, the electronic device 300 includes a memory 301, a processor 302, and a computer program stored in the memory 301 and executable on the processor 302.
[0110] In some examples, the processor 302 described above may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits that may be configured to implement the embodiments of this application.
[0111] Memory 301 may include read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical, or other physical / tangible memory storage device. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the data delivery method according to embodiments of this application.
[0112] The processor 302 reads the executable program code stored in the memory 301 to run the computer program corresponding to the executable program code, so as to implement the data distribution method in the above embodiment.
[0113] In some examples, the electronic device 300 may also include a communication interface 303 and a bus 304. For example, Figure 9As shown, the memory 301, processor 302, and communication interface 303 are connected through bus 304 and complete communication with each other.
[0114] The communication interface 303 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application. Input devices and / or output devices can also be connected through the communication interface 303.
[0115] Bus 304 includes hardware, software, or both, that couples components of electronic device 300 together. For example, and not limitingly, bus 304 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-E) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 304 may include one or more buses. Although specific buses are described and illustrated in the embodiments of this application, this application considers any suitable bus or interconnection.
[0116] A fourth aspect of this application also provides a computer-readable storage medium storing computer program instructions. When these computer program instructions are executed by a processor, they can implement the data distribution method described in the above embodiments and achieve the same technical effect. To avoid repetition, further details are omitted here. The aforementioned computer-readable storage medium may include non-transitory computer-readable storage media, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, etc., and is not limited thereto.
[0117] This application provides a computer program product. When the instructions in this computer program product are executed by the processor of an electronic device, the electronic device performs the data distribution method described in the above embodiments and achieves the same technical effect. To avoid repetition, further details are omitted here.
[0118] It should be clarified that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. For the device embodiments, equipment embodiments, computer-readable storage medium embodiments, and computer program product embodiments, the relevant parts can be referred to the description section of the method embodiments. This application is not limited to the specific steps and structures described above and shown in the figures. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application. Furthermore, for the sake of brevity, detailed descriptions of known methods and techniques are omitted here.
[0119] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0120] Those skilled in the art will understand that the above embodiments are exemplary and not restrictive. Different technical features appearing in different embodiments can be combined to achieve beneficial effects. Based on a study of the drawings, specification, and claims, those skilled in the art should be able to understand and implement other variations of the disclosed embodiments. In the claims, the term "comprising" does not exclude other means or steps; the quantifier "a" does not exclude a plurality; the terms "first" and "second" are used to identify names and not to indicate any particular order. No reference numerals in the claims should be construed as limiting the scope of protection. The functionality of multiple parts appearing in the claims can be implemented by a single hardware or software module. The appearance of certain technical features in different dependent claims does not mean that these technical features cannot be combined to achieve beneficial effects.
Claims
1. A data distribution method, characterized in that, include: Based on business data from multiple data sources, a generative adversarial network (GAN) is initialized. The GAN includes a generator and a discriminator. The data sources are obtained based on the data requirements of the user organization's system. Based on the obtained input requirement parameters, the correlation vector of the business data is extracted, and the correlation vector represents the distribution of the effective fields in the business data; The correlation vector is loaded into the generator, and the generative adversarial network is iteratively trained until the generative adversarial network meets the iteration requirements. The generator in the generative adversarial network that meets the iteration requirements is determined as the business data generation model. The synthetic data generated by the generator and the business data are used as inputs to the discriminator. The output information of the discriminator is used to train the generator. The synthetic data is correlated with the business data of the data source that matches the data requirements and conforms to the distribution of effective fields in the business data of the introduced data source, reflecting the characteristics of the business data required by the user organization system. Using the business data generation model and the acquired data output strategy, synthetic data is generated and distributed to the user organization system through the subscribed data distribution system.
2. The method according to claim 1, characterized in that, The step of extracting the correlation vector of the business data based on the acquired input requirement parameters includes: Extract the data features of the business data; Based on the input requirements parameters, the data features are integrated and filtered to obtain effective data features; The correlation vector is generated based on the effective data features.
3. The method according to claim 2, characterized in that, The process of integrating and filtering the data features to obtain effective data features includes: The data features are integrated and filtered using a long short-term memory network to obtain the effective data features.
4. The method according to claim 1, characterized in that, The step of loading the correlation vector into the generator and iteratively training the generative adversarial network until the generative adversarial network meets the iteration requirements includes: The correlation vector is loaded into the generator, and the generator generates the synthetic data. The synthesized data and the business data are mixed and input into the discriminator. The discriminator samples the data domains of the synthesized data and the business data, and randomly swaps the fields in the data domains of different data to generate combined data. The discriminator outputs the result information based on the combined data; The resulting information is used as training parameters to train the generator until the generative adversarial network meets the iteration requirements.
5. The method according to claim 4, characterized in that, The discriminator outputs the result information based on the combined data, including: The discriminator performs a correlation evaluation on the combined data to obtain data correlation parameters; The discriminator generates the result information based on the combined data with the highest correlation indicated by the data correlation parameter.
6. The method according to claim 1, characterized in that, The step of generating synthetic data using the business data generation model and the acquired data output strategy, and then distributing it to the user organization system through the subscribed data distribution system, includes: Using the business data generation model and data output strategy, the synthesized data conforming to the data output strategy is output; The synthesized data conforming to the data output strategy is formatted into a data format supported by the data distribution system; If the synthesized data conforms to the data distribution rules configured by the data distribution system, the data distribution system assembles the synthesized data according to the acquired data assembly rules and distributes the assembled synthesized data.
7. The method according to claim 6, characterized in that, Also includes: If the synthesized data does not conform to the data distribution rules, new synthesized data is obtained from the business data generation model and distributed to the user organization system through the subscribed data distribution system.
8. A data transmission device, characterized in that, include: The model training module is used to initialize a generative adversarial network (GAN) based on business data from multiple data sources. The GAN includes a generator and a discriminator. The data sources are obtained based on the data requirements of the user organization system. And, based on the acquired input requirement parameters, to extract the correlation vector of the business data, wherein the correlation vector represents the distribution of valid fields in the business data; And, for loading the correlation vector into the generator, iteratively training the generative adversarial network until the generative adversarial network meets the iteration requirements, and determining the generator in the generative adversarial network that meets the iteration requirements as the business data generation model, wherein the synthetic data generated by the generator and the business data are used as inputs to the discriminator, the result information output by the discriminator is used to train the generator, the synthetic data is correlated with the business data of the data source matching the data requirements, and conforms to the distribution of effective fields in the business data of the introduced data source, reflecting the characteristics of the business data required by the user organization system; The data distribution module is used to generate synthetic data using the business data generation model and the acquired data output strategy, and distribute it to the user organization system through the subscribed data distribution system.
9. An electronic device, characterized in that, include: Processor and memory storing computer program instructions; When the processor executes the computer program instructions, it implements the data delivery method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the data delivery method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Adversarial sample generation method of neural network model and related equipment
CN114677556A