Transaction detection message generation method and device, electronic equipment and storage medium

By acquiring real transaction data, determining transaction field characteristics, and using a data synthesizer to generate simulated transaction data, the problem of insufficient coverage in existing transaction detection messages is solved, and the accuracy of transaction system detection is improved.

CN120849294APending Publication Date: 2025-10-28CHINA UNIONPAY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511233174.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing rule-based transaction detection messages are insufficient to cover real-world transaction scenarios, resulting in inadequate detection accuracy.

Method used

By acquiring real transaction data, determining the data characteristics of transaction fields, using a data synthesizer to generate simulated transaction data, and then merging it into target transaction data, a transaction probe message is finally generated.

Benefits of technology

This improves the similarity between simulated trading scenarios and real-world trading scenarios, thereby enhancing the accuracy of testing the trading system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849294A_ABST
    Figure CN120849294A_ABST
Patent Text Reader

Abstract

The invention discloses a transaction detection message generation method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring first transaction data of a real transaction; based on transaction data corresponding to a plurality of transaction fields in the first transaction data, determining field data features corresponding to the plurality of transaction fields; splitting the first transaction data based on the plurality of field data features to obtain a plurality of second transaction data; for each piece of second transaction data, determining a data synthesizer corresponding to the second transaction data based on a corresponding relationship between the field data feature corresponding to the second transaction data and the data synthesizer, and generating simulation transaction data corresponding to the second transaction data through the data synthesizer corresponding to the second transaction data; fusing the simulation transaction data corresponding to the multiple pieces of second transaction data to obtain target transaction data; and generating a transaction detection message based on the target transaction data. According to the embodiment of the invention, the accuracy of transaction system detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data processing technology, and in particular relates to a method, apparatus, electronic device and storage medium for generating transaction probe messages. Background Art

[0002] As the core hub of financial transactions, the transaction system works closely with banks (fund clearing parties) and merchants (transaction initiators or goods / service providers) to jointly complete business processes such as payment, settlement, and risk control. To ensure the normal operation of the transaction system, transaction probe messages are typically used to test the system and detect potential problems in advance.

[0003] Currently, transaction probe messages are typically generated based on rules. However, rule-based transaction probe messages are relatively simple and cannot cover both normal and abnormal transaction scenarios in real-world situations. This results in a significant discrepancy between the transaction scenarios they reflect and the actual transaction scenarios, thereby reducing the accuracy of transaction system detection. Summary of the Invention

[0004] This application provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating transaction probe messages, which can improve the similarity between the simulated transaction scenario reflected by the transaction probe message and the actual transaction scenario, thereby improving the accuracy of detecting the transaction system.

[0005] In a first aspect, embodiments of this application provide a method for generating a transaction probe message, the method comprising:

[0006] Obtain first-hand transaction data of real transactions;

[0007] Based on the transaction data corresponding to multiple transaction fields in the first transaction data, determine the field data features corresponding to the multiple transaction fields respectively;

[0008] Based on the field data characteristics corresponding to the multiple transaction fields, the first transaction data is split to obtain multiple second transaction data;

[0009] For each of the second transaction data, a data synthesizer corresponding to the second transaction data is determined based on the correspondence between the field data features of the transaction field corresponding to the second transaction data and the data synthesizer;

[0010] For each second transaction data, simulated transaction data corresponding to the second transaction data is generated through a data synthesizer corresponding to the second transaction data;

[0011] The simulated transaction data corresponding to the multiple second transaction data are fused to obtain the target transaction data;

[0012] Based on the target transaction data, a transaction probe message is generated.

[0013] In some implementations, determining the field data features corresponding to each of the multiple transaction fields in the first transaction data includes:

[0014] For any two transaction fields, the dependency strength between the two transaction fields is calculated based on the transaction data corresponding to the two transaction fields in the first transaction data.

[0015] For any of the transaction fields, the data distribution corresponding to the transaction field is determined based on the transaction data corresponding to the transaction field in the first transaction data;

[0016] Based on the dependency strength and the data distribution, the field data characteristics corresponding to the multiple transaction fields are determined respectively.

[0017] In some implementations, determining the field data features corresponding to the plurality of transaction fields based on the dependency strength and the data distribution includes:

[0018] If the dependency strength between two transaction fields meets the dependency condition, the field data features corresponding to the two transaction fields are determined as field dependency features;

[0019] The remaining transaction fields that do not satisfy the field dependency characteristics among the plurality of transaction fields are identified as the first type of transaction fields;

[0020] If the difference between the data distribution corresponding to the first type of transaction field and the reference distribution meets the difference condition, the field data feature corresponding to the first type of transaction field is determined as the first distribution feature.

[0021] If the difference between the data distribution corresponding to the first type of transaction field and the reference distribution does not meet the difference condition, the field data feature corresponding to the first type of transaction field is determined as the second distribution feature.

[0022] In some implementations, the field data feature of the transaction field corresponding to the second transaction data is the field dependency feature, the data synthesizer corresponding to the second transaction data is a first Gaussian synthesizer, the simulated transaction data includes first simulated transaction data, and the step of generating simulated transaction data corresponding to the second transaction data through the data synthesizer corresponding to the second transaction data includes:

[0023] The first mean and covariance of the second transaction data are calculated using the first Gaussian synthesizer.

[0024] The first simulated transaction data is generated based on the first mean and the covariance.

[0025] In some implementations, the field data feature of the transaction field corresponding to the second transaction data is the first distribution feature, the data synthesizer corresponding to the second transaction data is a second Gaussian synthesizer, the simulated transaction data includes the second simulated transaction data, and the step of generating simulated transaction data corresponding to the second transaction data through the data synthesizer corresponding to the second transaction data includes:

[0026] The second mean and standard deviation of the second transaction data are calculated using the second Gaussian synthesizer.

[0027] The second simulated transaction data is generated based on the second mean and the standard deviation.

[0028] In some embodiments, the field data feature of the transaction field corresponding to the second transaction data is the second distribution feature, the data synthesizer corresponding to the second transaction data is a conditional table generative adversarial network synthesizer, the simulated transaction data includes third simulated transaction data, and the step of generating simulated transaction data corresponding to the second transaction data through the data synthesizer corresponding to the second transaction data includes:

[0029] The conditional table generative adversarial network synthesizer is trained based on the second transaction data to obtain a trained target conditional table generative adversarial network synthesizer.

[0030] Obtain random noise data;

[0031] The adversarial network synthesizer generates the third simulated transaction data based on the random noise data using the target condition table.

[0032] In some implementations, generating a transaction probe message based on the target transaction data includes:

[0033] Based on the target transaction data, an initial transaction probe message is generated;

[0034] The initial transaction probe message is optimized according to the target rules to obtain the transaction probe message.

[0035] In some implementations, the initial transaction probe message includes message fields corresponding to the plurality of transaction fields, and the optimization of the initial transaction probe message according to the target rule to obtain the transaction probe message includes:

[0036] Query the field constraint conditions corresponding to the multiple message fields from the field constraint knowledge base, wherein the field constraint knowledge base includes the multiple field constraint conditions;

[0037] If the content of the first target message field does not meet the target field constraint conditions corresponding to the first target message field, the content of the first target message field is modified based on the target field constraint conditions to obtain the transaction probe message, wherein the first target message field is any one of the plurality of message fields.

[0038] In some implementations, the initial transaction probe message includes message fields corresponding to the plurality of transaction fields, and the optimization of the initial transaction probe message according to the target rule to obtain the transaction probe message includes:

[0039] The personalized requirements for obtaining the second target message field are obtained, where the second target message field is any one of the plurality of message fields;

[0040] Based on the personalized requirements, the field content of the second target message field is rewritten to obtain the transaction detection message.

[0041] Secondly, embodiments of this application provide a device for generating transaction probe messages, the device comprising:

[0042] The acquisition module is used to acquire the first transaction data of the actual transaction.

[0043] The determination module is used to determine the field data features corresponding to the multiple transaction fields based on the transaction data corresponding to the multiple transaction fields in the first transaction data.

[0044] The splitting module is used to split the first transaction data based on the field data features corresponding to the multiple transaction fields to obtain multiple second transaction data.

[0045] The determining module is further configured to, for each second transaction data, determine the data synthesizer corresponding to the second transaction data based on the correspondence between the field data features of the transaction field corresponding to the second transaction data and the data synthesizer;

[0046] The generation module is used to generate simulated transaction data corresponding to each second transaction data by using a data synthesizer corresponding to the second transaction data;

[0047] The fusion module is used to fuse the simulated transaction data corresponding to the multiple second transaction data respectively to obtain the target transaction data;

[0048] The generation module is also used to generate a transaction probe message based on the target transaction data.

[0049] Thirdly, embodiments of this application provide an electronic device, which includes: a processor and a memory storing computer program instructions;

[0050] When the processor executes the computer program instructions, it implements any of the possible implementations of the first aspect described above.

[0051] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the method in any of the possible implementations of the first aspect described above.

[0052] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform a method as described in any of the possible implementations of the first aspect above.

[0053] This application embodiment first determines the field data features corresponding to multiple transaction fields based on the transaction data corresponding to multiple transaction fields in the first transaction data. Then, based on the field data features corresponding to the multiple transaction fields, the first transaction data is split to obtain multiple second transaction data that match the field data features of the transaction fields. Next, for each second transaction data, based on the correspondence between the field data features of the corresponding transaction fields and the data synthesizer, a data synthesizer corresponding to the second transaction data is determined, resulting in a data synthesizer that matches both the field data features and the second transaction features. Based on this, simulated transaction data corresponding to the second transaction data is generated using the data synthesizer corresponding to the second transaction data. That is, by using the data synthesizer that matches both the field data features and the second transaction data, data synthesis is performed on the second transaction data that matches the field data features to obtain simulated transaction data, thereby improving the similarity between the simulated transaction data and the second transaction data. Thus, by fusing the simulated transaction data corresponding to multiple second transaction data, simulated target transaction data is obtained, which improves the similarity between the simulated target transaction data and the first transaction data of the real transaction. Thus, by generating transaction probe messages based on target transaction data, the similarity between the simulated transaction scenario reflected in the transaction probe message and the actual transaction scenario can be improved, thereby improving the accuracy of detecting the transaction system. Attached Figure Description

[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0055] Figure 1 This is a schematic diagram of a transaction detection scenario provided in an embodiment of this application;

[0056] Figure 2 This is a flowchart illustrating a method for generating a transaction probe message according to an embodiment of this application;

[0057] Figure 3 This is a flowchart illustrating another method for generating transaction probe messages provided in an embodiment of this application;

[0058] Figure 4 This is a schematic diagram of a method for generating a transaction probe message provided in an embodiment of this application;

[0059] Figure 5 This is a schematic diagram of the structure of a transaction detection message generation device provided in an embodiment of this application;

[0060] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0061] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.

[0062] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.

[0063] Furthermore, the acquisition, storage, use, and processing of data in this application's technical solution all comply with relevant national laws and regulations.

[0064] As described in the background section, transaction probe messages are currently typically generated based on rules. However, rule-based transaction probe messages are relatively simple and cannot cover both normal and abnormal transaction scenarios in real-world situations. This results in a significant discrepancy between the transaction scenarios they reflect and the actual transaction scenarios, thereby reducing the accuracy of transaction system detection.

[0065] Therefore, to address the problems of existing technologies, embodiments of this application provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating transaction probe messages, thereby improving the quality of transaction probe messages and increasing the similarity between the simulated transaction scenario reflected by the transaction probe message and the actual transaction scenario, thus improving the accuracy of transaction system detection. The method for generating transaction probe messages can be applied to transaction probe scenarios. Furthermore, the method for generating transaction probe messages can be executed by a simulated transaction system. After generating the transaction probe message, the simulated transaction system can send the transaction probe message to a health detection module, enabling the health detection module to detect the transaction system based on the transaction probe message. The health detection module can include a health detection acquiring-side simulator and a health detection issuing-side simulator. The health detection acquiring-side simulator can be used to simulate merchants, and the health detection issuing-side simulator can be used to simulate issuing banks.

[0066] Based on this, the schematic diagram of a transaction detection scenario provided in the embodiments of this application can be as follows: Figure 1 As shown. The transaction probing process in this transaction probing scenario can be executed by the transaction probing system.

[0067] like Figure 1As shown, the transaction detection system may include a health detection acquiring-side simulator 11, a transaction system 12, and a health detection issuing-side simulator 13.

[0068] As an example, the health detection acquiring simulator 11 can first generate a transaction detection message P0, and then send the transaction detection message P0 to the transaction system 12. After receiving the transaction detection message P0, the transaction system 12 can process the transaction detection message P0 to obtain message P1, and send message P1 to the health detection issuing simulator 13. The health detection issuing simulator 13 can generate a response message P2 based on message P1, and send the response message P2 to the transaction system 12. Then, the transaction system 12 can process the response message P2 to obtain a response message P3, and return the response message P3 to the health detection acquiring simulator 11, thereby realizing the entire transaction detection process.

[0069] The method for generating transaction probe messages provided in the embodiments of this application will be described below.

[0070] Figure 2 The illustration shows a flowchart of a method for generating a transaction probe message according to an embodiment of this application.

[0071] like Figure 2 As shown in the embodiments of this application, the method for generating transaction probe messages includes the following steps:

[0072] S210, Obtain the first transaction data of the actual transaction;

[0073] S220. Based on the transaction data corresponding to multiple transaction fields in the first transaction data, determine the field data characteristics corresponding to the multiple transaction fields respectively;

[0074] S230. Based on the field data characteristics corresponding to multiple transaction fields, the first transaction data is split to obtain multiple second transaction data;

[0075] S240. For each second transaction data, based on the correspondence between the field data features of the transaction field corresponding to the second transaction data and the data synthesizer, determine the data synthesizer corresponding to the second transaction data.

[0076] S250. For each second transaction data, generate simulated transaction data corresponding to the second transaction data through the data synthesizer corresponding to the second transaction data;

[0077] S260. Merge the simulated transaction data corresponding to multiple second transaction data to obtain the target transaction data;

[0078] S270. Generate a transaction probe message based on the target transaction data.

[0079] This application embodiment first determines the field data features corresponding to multiple transaction fields based on the transaction data corresponding to multiple transaction fields in the first transaction data. Then, based on the field data features corresponding to the multiple transaction fields, the first transaction data is split to obtain multiple second transaction data that match the field data features of the transaction fields. Next, for each second transaction data, based on the correspondence between the field data features of the corresponding transaction fields and the data synthesizer, a data synthesizer corresponding to the second transaction data is determined, resulting in a data synthesizer that matches both the field data features and the second transaction features. Based on this, simulated transaction data corresponding to the second transaction data is generated using the data synthesizer corresponding to the second transaction data. That is, by using the data synthesizer that matches both the field data features and the second transaction data, data synthesis is performed on the second transaction data that matches the field data features to obtain simulated transaction data, thereby improving the similarity between the simulated transaction data and the second transaction data. Thus, by fusing the simulated transaction data corresponding to multiple second transaction data, simulated target transaction data is obtained, which improves the similarity between the simulated target transaction data and the first transaction data of the real transaction. Thus, by generating transaction probe messages based on target transaction data, the similarity between the simulated transaction scenario reflected in the transaction probe message and the actual transaction scenario can be improved, thereby improving the accuracy of detecting the transaction system.

[0080] The specific implementation methods for each of the above steps are described below.

[0081] In some embodiments, in S210, the first transaction data may be transaction data that has been preprocessed and can be used for subsequent data splitting, data synthesis and other processing.

[0082] As an example, a simulated trading system may include a transaction preprocessing module. This module can be used to acquire the raw transaction data corresponding to real transactions, and to preprocess the acquired raw transaction data to obtain the first transaction data.

[0083] As a more concrete example, the transaction preprocessing module can obtain raw transaction data corresponding to real transactions from different channels such as logs and real-time data sources. The raw transaction data obtained from logs can be log transaction data, while the raw transaction data obtained from real-time data sources can be streaming transaction data. Both log transaction data and streaming transaction data can be in message form, i.e., transaction messages. Furthermore, different transaction messages may follow different message specifications (or protocols), which may employ different data formats and have different field structures and content meanings. These different data formats can include Extensible Markup Language (XML) format, JavaScript Object Notation (JSON) format, binary format, etc.

[0084] Based on this, the transaction preprocessing module can first filter the raw transaction data obtained from different channels to remove irrelevant data from a large amount of transaction data and retain data with the same transaction specifications. Then, it can perform format conversion on the filtered raw transaction data to convert the raw transaction data with different data formats into structured data, thus obtaining the first transaction data. This structured data can be data in database table format, Excel format, comma-separated values ​​(CSV) format, etc.

[0085] In some embodiments, in S220, the first transaction data may include multiple transaction fields and transaction data corresponding to each of the multiple transaction fields. The field data characteristics corresponding to different transaction fields may be the same or different. These field data characteristics can be the attributes and properties exhibited by each transaction field. Field data characteristics can be used to describe the properties, structure, statistical attributes, etc., of the transaction data corresponding to each transaction field. Thus, by analyzing the properties, structure, statistical attributes, etc., of the transaction data corresponding to each of the multiple transaction fields in the first transaction data, the field data characteristics corresponding to each of the multiple transaction fields can be determined.

[0086] Furthermore, in the first transaction data, different transaction fields may have dependencies. If two transaction fields have dependencies, the common field data features corresponding to the two transaction fields can be determined based on the dependencies, thereby improving the accuracy of the field data features.

[0087] Therefore, in order to improve the accuracy of field data features, in some embodiments, the above-mentioned S220 may specifically include:

[0088] For any two transaction fields, calculate the dependency strength between the two transaction fields based on the transaction data corresponding to the two transaction fields in the first transaction data.

[0089] For any transaction field, the data distribution corresponding to the transaction field is determined based on the transaction data corresponding to the transaction field in the first transaction data;

[0090] Based on dependency strength and data distribution, the field data characteristics corresponding to multiple transaction fields are determined.

[0091] Here, the strength of the dependency between any two transaction fields can be assessed using correlation analysis, mutual information, causal discovery methods, etc. Correlation analysis can include correlation analysis based on Spearman's correlation coefficient, correlation analysis based on Pearson's correlation coefficient, etc. Spearman's correlation coefficient measures the monotonic relationship between two variables.

[0092] The following section uses Spearman correlation coefficient as an example to explain in detail how to assess the strength of dependence between any two transaction fields.

[0093] First, determine the transaction data (i.e., the field values) corresponding to the two transaction fields in the first transaction data. Then, for the same transaction record, establish the combination relationship between the field values ​​corresponding to the two transaction fields to obtain n sets of observations (x1, y1), (x2, y2), ... (x n ,y n Here, n can be the number of transaction records in the first transaction data, i.e., the sample size. Then, the field values ​​of each transaction field are sorted in ascending order and assigned a rank, and the rank difference d corresponding to each group of observations is calculated. i (i = 1, ..., n).

[0094] Based on this, the dependence strength r between the two transaction fields can be determined by substituting the sample size n and the n rank differences into the following formula (1). s :

[0095]

[0096] In addition, for any given transaction field, the data distribution corresponding to that single transaction field can be determined based on the transaction data corresponding to the transaction field in the first transaction data by means of statistical descriptive indicators, empirical cumulative distribution functions, etc.

[0097] Subsequently, by comprehensively analyzing the dependency strength and data distribution of multiple transaction fields in the first transaction data, the field data characteristics corresponding to each of the multiple transaction fields can be determined.

[0098] This application embodiment improves the accuracy of field data features by comprehensively considering the dependency strength between any two transaction fields and the data distribution of any single transaction field to jointly determine the field data features corresponding to multiple transaction fields.

[0099] Building upon this, to further improve the accuracy of field data features, in some embodiments, the above-mentioned determination of field data features corresponding to multiple transaction fields based on dependency strength and data distribution may specifically include:

[0100] If the dependency strength between two transaction fields meets the dependency condition, the field data features corresponding to the two transaction fields are determined as field dependency features;

[0101] The remaining transaction fields that do not satisfy the field dependency feature among multiple transaction fields are identified as the first type of transaction fields;

[0102] If the difference between the data distribution corresponding to the first type of transaction field and the reference distribution meets the difference condition, the field data feature corresponding to the first type of transaction field is determined as the first distribution feature;

[0103] If the difference between the data distribution corresponding to the first type of transaction field and the reference distribution does not meet the difference condition, the field data feature corresponding to the first type of transaction field is determined as the second distribution feature.

[0104] Here, the dependency strength between two transaction fields satisfies the dependency condition if the dependency strength between the two transaction fields is greater than a preset strength threshold. If the dependency strength between two transaction fields is greater than the preset strength threshold, then it can be determined that the two transaction fields are dependent fields, and the field data features corresponding to the two transaction fields are field dependency features. For example, assuming the preset strength threshold is 0.5, then in the above example, if the dependency strength |r s If |>0.5, then the dependency strength between the two transaction fields meets the dependency condition, that is, the dependency relationship between the two transaction fields is relatively strong, and thus the field data features corresponding to the two transaction fields can be determined as field dependency features.

[0105] As an example, this application embodiment can first determine whether any two transaction fields are dependent fields. If so, it can be determined that the field data characteristics of the two transaction fields are both field dependency characteristics. For the remaining transaction fields that do not have field dependency characteristics, they can be temporarily identified as first-type transaction fields, and the data distribution of each first-type transaction field can be determined by statistical descriptive indicators, empirical cumulative distribution functions, etc.

[0106] Furthermore, the first distribution characteristic can be a separable distribution characteristic of the marginal distribution, such as the theoretical normal distribution or the log-normal distribution. Therefore, the reference distribution can be the theoretical normal distribution, the log-normal distribution, etc. Additionally, the second distribution characteristic can be a complex distribution characteristic of the marginal distribution, such as a multimodal distribution, a heavy-tailed distribution, or a non-Gaussian distribution.

[0107] As an example, the difference between the data distribution corresponding to the first type of transaction field and the reference distribution can satisfy the difference condition if the distribution difference between the data distribution and the reference distribution is less than a preset difference threshold. If the distribution difference is less than the preset difference threshold, the field data feature corresponding to the first type of transaction field can be determined to be a first distribution feature. Conversely, if the distribution difference is greater than or equal to the preset difference threshold, the field data feature corresponding to the first type of transaction field can be determined to be a second distribution feature. That is, in this embodiment of the application, if a transaction field satisfies neither the field dependency feature nor the first distribution feature, it can be determined that the transaction field satisfies the second distribution feature.

[0108] For each type I transaction field, the distributional difference between the type I transaction field and the reference distribution can be determined by methods such as Kullback-Leibler Divergence (KL divergence) and Kolmogorov–Smirnov test (KS test).

[0109] The following section uses KL divergence as an example to explain in detail how to determine the field data characteristics of a single transaction field.

[0110] Assuming the data distribution of a certain transaction field is denoted as P(x) and the reference distribution is denoted as Q(x), the KL divergence D of the transaction field can be calculated using the following formula (2). KL (P|Q):

[0111]

[0112] Assuming a preset difference threshold of 0.5, if the KL divergence is less than 0.5 (close to 0), it can be determined that the data distribution of the transaction field differs little from the reference distribution, thus identifying the data feature of the transaction field as a first distribution feature. If the KL divergence is greater than or equal to 0.5, it can be determined that the data distribution of the transaction field differs significantly from the reference distribution, indicating a more complex data distribution, such as a heavy-tailed distribution, thus identifying the data feature of the transaction field as a second distribution feature.

[0113] This application embodiment first determines whether a transaction field satisfies the field dependency feature based on the dependency strength between different transaction fields, then performs single-field data distribution analysis on the transaction fields that do not satisfy the field dependency feature to determine whether the transaction field satisfies the first distribution feature, and then determines the transaction fields that neither satisfy the field dependency feature nor the first distribution feature as satisfying the second distribution feature. This can further refine the field data features of the transaction fields, thereby further improving the accuracy of the field data features.

[0114] In some embodiments, in S230, the simulated trading system may further include a data splitting module. The data splitting module can be used to split the first trading data according to field data characteristics to obtain multiple second trading data sets. Specifically, if two trading fields satisfy a field dependency characteristic, i.e., the two trading fields are dependent fields, the trading data corresponding to these two trading fields can be split together from the first trading data to obtain one second trading data set. If a trading field satisfies a first distribution characteristic, the trading data corresponding to that trading field can be split separately from the first trading data to obtain one second trading data set. After all trading data corresponding to all field dependency characteristics and the first distribution characteristic have been split from the first trading data, the remaining trading data in the first trading data can be collectively determined as one second trading data set.

[0115] Furthermore, the data characteristics of the second transaction data can be the same as the field data characteristics of the corresponding transaction field. For example, if the field data characteristics of the corresponding transaction field of the second transaction data satisfy the field dependency characteristic, then the second transaction data can satisfy the field dependency characteristic. If the field data characteristics of the corresponding transaction field of the second transaction data satisfy the first distribution characteristic, then the second transaction data can satisfy the first distribution characteristic. If the field data characteristics of the corresponding transaction field of the second transaction data satisfy the second distribution characteristic, then the second transaction data can satisfy the second distribution characteristic.

[0116] In some embodiments, in S240, the data synthesizer can be used to generate simulated transaction data similar to the second transaction data. The data synthesizer may include a Gaussian synthesizer (i.e., a Copulas synthesizer) and a Conditional Tabular Generative Adversarial Network (CTGAN) synthesizer. The core idea of ​​the Gaussian synthesizer is to separate the marginal distributions of random variables from their joint dependency structure, which is driven by a multivariate Gaussian distribution. The linear correlation between variables is characterized by a Gaussian correlation matrix. This synthesis is mainly used to learn the distribution of individual variables and the correlation between variables. That is, the Gaussian synthesizer mainly generates simulated transaction data by learning the distributions within different fields of the production transaction. The CTGAN synthesizer is a variant of the Generative Adversarial Network (GAN) specifically designed for tabular data. Its core objective is to enable the generator to learn the joint distribution of real data and to distinguish generated data from real data through a discriminator, ultimately generating synthetic data with statistical characteristics consistent with the real data. Compared to ordinary GAN networks, CTGAN introduces a conditional variable y into the inputs of its generator and discriminator. By using y (such as class labels, text descriptions, etc.) as additional input, the model can generate data that matches the conditions. The CTGAN synthesizer mainly generates simulated transaction data by learning the correlation information between different fields within a single transaction.

[0117] Based on the aforementioned characteristics of different data synthesizers, different data synthesizers can be applied to different use cases. Specifically, for single-field distribution characteristics, if the field data feature of the transaction field corresponding to the second transaction data is a first distribution feature, that is, the field data feature belongs to a marginally separable distribution feature (such as normal or log-normal), or a distribution feature that needs to strictly match a preset distribution, then a Gaussian synthesizer can be used to generate simulated transaction data corresponding to the second transaction data; if the field data feature of the transaction field corresponding to the second transaction data is a second distribution feature, that is, the field data feature belongs to a marginally complex distribution feature (multimodal, heavy-tailed, non-Gaussian), in order to generate realistic transaction details (such as transaction amount fluctuations), then a CTGAN synthesizer can be used to generate simulated transaction data corresponding to the second transaction data. Furthermore, regarding the strength of dependencies between fields, if the fields satisfy the field dependency characteristics, i.e., there is a strong linear / monotonic dependency (such as merchant number and merchant type), then a strictly conformal dependency relationship is required, and a Gaussian synthesizer can be used to generate simulated transaction data corresponding to the second transaction data. If the fields do not satisfy the field dependency characteristics, i.e., the dependency between fields is non-linear (such as card number), or the dependency relationship is ambiguous, then a Gaussian synthesizer can be used to generate simulated transaction data corresponding to the second transaction data. In other words, there can be a correspondence between field data characteristics and data synthesizers.

[0118] As an example, if the field data feature of the transaction field corresponding to the second transaction data is a field dependency feature, then the data synthesizer corresponding to the second transaction data can be a first Gaussian synthesizer. If the field data feature of the transaction field corresponding to the second transaction data is a first distribution feature, then the data synthesizer corresponding to the second transaction data can be a second Gaussian synthesizer. If the field data feature of the transaction field corresponding to the second transaction data is a second distribution feature, then the data synthesizer corresponding to the second transaction data can be a conditional table generative adversarial network synthesizer (i.e., a CTGAN synthesizer).

[0119] In some embodiments, in S250, the simulated trading system may further include a transaction generation module, which can generate simulated trading data corresponding to the second trading data through the aforementioned data synthesizer. The simulated trading data may include first simulated trading data corresponding to field dependency features, second simulated trading data corresponding to a first distribution feature, and third simulated trading data corresponding to a second distribution feature. Therefore, a first Gaussian synthesizer can be used to generate the first simulated trading data, a second Gaussian synthesizer can be used to generate the second simulated trading data, and a CTGAN synthesizer can be used to generate the third simulated trading data.

[0120] Based on this, in order to improve the data quality of the first simulated transaction data, in some embodiments, the above-mentioned generation of simulated transaction data corresponding to the second transaction data through a data synthesizer corresponding to the second transaction data may specifically include:

[0121] The first mean and covariance corresponding to the second transaction data are calculated using the first Gaussian synthesizer.

[0122] First simulated transaction data is generated based on the first mean and covariance.

[0123] Here, the transaction generation module can input the second transaction data corresponding to the field dependency features into the first Gaussian synthesizer. The first Gaussian synthesizer first calculates the first mean and covariance corresponding to the second transaction data, and then generates the first simulated transaction data based on the first mean and covariance. Each set of second transaction data satisfying the field dependency features can be independently input into the first Gaussian synthesizer. For example, if the first transaction data includes six second transaction data points, and among these six second transaction data points, two satisfy the field dependency features, three satisfy the first distribution feature, and one satisfies the second distribution feature, then the two second transaction data points satisfying the field dependency features can be independently input into the first Gaussian synthesizer to generate the first simulated transaction data corresponding to each of the two second transaction data points.

[0124] In this embodiment, by independently inputting the second transaction data that meets the field dependency characteristics into the first Gaussian synthesizer, the first Gaussian synthesizer can utilize its statistical conformity to accurately capture the dependency relationship between variables based on the data characteristics of the second transaction data, and perform targeted processing on the second transaction data to obtain the first simulated transaction data, thereby improving the data quality of the first simulated transaction data.

[0125] In addition, to improve the data quality of the second simulated transaction data, in some embodiments, the above-mentioned generation of simulated transaction data corresponding to the second transaction data through a data synthesizer corresponding to the second transaction data may specifically include:

[0126] The second mean and standard deviation corresponding to the second transaction data are calculated using the second Gaussian synthesizer.

[0127] Second simulated trading data is generated based on the second mean and standard deviation.

[0128] Here, the transaction generation module can input the second transaction data corresponding to the first distribution characteristic into the second Gaussian synthesizer. The second Gaussian synthesizer first calculates the second mean and standard deviation corresponding to the second transaction data, and then generates the second simulated transaction data based on the second mean and standard deviation. Specifically, the first second transaction data that satisfies the first distribution characteristic, i.e., whose marginal distribution is separable, can be independently input into the second Gaussian synthesizer. For example, continuing the previous example, three second transaction data that satisfy the first distribution characteristic can be input into the second Gaussian synthesizer separately and independently to generate second simulated transaction data corresponding to each of the three second transaction data.

[0129] In this embodiment of the application, by independently inputting the second transaction data that meets the first distribution characteristics into the second Gaussian synthesizer, the second Gaussian synthesizer can perform targeted processing on the second transaction data based on the data characteristics of the second transaction data to obtain the second simulated transaction data, thereby improving the data quality of the second simulated transaction data.

[0130] In addition, to improve the data quality of the third simulated transaction data, in some embodiments, the above-mentioned generation of simulated transaction data corresponding to the second transaction data through a data synthesizer corresponding to the second transaction data may specifically include:

[0131] The conditional table generative adversarial network synthesizer is trained based on the second transaction data to obtain the trained target conditional table generative adversarial network synthesizer.

[0132] Obtain random noise data;

[0133] An adversarial network synthesizer is generated using a target condition table, and third-level simulated transaction data is generated based on random noise data.

[0134] Here, the CTGAN synthesizer, through adversarial training, can capture complex nonlinear distributions, making the generated data closer to the distribution of real data. Thus, assuming the second transaction data includes transaction time, transaction type, and transaction institution, training the CTGAN synthesizer based on the second transaction data enables the CTGAN synthesizer to learn the transaction distribution across different time periods, transaction types, and institutions.

[0135] Thus, by inputting random noise data into the trained target CTGAN synthesizer, the target CTGAN synthesizer model can generate third simulated transaction data based on random noise data, which has the same or similar transaction distribution among different time periods, different transaction types, and different institutions as the second transaction data, thereby improving the data quality of the third simulated transaction data.

[0136] In some embodiments, in S260, the fusion of multiple simulated transaction data can be performed by merging the multiple simulated transaction data according to the position of the multiple second transaction data in the first transaction data, thereby obtaining the complete simulated transaction data corresponding to the first transaction data, i.e., the target transaction data.

[0137] In some embodiments, in S270, after obtaining the simulated target transaction data, the transaction generation module can further perform message conversion on the target transaction data to obtain transaction probe messages, thereby converting the target transaction data back into sendable transaction information. Each transaction probe message may include one transaction or multiple transactions; this is not limited here.

[0138] Some content in the aforementioned target transaction data may not conform to transaction specifications, resulting in non-compliance of transaction probe messages and affecting the accuracy of using these messages to detect transactions within the system. Transaction specifications define what the transaction fields should contain, what conditions they must meet, and whether different transaction fields should appear simultaneously or not. For example, a transaction specification might stipulate that an ID number consists of 18 digits, with the last digit being an "X," and that the ID number and ID type should appear simultaneously or not. Another example is that if the transaction field is for a transaction institution, the institution's name could include Institution A, Institution B, Institution C, etc.

[0139] Therefore, in order to improve the message quality of transaction probe messages and thus improve the accuracy of transaction system detection, in some embodiments, such as Figure 3 As shown, the above S270 may specifically include:

[0140] S271. Generate an initial transaction probe message based on the target transaction data;

[0141] S272. Optimize the initial transaction probe message according to the target rules to obtain the transaction probe message.

[0142] Here, generating an initial transaction probe message based on the target transaction data can be achieved by transforming the target transaction data into an initial transaction probe message. This initial transaction probe message can include message fields corresponding to multiple transaction fields. If a message field corresponds to a transaction field, both can have the same meaning. Furthermore, each initial transaction probe message can include one transaction or multiple transactions; this is not limited here. Assuming each initial transaction probe message can include one transaction, for transaction fields and message fields with the same meaning, the content of the message field can correspond to the transaction data of one transaction within the transaction field.

[0143] In addition, the target rules may include at least one of the rules related to transaction specifications and the rules related to users' personalized needs.

[0144] As an example, a simulated trading system may also include a trading correction module. This module may include a field constraint knowledge base and a Large Language Model (LLM). The LLM can optimize the initial trading probe message based on the field constraint knowledge base and according to target rules to obtain the final trading probe message.

[0145] In addition to the above, other steps of the method in the embodiments of this application can be found in the above text. Figure 2 The relevant descriptions of the embodiments shown will not be repeated here.

[0146] Thus, by optimizing the initial transaction probe message according to the target rules, a transaction probe message is obtained that can meet at least one of the transaction specifications and the user's personalized needs, thereby improving the message quality of the transaction probe message and thus improving the accuracy of the transaction system detection.

[0147] Therefore, in order to ensure that the transaction probe message meets the transaction specifications and thus improves the message quality, in some embodiments, the initial transaction probe message is optimized according to the target rules to obtain the transaction probe message, which may specifically include:

[0148] Query the field constraint conditions corresponding to multiple message fields from the field constraint knowledge base. The field constraint knowledge base includes multiple field constraint conditions.

[0149] If the content of the first target message field does not meet the target field constraint conditions corresponding to the first target message field, the content of the first target message field is modified based on the target field constraint conditions to obtain a transaction probe message. The first target message field is any one of multiple message fields.

[0150] Here, the field constraint knowledge base can be a knowledge base built based on transaction specification-related knowledge. This knowledge base includes multiple field constraints. Field constraints can be used to constrain information such as what the content of a message field is, what conditions the field content must meet, and whether different message fields must appear simultaneously or not.

[0151] As an example, based on knowledge of transaction regulations, a Word document in the form of question-and-answer pairs can be generated first. The question-and-answer pairs in the Word document can be as follows:

[0152] Q: What are the specifications for the ID number-IDNO field?

[0153] A: The ID card number consists of 18 digits, the last digit can be X, and the ID number (IDNO) and ID type (IDTP) should appear simultaneously or not.

[0154] Here, Q can represent the question, and A can represent the answer. Additionally, IDNO and IDTP can be message fields.

[0155] Based on this, the above Word document is input into the vector embedding model (i.e., the Embedding model). The text representation in the Word document can be converted into a vector representation through the Embedding model, resulting in the vector representation of the question-answer pair. This leads to a vector database containing vector representations of multiple question-answer pairs, i.e., a field constraint knowledge base.

[0156] Furthermore, in the field constraint knowledge base, the aforementioned question-and-answer pairs can exist in tabular form. If the question-and-answer pairs exist in tabular form, the tabular data can be structurally transformed to obtain data in JSON-schema format, thereby improving the quality of knowledge retrieved from the field constraint knowledge base.

[0157] As an example, using Retrieval-Augmented Generation (RAG) technology, field constraints corresponding to multiple message fields can be queried from a field constraint knowledge base. These queried field constraints, along with the initial transaction probe message, are then input into a large language model. To enhance the model's understanding, prompt words can also be included to explain technical terms within the transaction probe message. Based on this, the large language model can validate the initial transaction probe message against these field constraints, identifying the first target message field that does not meet the constraints. Then, based on the target field constraints corresponding to this first target message field, the model corrects its content, resulting in the final transaction probe message.

[0158] This application embodiment first queries the field constraint conditions corresponding to multiple message fields from the field constraint knowledge base, and then, when the field content of the first target message field does not meet the target field constraint conditions corresponding to the first target message field, it corrects the field content of the first target message field based on the target field constraint conditions to obtain a transaction probe message, so that the transaction probe message can meet the transaction specifications, thereby improving the message quality of the transaction probe message.

[0159] Based on this, in order to further improve the message quality of the transaction probe message, in some embodiments, the initial transaction probe message is optimized according to the target rule to obtain the transaction probe message, which may specifically include:

[0160] To obtain personalized requirements for the second target message field, which can be any one of multiple message fields;

[0161] Based on personalized requirements, the content of the second target message field is rewritten to obtain the transaction probe message.

[0162] Here, personalized requirements can be integrated into the prompt word information. That is, the prompt word information can also include some personalized requirements. For example, a personalized requirement could be to instruct the large language model to change the first few digits of the card number in the transaction probe message to "621335". Another example is to instruct the large language model to change the content of the institution field from institution A to institution C, where institution C could be an institution that does not exist in the first transaction data but was added to the field constraint knowledge base.

[0163] As an example, after optimizing the initial transaction probe message to obtain the final transaction probe message, the comparison between the initial transaction probe message and the final transaction probe message can be shown in Table 1 below:

[0164] Table 1

[0165]

[0166] This application embodiment adds personalized requirements to the prompt word information, enabling the large language model to rewrite the initial transaction probe message in a personalized way based on external knowledge, generating richer field content than the current production situation. This makes the transaction probe message not only have the characteristics of approximating the real production transaction, but also have richer semantic information than the real production transaction, thereby further improving the message quality of the transaction probe message.

[0167] To better understand the above solutions, some specific examples are given based on the above embodiments.

[0168] For example, a schematic diagram of a method for generating a transaction probe message provided in an embodiment of this application can be shown as follows: Figure 4 As shown.

[0169] like Figure 4 As shown, the simulation trading system may include a transaction preprocessing module 41, a data splitting module 42, a transaction generation module 43, and a transaction correction module 44.

[0170] Since the data synthesizer generates data based on data having the same fields, the transaction preprocessing module 41 can be used to perform transaction preprocessing such as data filtering and format conversion on the raw transaction data generated by the transaction system, so as to obtain the first transaction data that can be directly input into the data synthesizer.

[0171] The data splitting module 42 can be used to determine the field data features corresponding to the multiple transaction fields in the first transaction data, and to split the first transaction data into multiple second transaction data based on the field data features corresponding to the multiple transaction fields. The field data features include field dependency features, a first distribution feature, and a second distribution feature.

[0172] The transaction generation module 43 may include a first Gaussian synthesizer, a second Gaussian synthesizer, and a CTGAN synthesizer. The first Gaussian synthesizer can generate first simulated transaction data based on second transaction data that satisfies field dependency characteristics. The second Gaussian synthesizer can generate second simulated transaction data based on second transaction data that satisfies a first distribution characteristic. The CTGAN synthesizer can generate third simulated transaction data based on second transaction data that satisfies a second distribution characteristic. Subsequently, the transaction generation module 43 can perform data fusion on the first, second, and third simulated transaction data to obtain target transaction data, and perform message transformation on the target transaction data to obtain an initial transaction probe message.

[0173] The transaction correction module 44 may include an embedding model, a field constraint knowledge base, and a large language model. The transaction correction module 44 can input prompt word information and the initial transaction probe message into the embedding model, enabling the embedding model to vectorize the prompt word information and the initial transaction probe message. Then, the transaction correction module 44 can query multiple structured and vectorized field constraints from the field constraint knowledge base, and input the vectorized field constraints, prompt word information, and initial transaction probe message into the large language model. This allows the large language model to determine the target rule based on the prompt word information and optimize the initial transaction probe message according to the target rule, thus obtaining the transaction probe message.

[0174] In summary, this application proposes a batch transaction generation scheme based on multiple synthesizers and large models.

[0175] First, to improve the realism of generated transactions, multiple different synthesizer networks were designed to address the characteristics of transactions in the financial sector. Leveraging the unique features of each synthesizer, transactions were generated in batches to suit the different characteristics of financial transactions. Specifically, a Gaussian synthesizer was designed to learn the distribution of different fields within the generated transactions, and a CTGAN synthesizer was designed to learn the correlation information between different fields within a single transaction.

[0176] Secondly, a data splitting method is proposed. The characteristic preferences of different transaction fields are calculated using KL divergence and Spearman correlation coefficient. Taking advantage of the strong statistical conformity of the Gaussian synthesizer, it can accurately capture the dependencies between variables. The CTGAN synthesizer, through adversarial training, can capture complex nonlinear distributions, making the generated data closer to the distribution of real data. Different fields are fed into different synthesizers for training, and then the results generated by each synthesizer are merged, thereby enhancing the quality of the generated transaction data.

[0177] Furthermore, a method for correcting generated transactions using a large model and knowledge base is proposed. A vector knowledge base is constructed to store transaction knowledge, and RAG technology is used to retrieve the corresponding knowledge. The transaction and knowledge are then fed into a large language model for verification and correction. The large language model can utilize external knowledge accessed within it to generate richer field content than the production transaction under the current conditions. This results in corrected transactions that not only approximate the characteristics of the production transaction but also possess richer semantic information. While the large language model can configure complex rules in the knowledge base to correct generated transactions, it cannot understand the distribution of transactions and the correlation between field content distributions under different dates, times, and transactions in actual production conditions based on these rules. Therefore, a method is designed to first generate data using a synthesizer and then correct the data using a large model, thereby improving the quality of transaction detection messages and the accuracy of transaction system detection.

[0178] In addition, the target transaction data generated by this application can also be used for data analysis and scenario modeling to explore user consumption habits under the premise of privacy protection, and to verify system compliance or simulate audit processes in compliance and audit scenarios to avoid exposing real business details.

[0179] Based on the transaction probe message generation method provided in the above embodiments, this application also provides specific implementations of the transaction probe message generation apparatus. Please refer to the following embodiments.

[0180] like Figure 5 As shown, the transaction probe message generation apparatus 500 provided in this application embodiment includes the following modules:

[0181] Module 510 is used to acquire the first transaction data of the actual transaction;

[0182] The determination module 520 is used to determine the field data features corresponding to the multiple transaction fields based on the transaction data corresponding to the multiple transaction fields in the first transaction data.

[0183] The splitting module 530 is used to split the first transaction data based on the field data characteristics corresponding to multiple transaction fields to obtain multiple second transaction data.

[0184] The determination module 520 is also used to determine, for each second transaction data, the data synthesizer corresponding to the second transaction data based on the correspondence between the field data features of the transaction field corresponding to the second transaction data and the data synthesizer;

[0185] The generation module 540 is used to generate simulated transaction data corresponding to each second transaction data by using a data synthesizer corresponding to the second transaction data.

[0186] The fusion module 550 is used to fuse the simulated transaction data corresponding to multiple second transaction data to obtain the target transaction data;

[0187] The generation module 540 is also used to generate transaction probe messages based on the target transaction data.

[0188] The transaction probe message generation device 500 described above will be explained in detail below:

[0189] In some embodiments, the determining module 520 may specifically include:

[0190] The calculation submodule is used to calculate the dependency strength between any two transaction fields based on the transaction data corresponding to the two transaction fields in the first transaction data.

[0191] The determination submodule is used to determine the data distribution corresponding to any transaction field based on the transaction data corresponding to the transaction field in the first transaction data for any given transaction field.

[0192] The determination submodule is also used to determine the field data characteristics corresponding to multiple transaction fields based on dependency strength and data distribution.

[0193] In some embodiments, determining a submodule may specifically include:

[0194] Determine the unit, used for:

[0195] If the dependency strength between two transaction fields meets the dependency condition, the field data features corresponding to the two transaction fields are determined as field dependency features;

[0196] The remaining transaction fields that do not satisfy the field dependency feature among multiple transaction fields are identified as the first type of transaction fields;

[0197] If the difference between the data distribution corresponding to the first type of transaction field and the reference distribution meets the difference condition, the field data feature corresponding to the first type of transaction field is determined as the first distribution feature;

[0198] If the difference between the data distribution corresponding to the first type of transaction field and the reference distribution does not meet the difference condition, the field data feature corresponding to the first type of transaction field is determined as the second distribution feature.

[0199] In some embodiments, the field data feature of the transaction field corresponding to the second transaction data is a field dependency feature, the data synthesizer corresponding to the second transaction data is a first Gaussian synthesizer, and the simulated transaction data includes the first simulated transaction data.

[0200] Based on this, the generation module 540 may specifically include:

[0201] The calculation submodule is also used to calculate the first mean and covariance corresponding to the second transaction data through the first Gaussian synthesizer;

[0202] The generation submodule is used to generate the first simulated transaction data based on the first mean and covariance.

[0203] In some embodiments, the field data feature of the transaction field corresponding to the second transaction data is a first distribution feature, the data synthesizer corresponding to the second transaction data is a second Gaussian synthesizer, and the simulated transaction data includes the second simulated transaction data.

[0204] Based on this, the generation module 540 may specifically include:

[0205] The calculation submodule is also used to calculate the second mean and standard deviation corresponding to the second transaction data through the second Gaussian synthesizer;

[0206] The generation submodule is also used to generate second simulated trading data based on the second mean and standard deviation.

[0207] In some embodiments, the field data feature of the transaction field corresponding to the second transaction data is a second distribution feature, the data synthesizer corresponding to the second transaction data is a conditional table generative adversarial network synthesizer, and the simulated transaction data includes third simulated transaction data.

[0208] Based on this, the generation module 540 may specifically include:

[0209] The training submodule is used to train the conditional table generative adversarial network synthesizer based on the second transaction data, so as to obtain the trained target conditional table generative adversarial network synthesizer.

[0210] The acquisition submodule is used to acquire random noise data;

[0211] The generation submodule is also used to generate an adversarial network synthesizer based on a target condition table and to generate third-dimensional simulated transaction data based on random noise data.

[0212] In some embodiments, the generation module 540 may specifically include:

[0213] The generation submodule is also used to generate an initial transaction probe message based on the target transaction data;

[0214] The optimization submodule is used to optimize the initial transaction probe message according to the target rules to obtain the transaction probe message.

[0215] In some embodiments, the initial transaction probe message includes message fields that correspond to multiple transaction fields respectively.

[0216] Based on this, the optimization sub-modules may specifically include:

[0217] The query unit is used to query the field constraint conditions corresponding to multiple message fields from the field constraint knowledge base, which includes multiple field constraint conditions.

[0218] The correction unit is used to correct the content of the first target message field based on the target field constraint conditions when the content of the first target message field does not meet the target field constraint conditions corresponding to the first target message field, so as to obtain a transaction probe message. The first target message field is any one of multiple message fields.

[0219] In some embodiments, the initial transaction probe message includes message fields that correspond to multiple transaction fields respectively.

[0220] Based on this, the optimization sub-modules may specifically include:

[0221] The acquisition unit is used to acquire the personalized requirements of the second target message field, which is any one of multiple message fields;

[0222] The rewriting unit is used to rewrite the field content of the second target message based on personalized requirements to obtain the transaction probe message.

[0223] This application embodiment first determines the field data features corresponding to multiple transaction fields based on the transaction data corresponding to multiple transaction fields in the first transaction data. Then, based on the field data features corresponding to the multiple transaction fields, the first transaction data is split to obtain multiple second transaction data that match the field data features of the transaction fields. Next, for each second transaction data, based on the correspondence between the field data features of the corresponding transaction fields and the data synthesizer, a data synthesizer corresponding to the second transaction data is determined, resulting in a data synthesizer that matches both the field data features and the second transaction features. Based on this, simulated transaction data corresponding to the second transaction data is generated using the data synthesizer corresponding to the second transaction data. That is, by using the data synthesizer that matches both the field data features and the second transaction data, data synthesis is performed on the second transaction data that matches the field data features to obtain simulated transaction data, thereby improving the similarity between the simulated transaction data and the second transaction data. Thus, by fusing the simulated transaction data corresponding to multiple second transaction data, simulated target transaction data is obtained, which improves the similarity between the simulated target transaction data and the first transaction data of the real transaction. Thus, by generating transaction probe messages based on target transaction data, the similarity between the simulated transaction scenario reflected in the transaction probe message and the actual transaction scenario can be improved, thereby improving the accuracy of detecting the transaction system.

[0224] Based on the transaction probe message generation method provided in the above embodiments, this application also provides specific implementation methods for electronic devices. Figure 6 A schematic diagram of an electronic device provided in an embodiment of this application is shown.

[0225] like Figure 6 As shown, the electronic device 600 may include a processor 610 and a memory 620 storing computer program instructions.

[0226] Specifically, the processor 610 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0227] Memory 620 may include mass storage for data or instructions. For example, and not limitingly, memory 620 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where suitable, memory 620 may include removable or non-removable (or fixed) media. Where suitable, memory 620 may be internal or external to electronic device 600. In a particular embodiment, memory 620 is a non-volatile solid-state memory.

[0228] In a specific embodiment, the memory 620 can be implemented as ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 620 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 620 and executed by the processor 610. The processor 610 reads and executes the computer program instructions stored in the memory 620 to implement any of the transaction probe message generation methods in the above embodiments.

[0229] The processor 610 reads and executes computer program instructions stored in the memory 620 to implement any of the transaction probe message generation methods in the above embodiments.

[0230] In one example, electronic device 600 may further include communication interface 630 and bus 640. For example, Figure 6 As shown, the processor 610, memory 620, and communication interface 630 are connected via bus 640 and communicate with each other.

[0231] The communication interface 630 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0232] Bus 640 includes hardware, software, or both, that couples components of an electronic device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 640 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.

[0233] For example, the electronic device 600 can be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc.

[0234] The electronic device can execute the transaction probe message generation method in the embodiments of this application, thereby achieving the combination of Figures 1 to 4 The method for generating transaction probe messages is described, and the beneficial effects of the corresponding method embodiments are not elaborated here.

[0235] Furthermore, in conjunction with the transaction probe message generation method in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement any of the transaction probe message generation methods in the above embodiments.

[0236] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0237] The computer program instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the transaction probe message generation method as shown in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0238] Based on the transaction probe message generation method in the above embodiments, this application embodiment can provide a computer program product for implementation. When the instructions in this computer program product are executed by the processor of an electronic device, they implement any of the transaction probe message generation methods in the above embodiments.

[0239] The computer program products of the above embodiments are used to implement the transaction probe message generation method shown in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0240] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0241] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0242] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0243] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0244] The above description is only a specific embodiment of the present application. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the scope of protection of the present application.

Claims

1. A method for generating a transaction probe message, characterized in that, include: Obtain first-hand transaction data of real transactions; Based on the transaction data corresponding to multiple transaction fields in the first transaction data, determine the field data features corresponding to the multiple transaction fields respectively; Based on the field data characteristics corresponding to the multiple transaction fields, the first transaction data is split to obtain multiple second transaction data; For each of the second transaction data, a data synthesizer corresponding to the second transaction data is determined based on the correspondence between the field data features of the transaction field corresponding to the second transaction data and the data synthesizer; For each second transaction data, simulated transaction data corresponding to the second transaction data is generated through a data synthesizer corresponding to the second transaction data; The simulated transaction data corresponding to the multiple second transaction data are fused to obtain the target transaction data; Based on the target transaction data, a transaction probe message is generated.

2. The method according to claim 1, characterized in that, The step of determining the field data features corresponding to each of the multiple transaction fields in the first transaction data includes: For any two transaction fields, the dependency strength between the two transaction fields is calculated based on the transaction data corresponding to the two transaction fields in the first transaction data. For any of the transaction fields, the data distribution corresponding to the transaction field is determined based on the transaction data corresponding to the transaction field in the first transaction data; Based on the dependency strength and the data distribution, the field data characteristics corresponding to the multiple transaction fields are determined respectively.

3. The method according to claim 2, characterized in that, The step of determining the field data features corresponding to the multiple transaction fields based on the dependency strength and the data distribution includes: If the dependency strength between two transaction fields meets the dependency condition, the field data features corresponding to the two transaction fields are determined as field dependency features; The remaining transaction fields that do not satisfy the field dependency characteristics among the plurality of transaction fields are identified as the first type of transaction fields; If the difference between the data distribution corresponding to the first type of transaction field and the reference distribution meets the difference condition, the field data feature corresponding to the first type of transaction field is determined as the first distribution feature. If the difference between the data distribution corresponding to the first type of transaction field and the reference distribution does not meet the difference condition, the field data feature corresponding to the first type of transaction field is determined as the second distribution feature.

4. The method according to claim 3, characterized in that, The field data feature of the transaction field corresponding to the second transaction data is the field dependency feature, the data synthesizer corresponding to the second transaction data is a first Gaussian synthesizer, the simulated transaction data includes first simulated transaction data, and the step of generating simulated transaction data corresponding to the second transaction data through the data synthesizer corresponding to the second transaction data includes: The first mean and covariance of the second transaction data are calculated using the first Gaussian synthesizer. The first simulated transaction data is generated based on the first mean and the covariance.

5. The method according to claim 3, characterized in that, The field data feature of the transaction field corresponding to the second transaction data is the first distribution feature, the data synthesizer corresponding to the second transaction data is a second Gaussian synthesizer, the simulated transaction data includes the second simulated transaction data, and the step of generating simulated transaction data corresponding to the second transaction data through the data synthesizer corresponding to the second transaction data includes: The second mean and standard deviation of the second transaction data are calculated using the second Gaussian synthesizer. The second simulated transaction data is generated based on the second mean and the standard deviation.

6. The method according to claim 3, characterized in that, The field data feature of the transaction field corresponding to the second transaction data is the second distribution feature. The data synthesizer corresponding to the second transaction data is a conditional table generative adversarial network synthesizer. The simulated transaction data includes third simulated transaction data. Generating simulated transaction data corresponding to the second transaction data through the data synthesizer corresponding to the second transaction data includes: The conditional table generative adversarial network synthesizer is trained based on the second transaction data to obtain a trained target conditional table generative adversarial network synthesizer. Obtain random noise data; The adversarial network synthesizer generates the third simulated transaction data based on the random noise data using the target condition table.

7. The method according to any one of claims 1-6, characterized in that, The step of generating a transaction probe message based on the target transaction data includes: Based on the target transaction data, an initial transaction probe message is generated; The initial transaction probe message is optimized according to the target rules to obtain the transaction probe message.

8. The method according to claim 7, characterized in that, The initial transaction probe message includes message fields corresponding to the plurality of transaction fields respectively. The optimization of the initial transaction probe message according to the target rule to obtain the transaction probe message includes: Query the field constraint conditions corresponding to the multiple message fields from the field constraint knowledge base, wherein the field constraint knowledge base includes the multiple field constraint conditions; If the content of the first target message field does not meet the target field constraint conditions corresponding to the first target message field, the content of the first target message field is modified based on the target field constraint conditions to obtain the transaction probe message, wherein the first target message field is any one of the plurality of message fields.

9. The method according to claim 7, characterized in that, The initial transaction probe message includes message fields corresponding to the plurality of transaction fields respectively. The optimization of the initial transaction probe message according to the target rule to obtain the transaction probe message includes: The personalized requirements for obtaining the second target message field are obtained, where the second target message field is any one of the plurality of message fields; Based on the personalized requirements, the field content of the second target message field is rewritten to obtain the transaction detection message.

10. A device for generating transaction detection messages, characterized in that, The device includes: The acquisition module is used to acquire the first transaction data of the actual transaction. The determination module is used to determine the field data features corresponding to the multiple transaction fields based on the transaction data corresponding to the multiple transaction fields in the first transaction data. The splitting module is used to split the first transaction data based on the field data features corresponding to the multiple transaction fields to obtain multiple second transaction data. The determining module is further configured to, for each second transaction data, determine the data synthesizer corresponding to the second transaction data based on the correspondence between the field data features of the transaction field corresponding to the second transaction data and the data synthesizer; The generation module is used to generate simulated transaction data corresponding to each second transaction data by using a data synthesizer corresponding to the second transaction data; The fusion module is used to fuse the simulated transaction data corresponding to the multiple second transaction data respectively to obtain the target transaction data; The generation module is also used to generate a transaction probe message based on the target transaction data.

11. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the method for generating transaction probe messages as described in any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the method for generating transaction probe messages as described in any one of claims 1-9.

13. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device performs the transaction probe message generation method as described in any one of claims 1-9.