Bill data generation method and device
By combining real-time and offline data streams, and using Flink and Paimon to generate billing data tables, the needs of merchants for personalized billing time ranges are addressed, achieving real-time and flexible billing data and improving user experience.
Patent Information
- Application Number
- CN202511005886.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-07
AI Technical Summary
Existing offline methods for generating billing data cannot meet merchants' personalized needs for billing time ranges, and traditional streaming computing engines suffer from state bloat and uncontrollable latency issues.
By combining real-time and offline data streams, and using the streaming computing engine Flink and the streaming data warehouse Paimon, a target wide table is generated. Data records are then filtered according to the merchant's customized billing time range to generate a billing data table.
It achieves real-time and flexible generation of billing data based on the merchant's customized billing time range, improves user experience, and solves the problems of high latency and poor flexibility of traditional solutions.
Smart Images

Figure CN120912281A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] One or more embodiments of the present specification relate to the field of computer information processing, and in particular, to a bill data generation method and device. BACKGROUND
[0002] At present, the bill data is mainly generated through an offline manner, and the bill data billing time range is usually fixed, for example, 0:00-23:59 every day. In practice, there may be some merchants who need to generate bill data across days (for example, 10:00 of the previous day to 9:59 of the current day) or across hours (10:10-11:10 of the current day). Obviously, this offline bill data generation method cannot meet the personalized needs of merchants for billing time range.
[0003] Therefore, it is urgent to provide a more effective bill data generation scheme to meet the personalized needs of merchants for billing time range. SUMMARY
[0004] One or more embodiments of the present specification describe a bill data generation method and device, which can generate bill data according to the billing time range defined by the merchant, which can greatly improve the user experience.
[0005] In a first aspect, a bill data generation method is provided, which is executed through an offline computing platform, comprising:
[0006] reading a target wide table from a streaming data warehouse, wherein the target wide table includes a primary key column and a plurality of column clusters corresponding to a plurality of online single tables; a single column cluster synchronizes a first data record newly added in the corresponding online single table in real time by updating the target wide table in part; the plurality of online single tables includes a transaction detail table for recording transaction data, and the transaction detail table at least includes a transaction time field and a merchant field;
[0007] determining each second data record from the target wide table according to a target filtering condition, and generating a bill data table based on the each second data record; the target filtering condition includes that the field value of the transaction time field is within the billing time range defined by the target merchant, and the field value of the merchant field is the target merchant.
[0008] In a second aspect, a bill data generation device is provided, which is arranged in an offline computing platform, comprising:
[0009] The read unit is configured to read a target wide table from the stream warehouse, wherein the target wide table comprises a primary key column and a plurality of column clusters corresponding to a plurality of online single tables; a single column cluster synchronizes a first data record newly added in a corresponding online single table in real time by partially updating the target wide table; the plurality of online single tables comprises a transaction detail table used for recording transaction data, and the transaction detail table at least comprises a transaction time field and a merchant field;
[0010] The generation unit is configured to determine each second data record from the target wide table according to a target filtering condition, and generate a bill data table based on the each second data record; the target filtering condition comprises that a field value of the transaction time field is within a bill issuing time range defined by a target merchant, and a field value of the merchant field is the target merchant.
[0011] In a third aspect, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. When the computer program is executed in a computer, the computer program causes the computer to execute the method in the first aspect.
[0012] In a fourth aspect, a computing device is provided, and the computing device comprises a memory and a processor. The memory stores executable code, and the processor executes the executable code to implement the method in the first aspect.
[0013] The bill data generation method provided by one or more embodiments of the present specification generates a bill data table of a merchant by combining a real-time link and an offline link. The real-time link is used to generate a target wide table, and the target wide table stores data records newly added in a plurality of online single tables. The offline link is used to filter each data record related to a merchant from the target wide table according to a bill issuing time range defined by the merchant, and generate a bill data table of the merchant based on the each data record. It should be noted that the real-time link can be used to collect data in the online single table in real time, so as to meet the real-time requirement of data when generating the bill data according to the bill issuing time range defined by the merchant. In addition, the bill data of the merchant is generated according to the bill issuing time range defined by the merchant, which can greatly improve the flexibility of the present solution and improve the user experience. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present specification, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present specification, and other drawings can be obtained by those skilled in the art without creating laborious work.
[0015] Figure 1 The embodiment of the present specification is an embodiment of the embodiment of the present specification.
[0016] Figure 2 An updating method of a target wide table is shown in one example of the present specification.
[0017] Figure 3 A flow chart of a bill data generation method according to one embodiment of the present specification is shown.
[0018] Figure 4 An updating method of a target wide table is shown in one example of the present specification.
[0019] Figure 5 An updating method of a target wide table is shown in one example of the present specification. DETAILED DESCRIPTION
[0020] The scheme provided in the present specification will be described below in combination with the accompanying drawings.
[0021] As mentioned above, the bill data is currently mainly generated in an offline manner, however, the offline generation manner usually relies on offline data tables which are updated periodically (such as daily or hourly), and these offline data tables have high delay, thus cannot support the generation of bill data according to the custom billing time range of a merchant.
[0022] In order to meet the real-time requirement of data, part of the schemes attempt to use a stream computing engine Flink to generate bill data, however, this kind of manner needs to save a large amount of data in the Flink state, and is prone to problems such as state expansion and IO bottleneck. In addition, in the case that the saved data scale is large, the delay is also difficult to control.
[0023] Therefore, the present scheme proposes to generate the bill data table of a merchant by combining the use of real-time link and offline link. The real-time link is used to generate a target wide table, in which a plurality of stream newly added data records in online single tables are associated and stored. The offline link is used to filter out each data record related to a certain merchant from the target wide table according to the custom billing time range of the merchant, and generate the bill data table of the merchant based on it. It should be noted that, since the data in the online single table can be collected in real time through the real-time link, the real-time requirement of data when generating the bill data according to the custom billing time range of the merchant can be met. In addition, the bill data of the merchant is generated according to the custom billing time range of the merchant in the present scheme, which can greatly improve the flexibility of the present scheme, and can improve the user experience.
[0024] Figure 1 An implementation scenario diagram of one embodiment disclosed in the present specification. Figure 1In the target wide table, the data records newly added in the online single tables in real time can be stored in association through a real-time link. Through an offline link, the data records related to a certain merchant can be filtered from the target wide table according to a custom billing time range of the merchant, and the billing data table of the merchant can be generated based on the data records.
[0025] It should be understood that, Figure 1 It should be understood that,
[0026] The target wide table will be described in detail below.
[0027] The target wide table described in the specification can include a primary key column and a plurality of column clusters corresponding to a plurality of online single tables. In one example, the target wide table can be as shown in Table 1.
[0028] Table 1
[0029] Primary key column Column cluster 1 Column cluster 2 Column cluster 3 … PK {f11,f12,…} {f21,f22,…} {f31,f32,…} …
[0030] Among them, {f11, f12, …} is an array corresponding to a data record in the online single table corresponding to column cluster 1, and each element value in the array is a field value in the corresponding data record. Similarly, {f21, f22, …} is an array corresponding to a data record in the online single table corresponding to column cluster 2, and {f31, f32, …} is an array corresponding to a data record in the online single table corresponding to column cluster 3. It should be understood that the three data records all contain the primary key value PK.
[0031] It should be noted that each data record in Table 1 can be recorded in the corresponding column cluster in the target wide table through partial update of the target wide table, which will be described in detail below.
[0032] Figure 2 The updating method of the target wide table in one example of the specification is shown in the schematic diagram. Figure 2 In the target intermediate, when a new data record is added in a certain online single table, a corresponding data change log is generated. The target middleware (such as SLS) monitors the generation of the data change log, analyzes the data change log, thereby obtaining the newly added data record, and records it in the intermediate table corresponding to the online single table. In one example, the intermediate table can also be stored in the streaming data warehouse Paimon, so that the intermediate table can be regarded as a kind of Paimon table.
[0033] When it is monitored that there is a new data record in the intermediate table, the stream computing engine Flink can read the data record from the intermediate table, and record the data record to the corresponding column cluster in the target wide table according to the primary key value of the data wide table contained in the data record. Thus, the real-time synchronization recording of the data record in the single online single table is realized.
[0034] The target wide table can also be stored in the stream data warehouse Paimon, that is, the target wide table can also be regarded as a Paimon table.
[0035] It should be noted that the stream data warehouse Paimon supports high-throughput, low-latency data collection, stream subscription and real-time query. In addition, the stream computing engine Flink supports real-time data analysis and processing. Therefore, the target wide table is generated based on Flink and Paimon in the present scheme, which can solve the problems of state expansion and high latency. In addition, since Flink supports delay monitoring and abnormal alarm, and Paimon supports back flushing and reset point, the present scheme can also realize abnormal processing and monitoring. In summary, by using Flink and Paimon together, the present scheme can improve the generation efficiency and stability of the target wide table.
[0036] Of course, in practice, the target wide table can also be created and updated by other real-time computing engines (such as Spark, etc.), and can also be stored in other stream data warehouses (such as Hudi, etc.).
[0037] Since the data record is only added under one column cluster in the target wide table, it can be regarded as a partial update of the target wide table.
[0038] Additionally, before the stream computing engine Flink records the data record under the corresponding column cluster in the target wide table, the field values in the data record can be preprocessed as follows: standardization processing (such as converting data of different scales or dimensions into uniformly standardized data), character conversion (such as converting English characters into Chinese characters), etc.
[0039] It should be noted that the present scheme supports storing multiple online single tables with different data relationships in the target wide table. For example, the multiple online single tables can include a transaction detail table for recording transaction data, a payment detail table describing payment information of the transaction, and a settlement detail table obtained by aggregating transaction data according to settlement dates. Among them, the transaction detail table and the payment detail table can be in a 1:1 data relationship, and the transaction detail table and the settlement detail table and the payment detail table and the settlement detail table are in a N:1 data relationship.
[0040] In one example, the transaction detail table can be as shown in Table 2.
[0041] Table 2
[0042]
[0043] In Table 2, the trade_no field can be a primary key field.
[0044] In practice, the above transaction detail table can include fewer fields, for example, it can only include the trade_no field, the transaction time field and the merchant id field. It can also include more fields, for example, it can also include the out_order_no field and the status field, etc., which are not limited in the present specification.
[0045] The above payment detail table can be as shown in Table 3.
[0046] Table 3
[0047]
[0048] In Table 3, the payment_id field can be a primary key field.
[0049] In practice, the above payment detail table can also include more fields, for example, it can also include the fund_status field and the asset_type field, etc.
[0050] The above settlement detail table can be as shown in Table 4.
[0051] Table 4
[0052]
[0053] In Table 4, the settle_bill_id field can be a primary key field.
[0054] Since the settlement detail table is obtained by settling multiple transaction data, the field value of the above biz_no field can be multiple trade_no corresponding to multiple transaction data (i.e. multiple transaction records).
[0055] Wherein, in the case that the above several online single tables include the transaction detail table, the payment detail table and the settlement detail table, the target wide table constructed can be as shown in Table 5.
[0056] Table 5
[0057]
[0058] That is, the primary key column of the target wide table is the order number, and in the first row (or the first data record) of the target wide table, the value of column cluster 1 is array 1, each element value of which is the field value of each field in the data record identified by trade_nox in the transaction details table, except the order number field; the value of column cluster 2 is array 2, each element value of which is the field value of each field in the data record identified by trade_nox in the payment details table, except the order number field; the value of column cluster 3 is array 3, each element value of which is the field value of each field in the data record identified by trade_nox in the settlement details table, except the order number field.
[0059] It should be understood that since the order number field is not the primary key field of the payment details table and the settlement details table, there may be multiple data records with the same order number in them, in which case the data records can be aggregated according to the order number field (i.e., taking the order number as a grouping field). In a specific example, the aggregation method for each field in any online single table can be configured in advance, and then when multiple data records with the same order number in the online single table are aggregated, the field values of each field can be aggregated according to the aggregation method configured in advance for each field, thereby obtaining an aggregated record.
[0060] It should also be noted that since the data records in the online single table are streamingly growing, and these streamingly growing data records are synchronized to the target wide table in real time, if the data in the target wide table is not reduced, it will cause the problem of excessive data volume in the target wide table.
[0061] In an embodiment, the data in the target wide table can be reduced according to the transaction time, for example, periodically (e.g., daily) deleting data records corresponding to transaction times exceeding a target time range (e.g., 48 hours) from the target wide table.
[0062] In addition, the data in the target wide table can also be reduced according to the merchant field, for example, only retaining data records related to a specified merchant in the target wide table.
[0063] In a more specific embodiment, the number of specified merchants is multiple, and their records are recorded into a target data table, at which time a lookup join (Lookup Join) can be established between the target wide table and the target data table, and then data records related to the specified merchants are filtered out. Lookup Join is a core technology for real-time dimension association in stream data warehouse, and reasonable use can balance between low latency and high consistency.
[0064] In summary, the present scheme stores multiple online single tables in association in the target wide table, which can improve the flexibility of data processing.
[0065] In the case where the several online single tables include the transaction detail table, the bill data of a certain merchant within the self-defined payment time range of the merchant can be generated based on the constructed target wide table, and the generation process is described below.
[0066] Figure 3 A flowchart of a bill data generation method according to one embodiment of the present specification is shown, and the execution subject of the method can be an offline computing platform. As shown in the figure, the method can include the following steps: Figure 3
[0067] Step S302, reading the target wide table from the streaming data warehouse.
[0068] Here, the streaming data warehouse can be Paimon, Hudi, etc.
[0069] Taking Paimon as an example, it can include a data source layer ODS for storing each data record r read from the intermediate table, a detail data layer DWD for storing each data record r after data transcoding, a summary data layer DWS for integrating each data record r after transcoding into a wide table, and a data application layer ADS for sending the wide table to the offline computing platform.
[0070] The target wide table can be obtained by the method shown in the figure, that is, it includes the primary key column and several column clusters corresponding to several online single tables. Each column cluster synchronizes the newly added data records r in the corresponding online single table in real time by partially updating the target wide table. Figure 3
[0071] In addition, it can also include a payment detail table and a settlement detail table, etc.
[0072] Step S304, determining each data record R from the target wide table according to the target filtering condition, and generating a bill data table based on the read each data record R.
[0073] Additionally, the target wide table can be verified for correctness, which specifically includes: obtaining a plurality of offline single tables corresponding to a plurality of online single tables, each offline single table is periodically (such as daily or hourly) updated, and each field included therein is the same as the corresponding online single table. Based on the plurality of offline single tables, the correctness of the target wide table is verified.
[0074] Taking the transaction detail table shown in Table 2 as an example, the offline transaction detail table corresponding thereto also includes 8 fields such as trade_no, except that it is periodically updated. Taking the offline transaction detail table updated by the hour as an example, assuming that the current time is 9:42, the offline transaction detail table can record transaction data of each hour before the current time, i.e., 9 o'clock, 8 o'clock, 7 o'clock, and the like, but does not record transaction data after 9 o'clock. The online transaction detail table records each transaction data up to the current time, including transaction data after 9 o'clock.
[0075] Taking the correctness verification of the target wide table based on the offline transaction detail table as an example, each offline data record r' whose field value of the transaction time field is within the target time range can be read from the offline transaction detail table. Here, the target time range can refer to the time range (e.g., 48 hours) to which the transaction time in the target wide table belongs. Each offline data record r' is compared with each data record r under the target column cluster corresponding to the offline transaction detail table (i.e., the column cluster corresponding to the online transaction detail table) one by one, to obtain a comparison result.
[0076] In a more specific embodiment, each offline data record r' can be divided into multiple partitions according to the update time of the offline transaction detail table, and then compared by partition. For example, each offline data record r' is divided into multiple partitions according to the hour, and then each data record r' in each hourly partition is compared with each data record r under the target column cluster one by one.
[0077] The comparison of each offline data record r' with each data record r under the target column cluster one by one described above can be understood as a flow check. In practice, in addition to the flow check, the following correctness verification can also be performed: unified payment check (i.e., comparing the amount of each transaction data with the total unified payment amount), number of transactions check (i.e., determining whether the settlement batch ID in the settlement data can be associated with all the settlement batches of the transaction data), amount check (i.e., determining whether the amount in a batch is equal to the batch settlement amount or the transaction amount actually received is equal to the batch settlement amount), and the like.
[0078] After the correctness verification of the target wide table based on each offline single table, multiple comparison results corresponding to multiple offline single tables can be obtained, and finally, the multiple comparison results can be integrated to obtain a final correctness verification result.
[0079] Of course, in practice, the correctness verification of the target wide table can also be performed according to expert experience or by means of random sampling, which is not limited in the present specification.
[0080] It should be understood that in the case of correctness checking of the target wide table, each data record R can be determined from the target wide table according to the target filtering condition after the correctness checking is passed.
[0081] The target filtering condition described above at least includes that the field value of the transaction time field is within the target merchant self-defined posting time range, and the field value of the merchant field is the target merchant. That is, based on the target filtering condition, each data record R related to the target merchant and having a transaction time within the merchant self-defined posting time range can be filtered.
[0082] In addition, the target filtering condition described above can also include a plurality of target fields configured for the target merchant.
[0083] It should be understood that in the case where the target filtering condition only limits the transaction time field and the merchant field, the bill data table can be directly generated based on the filtered data records R.
[0084] In the case where the target filtering condition also limits the target field described above, for any data record R, each target field value corresponding to the plurality of target fields can be extracted from each column cluster contained therein, and a target data record R' can be formed based on each target field value. Finally, based on each target data record R' corresponding to each data record R, the bill data table is generated.
[0085] In one example, the bill data table can be shown in Table 6.
[0086] Table 6
[0087]
[0088] Figure 4 The bill data generation method shown in one example of the present specification is shown in the schematic diagram. Figure 4 In the method, the offline computing platform first reads the target wide table from the streaming data warehouse, wherein the target wide table is associated with the data records newly added in the plurality of online single tables. Then, the plurality of offline single tables corresponding to the plurality of online single tables are obtained, and the target wide table is subjected to multiple verifications such as payment checking, flow checking and amount checking based on the plurality of offline single tables. After the multiple verifications are passed, each data record R related to the target merchant and having a transaction time within the merchant self-defined posting time range can be filtered from the target wide table. In addition, field filtering can be performed on each data record R, thereby obtaining each final data record R', and a bill data table of the merchant is generated based thereon.
[0089] In summary, the scheme can solve the problems of high delay, difficult operation and maintenance, and poor flexibility of the traditional offline scheme by combining Flink and Paimon, and overcome the defects of state expansion and uncontrollable delay of the pure Flink scheme. In addition, the partial update mechanism and wide table merging capability of Paimon make the multi-table data fusion more efficient. In summary, the scheme generally supports flexible billing time range configuration, easy operation and maintenance, and exception handling, greatly improving the real-time performance, accuracy and maintainability of billing data.
[0090] Corresponding to the above billing data generation method, an embodiment of the present specification also provides a billing data generation device arranged in an offline computing platform. As shown in Figure 5 The device can include:
[0091] The reading unit 502 is configured to read a target wide table from the streaming data warehouse, wherein the target wide table includes a primary key column and a plurality of column clusters corresponding to a plurality of online single tables, and each column cluster synchronizes a first data record newly added in the corresponding online single table in real time by partially updating the target wide table. The plurality of online single tables includes a transaction detail table for recording transaction data, and the transaction detail table at least includes a transaction time field and a merchant field.
[0092] The generating unit 504 is configured to determine each second data record from the target wide table according to a target filtering condition, and generate a billing data table based on each second data record. The target filtering condition includes that the field value of the transaction time field is within the billing time range customized by the target merchant, and the field value of the merchant field is the target merchant.
[0093] In one embodiment, the target filtering condition further includes a plurality of target fields configured for the target merchant.
[0094] The generating unit 504 includes:
[0095] The extraction submodule 5042 is configured to extract, for any second data record, each target field value corresponding to each target field from each column cluster included in the second data record, and form a target data record based on each target field value.
[0096] The generating submodule 5044 is configured to generate a billing data table based on each target data record corresponding to each second data record.
[0097] In one embodiment, the device further includes:
[0098] The obtaining unit 506 is configured to obtain a plurality of offline single tables corresponding to a plurality of online single tables, and each offline single table is periodically updated and includes the same fields as the corresponding online single table.
[0099] The checking unit 508 is configured to check the correctness of the target wide table based on the plurality of offline single tables.
[0100] The generating unit 504 is specifically configured to:
[0101] After the correctness check passes, each second data record is determined from the target wide table.
[0102] In one embodiment, only the data records that are newly added in the target time range in the corresponding online single table are recorded under a single column cluster, and the target time range includes the above-mentioned posting time range.
[0103] The checking unit 508 is specifically configured to:
[0104] For the target offline single table corresponding to the transaction detail table, each offline data record whose field value of the transaction time field is within the target time range is read from the target offline single table;
[0105] Each offline data record is compared with each first data record under the target column cluster corresponding to the target offline single table, and a first comparison result is obtained.
[0106] According to the comparison results corresponding to the plurality of offline single tables, a checking result of the correctness check is determined.
[0107] In one embodiment, the target wide table is created by the stream computing engine.
[0108] In one embodiment, the first data record newly added in the single online single table is read by the stream computing engine from the corresponding intermediate table, and each data record in the intermediate table is obtained by parsing the data change log of the online single table by the target middleware.
[0109] In one embodiment, the stream data warehouse is Paimon, and the stream computing engine is Flink.
[0110] In one embodiment, the plurality of online single tables further include one or more of the following:
[0111] A payment detail table describing payment information of the transaction;
[0112] A settlement detail table obtained by aggregating the transaction data according to the settlement date.
[0113] The functions of each functional unit of the above-mentioned embodiment device can be realized by each step of the above-mentioned method embodiment, and therefore, the specific working process of the device provided by the embodiment of the present specification will not be repeated here.
[0114] The bill data generation apparatus provided by one embodiment of the present specification can generate bill data according to the account time range defined by the merchant, which can greatly improve the user experience.
[0115] According to another aspect, an embodiment also provides a computer readable storage medium having stored thereon a computer program which, when executed in a computer, causes the computer to perform the method described in conjunction with Figure 4 the described method.
[0116] According to another aspect, an embodiment also provides a computer readable storage medium having stored thereon a computer program which, when executed in a computer, causes the computer to perform the method described in conjunction with Figure 4 the described method.
[0117] Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, the medium or device embodiments are described simply because they are basically similar to the method embodiments, and the relevant parts can be referred to the part of the method embodiments.
[0118] The above describes specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.
[0119] The above detailed description of the specific implementation further describes the purpose, technical solutions and beneficial effects of the present specification. It should be understood that the above description is only for the specific implementation of the present specification and is not used to limit the protection scope of the present specification. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solutions of the present specification shall be included in the protection scope of the present specification.
Claims
1. A method for generating billing data, executed by an offline computing platform, comprising: reading a target wide table from a streaming data warehouse, wherein the target wide table comprises a primary key column and a plurality of column clusters corresponding to a plurality of online single tables; synchronizing, in real time, a plurality of first data records newly added in the corresponding online single tables into the target wide table by partially updating the target wide table, wherein the plurality of online single tables comprises a transaction detail table for recording transaction data, and the transaction detail table comprises at least a transaction time field and a merchant field; determining a plurality of second data records from the target wide table according to a target filtering condition, and generating a billing data table based on the plurality of second data records; the target filtering condition comprises that the field value of the transaction time field is within a target merchant-defined billing time range, and the field value of the merchant field is the target merchant.
2. The method of claim 1, wherein, the target filtering condition further comprises a plurality of target fields configured for the target merchant; the generating of the billing data table comprises: extracting a plurality of target field values corresponding to the plurality of target fields from each column cluster included in any second data record, and forming a target data record based on the plurality of target field values; generating a billing data table based on the target data records corresponding to the plurality of second data records. 3.The method of claim 1, further comprising: obtaining a plurality of offline single tables corresponding to the plurality of online single tables, wherein each offline single table is periodically updated, and each field included in the offline single table is the same as the corresponding online single table; performing a correctness check on the target wide table based on the plurality of offline single tables; the determining of the plurality of second data records from the target wide table comprises: determining the plurality of second data records from the target wide table after the correctness check is passed.
4. The method of claim 3, wherein, only the data records newly added in the corresponding online single table within a target time range are recorded under each column cluster, and the target time range includes the billing time range; the correctness check on the target wide table comprises: reading a plurality of offline data records with the field value of the transaction time field within the target time range from a target offline single table corresponding to the transaction detail table; comparing each offline data record with each first data record under a target column cluster corresponding to the target offline single table one by one to obtain a first comparison result; determining a check result of the correctness check according to the comparison results corresponding to the plurality of offline single tables.
5. The method of claim 1, wherein, the target wide table is created by a streaming computing engine.
6. The method of claim 5, wherein, the first data records newly added in each online single table are read from a corresponding intermediate table by the streaming computing engine, and each data record in the intermediate table is obtained by parsing a data change log of the online single table by a target middleware.
7. The method of claim 5, wherein, the streaming data warehouse is Paimon, and the streaming computing engine is Flink.
8. The method of claim 1, wherein, the plurality of online single tables further comprises one or more of the following: a payment detail table for describing payment information of a transaction; a settlement detail table obtained by aggregating the transaction data according to a settlement date. 9.An apparatus for generating billing data, disposed on an offline computing platform, comprising: A reading unit is configured to read a target wide table from a streaming data warehouse, wherein the target wide table includes a primary key column and a plurality of column clusters corresponding to a plurality of online single tables; A single column cluster synchronizes, in real time, a first data record added to the corresponding online single table in streaming mode by updating the target wide table in a partial manner; the plurality of online single tables includes a transaction detail table configured to record transaction data, and the transaction detail table includes at least a transaction time field and a merchant field; A generating unit is configured to determine each second data record from the target wide table according to a target filtering condition, and generate a bill data table based on the each second data record; The target filtering condition includes that a field value of the transaction time field is within a self-defined billing time range of a target merchant, and a field value of the merchant field is the target merchant.
10. A computer readable storage medium having stored thereon a computer program, wherein, When the computer program is executed in the computer, the computer is caused to perform the method in any one of claims 1-8.
11. A computing device comprising a memory and a processor, wherein, The memory stores executable code, and the processor executes the executable code to implement the method in any one of claims 1-8.