Business index acquisition method and device based on distributed processing and readable medium

By collecting and processing backend transaction data in a distributed processing system and designing a hierarchical database structure, the problems of slow and inaccurate indicator calculation in traditional marketing operations are solved, enabling more accurate marketing recommendations and more efficient data processing.

CN115841273BActive Publication Date: 2025-11-18CHINA ASSET MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211510984.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2025-11-18
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

Traditional marketing methods struggle to provide accurate and customer-centric marketing recommendations, and they suffer from slow computation and poor timeliness and stability of metrics when dealing with large datasets.

Method used

A distributed processing system is adopted to collect backend transaction data from the source database, process it based on partitioning and storage strategies, design a hierarchical structure for the database tables, and adopt a differentiated storage scheme to generate marketing metrics through distributed processing.

Benefits of technology

It enables more accurate marketing recommendations, improves the timeliness and stability of metrics, simplifies the calculation logic, and solves the problem of slow calculation of large amounts of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115841273B_ABST
    Figure CN115841273B_ABST
Patent Text Reader

Abstract

The application discloses a business index acquisition method and device based on distributed processing and a readable medium. The method comprises the following steps: collecting background transaction data from a source database to a distributed processing system as source layer data; partitioning and storing the source layer data to obtain partition storage data corresponding to different business requirements and / or functions as model layer data; processing the partition storage data corresponding to different business requirements and / or functions in the model layer data based on a distributed processing mode to obtain corresponding marketing indexes as data mart layer data; and finally outputting the data mart layer data to a downstream business system. The application can make more accurate and customer demand-conforming marketing recommendations by introducing background transaction data into a marketing scenario. By introducing a distributed processing system and designing a hierarchical structure of a database table for data storage and processing, the application solves the problems of slow calculation, poor accuracy and stability of indexes and the like in the traditional mode for large data volume.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data acquisition and processing technology, and in particular relates to a method, apparatus and readable medium for obtaining business indicators based on distributed processing. Background Technology

[0002] In the past, marketing operations mainly relied on data such as market conditions and distribution channels to calculate marketing metrics. However, this method is difficult to achieve accurate marketing recommendations that better meet customer needs.

[0003] In addition, traditional technologies also have problems with slow processing of large amounts of data when calculating marketing metrics based on data processing, making it difficult to meet the timeliness requirements of the metrics, as well as poor accuracy and stability of the metrics. Summary of the Invention

[0004] In view of this, this application provides a method, apparatus and readable medium for obtaining business metrics based on distributed processing, which combines distributed technology and information technology to solve at least some of the technical problems existing in traditional technologies when obtaining the required marketing metrics based on data processing for marketing recommendations and other applications.

[0005] The specific plan is as follows:

[0006] A method for obtaining business metrics based on distributed processing includes:

[0007] Backend transaction data is collected from the source database and sent to the distributed processing system as source-layer data.

[0008] Based on the partitioning and storage strategies corresponding to different business requirements and / or functions required for marketing, the source layer data is partitioned and stored to obtain partitioned storage data corresponding to different business requirements and / or functions, which serves as model layer data.

[0009] Based on a distributed processing approach, the partitioned storage data of different business requirements and / or functions in the model layer data is processed to obtain the marketing indicators required by each different business requirement and / or function, which serve as the data mart layer data.

[0010] The obtained data mart layer data is output to the corresponding downstream business system.

[0011] Optionally, the step of collecting backend transaction data from the source database to the distributed processing system as source-layer data includes:

[0012] Fund transaction and holding data from the source database are collected and sent to the distributed processing system as source-layer data; the fields in the source-layer data are consistent with those in the source data.

[0013] Optionally, based on the partitioning and storage strategies corresponding to different business requirements and / or functions needed for marketing, the source layer data is partitioned and stored to obtain partitioned storage data corresponding to different business requirements and / or functions, including:

[0014] Determine the time segmentation method required for calculating marketing metrics under different business requirements and / or functions;

[0015] Calculate the corresponding time sharding method based on marketing metrics under different business requirements and / or functions, and partition and store the database table of the source layer data.

[0016] Optionally, the source layer data includes at least a portion of the following types of data:

[0017] Transaction data, holdings data and fund data, user data, and dimensional data.

[0018] Optionally, the step of calculating the corresponding time sharding method based on marketing metrics under different business requirements and / or functions, and partitioning and storing the database table of the source layer data, includes:

[0019] Transaction data and user data are partitioned and stored according to transaction date or confirmation date;

[0020] Save a complete historical record of holdings and fund data from the start date to the present day;

[0021] The dimensional data is refreshed and stored once a day.

[0022] Optionally, the step of processing the partitioned storage data of the model layer data based on a distributed processing method to obtain the marketing indicators required by each different business requirement and / or function includes:

[0023] Based on the calculation logic corresponding to the marketing metrics required for different business requirements and / or functions, the corresponding partitioned storage data in the model layer data is processed, calculated, and adapted to the data format using a distributed processing method to generate the marketing metrics required for different business requirements and / or functions.

[0024] Optionally, outputting the obtained data mart layer data to the corresponding downstream business system includes:

[0025] The data from the marketplace layer is exported to downstream business systems using HDFS file operations.

[0026] Optionally, the step of exporting market layer data to downstream business systems using HDFS file operations includes:

[0027] A scheduled task outputs the HDFS files of the corresponding tables in the data mart layer to the relay server. Once the relay server detects that the files are ready for the corresponding business requirements and / or functions, it sends the corresponding files to the downstream marketing platform.

[0028] A business metric acquisition device based on distributed processing, comprising:

[0029] The data acquisition module is used to collect backend transaction data from the source database and send it to the distributed processing system as source layer data.

[0030] The partitioning and storage processing module is used to partition and store the source layer data based on the partitioning and storage strategies corresponding to different business requirements and / or functions required for marketing, so as to obtain partitioned storage data corresponding to different business requirements and / or functions, which serves as model layer data.

[0031] The data processing module is used to process the partitioned storage data of different business requirements and / or functions in the model layer data based on a distributed processing method, so as to obtain the marketing indicators required by each different business requirement and / or function, which serve as the data mart layer data.

[0032] The output module is used to output the obtained data mart layer data to the corresponding downstream business system.

[0033] A computer-readable medium having a computer program stored thereon, the computer program comprising program code for performing the methods described in any of the preceding descriptions.

[0034] In summary, the business metric acquisition method, apparatus, and readable medium based on distributed processing provided in this application collect backend transaction data from a source database to a distributed processing system as source-layer data. Based on a preset partitioning and storage strategy, the source-layer data is partitioned and stored to obtain partitioned storage data corresponding to different business requirements and / or functions, which serves as model-layer data. Furthermore, based on distributed processing, the partitioned storage data of different business requirements and / or functions in the model-layer data is processed to obtain the marketing metrics required by each different business requirement and / or function, which serve as data mart-layer data. Finally, the obtained data mart-layer data is output to the corresponding downstream business system.

[0035] This application introduces backend transaction data, such as fund trading and holding data, into the marketing scenario. This helps to quantitatively analyze customer characteristics and preferences, enabling more precise marketing recommendations that better meet customer needs. Furthermore, by introducing a distributed processing system for data storage and processing, the slow computation speed of traditional methods for large datasets is effectively addressed, meeting the timeliness requirements of the metrics. Moreover, by designing a hierarchical database table structure, marketing data is divided into a source layer, a model layer, and a data mart layer, with differentiated storage schemes for each layer. This simplifies the metric calculation logic and improves the accuracy and stability of the metrics. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0037] Figure 1 This is a flowchart of the business metric acquisition method based on distributed processing provided in this application;

[0038] Figure 2 This application provides an exemplary processing flow for calculating marketing metrics based on distributed processing of transaction and position data;

[0039] Figure 3 This is a structural diagram of the business indicator acquisition device based on distributed processing provided in this application. Detailed Implementation

[0040] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0041] This application discloses a method, apparatus, and readable medium for obtaining business metrics based on distributed processing. See [link to relevant documentation]. Figure 1 The flowchart shown illustrates a business metric acquisition method based on distributed processing. The method provided in this application includes the following processing steps:

[0042] Step 101: Collect backend transaction data from the source database to the distributed processing system as source layer data.

[0043] The back-end transaction data can be determined according to the actual business scenario; for example, it could be fund transaction holding data.

[0044] Previously, marketing operations primarily relied on market data and distribution channels to calculate marketing metrics. The applicant found that this approach was insufficient for precise and customer-centric marketing recommendations. The main reason was that the trading and holding information of existing customers who purchased the company's funds and other products was not given sufficient attention. Therefore, the applicant proposed introducing internal back-end trading and holding data into the marketing process to help quantitatively analyze customer characteristics and preferences. For example, based on the customer's trading instruments, trading frequency, and timing, the applicant could determine the customer's trading style and recommend products the customer is interested in during their preferred trading periods, thus enabling more precise and customer-centric marketing recommendations.

[0045] Correspondingly, this step can collect fund trading and holding data from the source database and send it to the distributed processing system as source layer data.

[0046] The fields and field data in the source layer data are consistent with those in the source data.

[0047] The collected source data, also known as the source layer data, includes, but is not limited to, some or all of the following types of data:

[0048] a. Transaction data

[0049] Such as user transaction records and user fixed investment records.

[0050] b. Holdings and Fund Data

[0051] Such as the user's daily holdings record.

[0052] c. User data

[0053] Such as the user's account information.

[0054] d. Dimensional data

[0055] Such as transaction channels and business types.

[0056] Step 102: Based on the partitioning and storage strategies corresponding to different business requirements and / or functions required for marketing, partition and store the source layer data to obtain partitioned storage data corresponding to different business requirements and / or functions, which will be used as model layer data.

[0057] To address the challenge of storing and processing massive amounts of internal back-end trading and position data using traditional databases, this application introduces a distributed processing system for data storage and computation. This system boasts advantages such as strong data storage capacity and high computation speed, effectively solving the problem of slow computation for large datasets and meeting the timeliness requirements of the indicators.

[0058] Meanwhile, addressing the challenges of calculating marketing metrics and utilizing a wide range of data dimensions, a hierarchical database table structure was designed, dividing marketing data into the aforementioned source layer, as well as the model layer and data mart layer. Furthermore, differentiated storage schemes were adopted for each layer to simplify metric calculation logic, improve metric accuracy and stability, and optimize data flow between different systems.

[0059] After obtaining the source layer data, this application embodiment further partitions and stores the source layer data based on the partitioning and storage strategies corresponding to different business requirements and / or functions required for marketing, and uses the partitioned storage data corresponding to different business requirements and / or functions as model layer data.

[0060] The aforementioned partitioning and storage strategies can be strategies related to time sharding.

[0061] Specifically, different types of data are used differently in the calculation of marketing metrics under the corresponding business requirements and / or functions. Based on this, the embodiments of this application match the source layer database tables with different time sharding methods according to different business requirements and functions, such as transaction data, holding data and fund data, user data, and dimension data in the source layer, and partition and store different data in the source layer according to the corresponding time sharding methods.

[0062] In this step, the time sharding method required for calculating marketing metrics for different business requirements and / or functions can be further determined, and the database table of the source layer data can be partitioned and stored according to the corresponding time sharding method calculated based on the marketing metrics for different business requirements and / or functions.

[0063] For example, transaction data and user data can be partitioned and stored according to transaction date or confirmation date. A full historical record of holding data and fund data from the start date to the present can be saved daily to calculate the holding amount through fund net value. Dimensional data can be fully refreshed and stored once a day. Model layer data built with this logic can meet the calculation of marketing indicators.

[0064] Step 103: Process the partitioned storage data of different business requirements and / or functions in the model layer data based on the distributed processing method to obtain the marketing indicators required by each different business requirement and / or function, which are used as the data mart layer data.

[0065] Among them, see reference Figure 2 The exemplary processing flow shown is based on distributed processing of transaction holding data to calculate marketing metrics. Specifically, according to the calculation logic corresponding to the marketing metrics required by different business requirements and / or functions, the corresponding partitioned storage data in the model layer data is processed, calculated and adapted in a distributed processing manner to generate the marketing metrics required by different business requirements and / or functions. The generated marketing metrics required by different business requirements and / or functions are then used as data mart layer data.

[0066] For example, common user daily transaction and position details indicators can be calculated only within the current day's partition, significantly improving computational efficiency compared to existing technologies that store the entire dataset in traditional databases. For more complex indicators such as historical transaction data, position details, and client profit and loss, the model layer first constructs time-partitioned data and performs backtracking calculations according to different historical dimensions. This addresses the pain point of existing technologies where historical data is not saved in traditional database designs or is stored using full tables or linked tables, resulting in extremely complex calculations of historical indicators and the inability to use historical data.

[0067] Step 104: Output the obtained data mart layer data to the corresponding downstream business system.

[0068] The data in the data mart layer can be directly applied to downstream business systems. After obtaining the data in the data mart layer, this embodiment uses HDFS (Hadoop Distributed File System) file operations to export the data to the downstream business system. That is, HDFS file operations are used instead of computing engine processing for data transfer between different systems. This embodiment relies on a distributed processing system and uses HDFS file operations to efficiently export data from the data mart layer directly to the downstream business system.

[0069] Specifically, the process of exporting data from the data mart layer to downstream business systems using HDFS file operations can be implemented, but is not limited to, as follows: A scheduled task reads the HDFS files from the corresponding tables in the data mart layer to a relay server; after the relay server detects that all event files for the corresponding business requirements and / or functions are ready, it packages the files and sends them to the downstream marketing platform; the downstream marketing platform unpacks the files and places them in the corresponding directory, mounts partitions, and creates partition tables for storing indicator data. Based on this, marketing indicator data can be queried and used in the downstream marketing platform.

[0070] In summary, the business metric acquisition method based on distributed processing provided in this application collects backend transaction data from the source database to the distributed processing system as source layer data. Based on a preset partitioning and storage strategy, the source layer data is partitioned and stored to obtain partitioned storage data corresponding to different business requirements and / or functions, which serves as model layer data. Furthermore, based on the distributed processing method, the partitioned storage data of different business requirements and / or functions in the model layer data is processed to obtain the marketing metrics required by each different business requirement and / or function, which serve as data mart layer data. Finally, the obtained data mart layer data is output to the corresponding downstream business system.

[0071] This application introduces backend transaction data, such as fund trading and holding data, into the marketing scenario. This helps to quantitatively analyze customer characteristics and preferences, enabling more precise marketing recommendations that better meet customer needs. Furthermore, by introducing a distributed processing system for data storage and processing, the slow computation speed of traditional methods for large datasets is effectively addressed, meeting the timeliness requirements of the metrics. Moreover, by designing a hierarchical database table structure, marketing data is divided into a source layer, a model layer, and a data mart layer, with differentiated storage schemes for each layer. This simplifies the metric calculation logic and improves the accuracy and stability of the metrics.

[0072] Corresponding to the above-mentioned business indicator acquisition method based on distributed processing, this application also provides a business indicator acquisition device based on distributed processing.

[0073] The device's structure is as follows: Figure 3 As shown, it includes a data acquisition module 10, a partitioning and storage processing module 20, a data processing module 30, and an output module 40.

[0074] The acquisition module 10 is used to collect backend transaction data from the source database to the distributed processing system as source layer data.

[0075] The back-end transaction data can be determined according to the actual business scenario; for example, it could be fund transaction holding data.

[0076] Previously, marketing operations primarily relied on market data and distribution channels to calculate marketing metrics. The applicant found that this approach was insufficient for precise and customer-centric marketing recommendations. The main reason was that the trading and holding information of existing customers who purchased the company's funds and other products was not given sufficient attention. Therefore, the applicant proposed introducing internal back-end trading and holding data into the marketing process to help quantitatively analyze customer characteristics and preferences. For example, based on the customer's trading instruments, trading frequency, and timing, the applicant could determine the customer's trading style and recommend products the customer is interested in during their preferred trading periods, thus enabling more precise and customer-centric marketing recommendations.

[0077] Correspondingly, the acquisition module 10 can collect fund trading and holding data from the source database and send it to the distributed processing system as source layer data.

[0078] The fields and field data in the source layer data are consistent with those in the source data.

[0079] The collected source data, also known as the source layer data, includes, but is not limited to, some or all of the following types of data:

[0080] a. Transaction data

[0081] Such as user transaction records and user fixed investment records.

[0082] b. Holdings and Fund Data

[0083] Such as the user's daily holdings record.

[0084] c. User data

[0085] Such as the user's account information.

[0086] d. Dimensional data

[0087] Such as transaction channels and business types.

[0088] The partitioning and storage processing module 20 is used to partition and store the source layer data based on a preset partitioning and storage strategy to obtain partitioned storage data corresponding to different business requirements and / or functions, which is then used as model layer data.

[0089] To address the challenge of storing and processing massive amounts of internal back-end trading and position data using traditional databases, this application introduces a distributed processing system for data storage and computation. This system boasts advantages such as strong data storage capacity and high computation speed, effectively solving the problem of slow computation for large datasets and meeting the timeliness requirements of the indicators.

[0090] Meanwhile, addressing the challenges of calculating marketing metrics and utilizing a wide range of data dimensions, a hierarchical database table structure was designed, dividing marketing data into the aforementioned source layer, as well as the model layer and data mart layer. Furthermore, differentiated storage schemes were adopted for each layer to simplify metric calculation logic, improve metric accuracy and stability, and optimize data flow between different systems.

[0091] After obtaining the source layer data, this application embodiment further partitions and stores the source layer data based on the partitioning and storage strategies corresponding to different business requirements and / or functions required for marketing, and uses the partitioned storage data corresponding to different business requirements and / or functions as model layer data.

[0092] The aforementioned partitioning and storage strategies can be strategies related to time sharding.

[0093] Specifically, different types of data are used differently in the calculation of marketing metrics under the corresponding business requirements and / or functions. Based on this, the embodiments of this application match the source layer database tables with different time sharding methods according to different business requirements and functions, such as transaction data, holding data and fund data, user data, and dimension data in the source layer, and partition and store different data in the source layer according to the corresponding time sharding methods.

[0094] The partitioning and storage processing module 20 can further determine the time sharding method required for calculating marketing indicators for different business requirements and / or functions, and partition and store the database table of the source layer data according to the corresponding time sharding method calculated based on the marketing indicators for different business requirements and / or functions.

[0095] For example, transaction data and user data can be partitioned and stored according to transaction date or confirmation date. A full historical record of holding data and fund data from the start date to the present can be saved daily to calculate the holding amount through fund net value. Dimensional data can be fully refreshed and stored once a day. Model layer data built with this logic can meet the calculation of marketing indicators.

[0096] The data processing module 30 is used to process the partitioned storage data of different business requirements and / or functions in the model layer data based on the distributed processing method, so as to obtain the marketing indicators required by each different business requirement and / or function, as the data mart layer data.

[0097] Specifically, according to the calculation logic corresponding to the marketing indicators required by different business requirements and / or functions, the corresponding partitioned storage data in the model layer data can be processed, calculated and adapted in a distributed processing manner to generate the marketing indicators required by different business requirements and / or functions, and the generated marketing indicators required by different business requirements and / or functions can be used as the data mart layer data.

[0098] For example, common user daily transaction and position details indicators can be calculated only within the current day's partition, significantly improving computational efficiency compared to existing technologies that store the entire dataset in traditional databases. For more complex indicators such as historical transaction data, position details, and client profit and loss, the model layer first constructs time-partitioned data and performs backtracking calculations according to different historical dimensions. This addresses the pain point of existing technologies where historical data is not saved in traditional database designs or is stored using full tables or linked tables, resulting in extremely complex calculations of historical indicators and the inability to use historical data.

[0099] Output module 40 is used to output the obtained data mart layer data to the corresponding downstream business system.

[0100] The data in the data mart layer can be directly applied to downstream business systems. After obtaining the data in the data mart layer, this embodiment uses HDFS (Hadoop Distributed File System) file operations to export the data to the downstream business system. That is, HDFS file operations are used instead of computing engine processing for data transfer between different systems. This embodiment relies on a distributed processing system and uses HDFS file operations to efficiently export data from the data mart layer directly to the downstream business system.

[0101] Specifically, the process of exporting data from the data mart layer to downstream business systems using HDFS file operations can be implemented, but is not limited to, as follows: A scheduled task reads the HDFS files from the corresponding tables in the data mart layer to a relay server; after the relay server detects that all event files for the corresponding business requirements and / or functions are ready, it packages the files and sends them to the downstream marketing platform; the downstream marketing platform unpacks the files and places them in the corresponding directory, mounts partitions, and creates partition tables for storing indicator data. Based on this, marketing indicator data can be queried and used in the downstream marketing platform.

[0102] In summary, the business indicator acquisition device based on distributed processing provided in this application collects backend transaction data from the source database to the distributed processing system as source layer data. Based on a preset partitioning and storage strategy, the source layer data is partitioned and stored to obtain partitioned storage data corresponding to different business requirements and / or functions, which serves as model layer data. The partitioned storage data of different business requirements and / or functions in the model layer data is then processed using a distributed processing method to obtain the marketing indicators required by each different business requirement and / or function, which serve as data mart layer data. Finally, the obtained data mart layer data is output to the corresponding downstream business system.

[0103] This application introduces backend transaction data, such as fund trading and holding data, into the marketing scenario. This helps to quantitatively analyze customer characteristics and preferences, enabling more precise marketing recommendations that better meet customer needs. Furthermore, by introducing a distributed processing system for data storage and processing, the slow computation speed of traditional methods for large datasets is effectively addressed, meeting the timeliness requirements of the metrics. Moreover, by designing a hierarchical database table structure, marketing data is divided into a source layer, a model layer, and a data mart layer, with differentiated storage schemes for each layer. This simplifies the metric calculation logic and improves the accuracy and stability of the metrics.

[0104] This application also provides a computer-readable medium having a computer program stored thereon, the computer program including program code for performing the distributed processing-based business metric acquisition method provided in the above method embodiments.

[0105] In the context of this application, a computer-readable medium (machine-readable medium) can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0106] It should be noted that the computer-readable medium described above in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0107] The aforementioned computer-readable medium may be contained within an electronic device or may exist independently without being assembled into an electronic device.

[0108] In summary, the business metric acquisition method, apparatus, and readable medium based on distributed processing provided in this application have at least the following technical advantages:

[0109] 1) The innovative introduction of back-end transaction and position data for marketing analysis scenarios allows for the creation of more diversified and precise marketing metrics compared to single market data and distribution channel data.

[0110] 2) To meet the requirements of marketing metric calculation, the data flow in the distributed processing system is designed according to the concept of layering. Layering helps to improve the accuracy and stability of the data and facilitates traceability.

[0111] 3) To calculate marketing metrics for complex scenarios or those that rely on historical data, different time partitions are used for storage and calculation based on different types of data;

[0112] 4) It innovatively adopts a highly efficient HDFS file operation method to replace the computing engine processing method for data transfer between different systems. Compared with processing by the computing engine, HDFS file operation has advantages such as faster processing speed, higher accuracy, and less susceptibility to errors in long-term tasks.

[0113] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0114] For ease of description, the above systems or devices are described separately as various modules or units based on their functions. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware components.

[0115] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0116] Finally, it should be noted that in this document, relational terms such as first, second, third, and fourth are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0117] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for obtaining business metrics based on distributed processing, characterized in that, include: Backend transaction data is collected from the source database and sent to the distributed processing system as source-layer data; wherein, the source-layer data includes at least a portion of transaction data, position data and fund data, user data, and dimensional data; Determine the time sharding method required for calculating marketing metrics under different business requirements and / or functions; partition and store transaction data and user data according to transaction date or confirmation date; save a full historical record of holding data and fund data from the start date to the present day; and refresh and store dimensional data once a day as model layer data. Based on the calculation logic corresponding to the marketing indicators required by different business requirements and / or functions, the corresponding partitioned storage data in the model layer data is processed, calculated and adapted in a distributed processing manner to generate the marketing indicators required by different business requirements and / or functions, which serve as the data mart layer data. The obtained data mart layer data is output to the corresponding downstream business system.

2. The method according to claim 1, characterized in that, The process of collecting backend transaction data from the source database and transmitting it to the distributed processing system as source-layer data includes: Fund transaction and holding data from the source database are collected and sent to the distributed processing system as source-layer data; the fields in the source-layer data are consistent with those in the source data.

3. The method according to claim 1, characterized in that, The step of outputting the obtained data mart layer data to the corresponding downstream business system includes: The data from the marketplace layer is exported to downstream business systems using HDFS file operations.

4. The method according to claim 3, characterized in that, The method of exporting market layer data to downstream business systems using HDFS file operations includes: A scheduled task outputs the HDFS files of the corresponding tables in the data mart layer to the relay server. Once the relay server detects that the files are ready for the corresponding business requirements and / or functions, it sends the corresponding files to the downstream marketing platform.

5. A business indicator acquisition device based on distributed processing, characterized in that, include: The data acquisition module is used to collect backend transaction data from the source database to the distributed processing system as source-layer data. The source-layer data includes at least a portion of transaction data, position data and fund data, user data, and dimensional data. The partitioning and storage processing module is used to determine the time sharding method required for the calculation of marketing indicators under different business requirements and / or functions. It partitions and stores transaction data and user data according to transaction date or confirmation date, saves a full historical record of holding data and fund data from the start date to the present every day, and refreshes and stores the dimension data every day as model layer data. The data processing module is used to process, calculate and adapt the corresponding partitioned storage data in the model layer data according to the calculation logic corresponding to the marketing indicators required by different business requirements and / or functions, based on the distributed processing method, so as to generate the marketing indicators required by different business requirements and / or functions, as the data mart layer data. The output module is used to output the obtained data mart layer data to the corresponding downstream business system.

6. A computer-readable medium, characterized in that, It stores a computer program thereon, the computer program containing program code for performing the method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Banking business data processing and querying method and device

    CN111161047A

  • User behavior statistical analysis method based on Flink streaming processing

    CN112000636A