Data processing method and related product

By reusing data models and metadata processing, the problem of information misalignment in the data query link is solved, the query efficiency and asset utilization are improved, and the consistency and efficient query of the data production link are achieved.

CN120705231APending Publication Date: 2025-09-26MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510837594.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In existing technologies, data production, analysis and processing, storage, and query are completed by different systems, resulting in information misalignment, low query efficiency, and repeated operations.

Method used

By reusing data models, we can uniformly and coherently process indicator data query and production links, use metadata to describe query requests, query and reuse data models and business logs, and perform online analysis and processing to obtain matching indicator data.

Benefits of technology

It achieves information alignment between indicator data query and production links, improves query efficiency and asset reuse rate, avoids duplicate production, and improves data delivery efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705231A_ABST
    Figure CN120705231A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and a related product, index data to be queried is obtained by multiplexing a data model, and links such as index data query and production are unified and coherent, so that the query efficiency and the asset reuse rate are improved, and repeated operation is avoided. The data processing method comprises the steps that in response to a query request, first index data matched with the query request is queried, and in response to query failure, first metadata used for describing the query request is determined based on the query request; querying a first data model matched with the first metadata; analyzing a business log generated by a data source based on the first data model to obtain first business data; and performing online analysis processing on the first business data based on the first metadata to obtain second index data matched with the query request.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data processing method and related products. Background Art

[0002] In order to continuously track the performance of specific aspects of the business, it is usually necessary to obtain real-time indicator data related to the business and use the Business Intelligence (BI) system to visualize, dynamically monitor and intelligently analyze the indicator data related to the business.

[0003] However, the above process usually involves various links such as data production, analysis and processing, storage and query. Currently, each link is completed by a different system. There are problems of single points and information misalignment between these systems, as well as problems such as low query efficiency and repeated operations. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a data processing method and related products, which obtain the indicator data to be queried by reusing the data model, unify and connect the indicator data query, production and other links, so as to improve the query efficiency and asset reuse rate and avoid repeated operations.

[0005] In order to achieve the above objectives, the embodiments of the present application adopt the following technical solutions: In a first aspect, an embodiment of the present application provides a data processing method, comprising: In response to a query request, querying first indicator data matching the query request, and in response to a query failure, determining first metadata for describing the query request based on the query request; Querying a first data model that matches the first metadata; Parsing the business log generated by the data source based on the first data model to obtain first business data; Based on the first metadata, the first business data is analyzed and processed online to obtain second indicator data that matches the query request.

[0006] In a second aspect, an embodiment of the present application provides a data processing device, including: a first query module configured to, in response to a query request, query first indicator data matching the query request, and, in response to a query failure, determine first metadata describing the query request based on the query request; A first determining module, configured to query a first data model that matches the first metadata; A second query module, configured to query a first data model that matches the first metadata; an acquisition module, configured to parse the business log generated by the data source based on the first data model to obtain first business data; An analysis module is used to perform online analysis and processing on the first business data based on the first metadata to obtain second indicator data that matches the query request.

[0007] In a third aspect, an embodiment of the present application provides an electronic device, including: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the data processing method provided in the first aspect.

[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the data processing method provided in the first aspect.

[0009] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute some or all of the steps in the data processing method provided in the first aspect.

[0010] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: In response to a query request, the system queries for first indicator data that matches the query request. If the query fails, the system determines first metadata describing the query request, such as indicator items and dimensions, and queries a first data model that matches the first metadata. If the query succeeds, the system directly reuses the first data model to obtain first business data, and performs online analysis and processing on the first business data based on the first metadata to obtain second indicator data that matches the query request. As a crucial asset in the data production process, reusing data models unifies and connects the query and production stages of indicator data, enabling information alignment between these stages. This avoids duplicate infrastructure data production compared to implementing different stages in separate systems, improving query efficiency and asset reuse. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A schematic diagram of the structure of a data platform provided for one embodiment of the present application; Figure 2 A schematic diagram of the structure of a data platform provided in another embodiment of the present application; Figure 3 A schematic diagram of the structure of a data platform provided in yet another embodiment of the present application; Figure 4 A flowchart of a data processing method provided in one embodiment of the present application; Figure 5 A schematic diagram of a query interface provided for one embodiment of the present application; Figure 6 A schematic diagram of a filling interface provided for an embodiment of the present application; Figure 7 A schematic diagram of a metadata configuration interface provided for one embodiment of the present application; Figure 8 A flowchart of a data processing method provided in another embodiment of the present application; Figure 9 A flowchart of a data processing method provided in yet another embodiment of the present application; Figure 10 A schematic structural diagram of a data processing device provided in accordance with an embodiment of the present application; Figure 11 A schematic structural diagram of an electronic device provided in accordance with an embodiment of the present application. DETAILED DESCRIPTION

[0012] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0013] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0014] Some concept descriptions: DataOps (Data Operations): A process-oriented approach used by data teams to improve R&D efficiency, data quality, and delivery efficiency. DataOps unifies data users and data production teams to deliver analytical solutions and products faster and more accurately, achieving goals such as continuous data iteration, continuous integration, and continuous delivery. DataOps has evolved into a unique approach to data processing, preparation, and analysis.

[0015] Data R&D Center: Provides various data production capabilities including offline and real-time production capabilities on the data production side.

[0016] Data Operation and Maintenance Center: Provides data quality assurance capabilities around the data operation and maintenance side, ensuring data timeliness and quality from two perspectives: Service-Level Agreement (SLA) and Dynamic Quality Control (DQC).

[0017] Data Asset Center: Focusing on data asset management, it provides users with a better platform capability to find and use data.

[0018] Real-time BI pipeline: A pipeline method for producing real-time BI based on business data needs. The real-time BI pipeline capability generally includes the following related links: 1. Business data: such as transaction data, activity data, user operation tracking data, and other business-related dimension data; 2. Data integration: unified collection and distribution of various business logs (such as binlog) and embedded data, as well as data transmission including Kafka to Doris; 3. Real-time data warehouse: Based on data integration, real-time detailed data is accessed, cleaned, trimmed, converted, merged, and output; 4. Real-time Online Analytical Processing (OLAP): Based on detailed data from a real-time data warehouse, different models are used for modeling, such as detail models, aggregate models, and primary key models, depending on the analysis direction. 5. Model management: logical model management of OLAP data and binding management between model and indicators, and between model and dimensions; 6. Metadata management: unified dimension management, indicator management, business domain management, and other metadata management for data applications; 7. BI Report: Real-time BI report configuration and delivery based on OLAP data tables; 8. Data Operations and Maintenance: This includes ensuring the timeliness (SLA) and data quality (DQC) capabilities of the entire data production chain.

[0019] As mentioned above, the process of obtaining real-time indicator data related to the business from the business platform for visualization, dynamic monitoring and intelligent analysis involves various links such as data production, analysis and processing, storage and query. Currently, each link is completed by a different system. There are problems such as single points and information misalignment between these systems, as well as low query efficiency and repeated operations.

[0020] In view of this, an embodiment of the present application proposes a data processing method. When a data user (such as an operator) requests to query the required indicator data, if there is reusable first indicator data in the asset center's existing indicator data, the first indicator data is directly returned as the query result, thereby reusing the asset center's existing indicator data, avoiding duplication of indicator data, and improving query efficiency and asset reuse rate. If the first indicator data does not exist in the asset center's existing indicator data, first metadata is constructed to describe the query request, such as indicator items and dimensions, and the asset center's existing infrastructure data is queried to see if there is a reusable first data model. If the first data model exists, the first data model is directly reused to parse the business log of the data source to obtain the first business data, and the first business data is analyzed and processed online based on the first metadata to obtain second indicator data that matches the query request. The data model is an important asset in the data production process. By reusing the data model, the query and production links of the indicator data are unified and coherent, and information alignment between these links can be achieved. Compared with implementing different links in different systems, this avoids duplication of infrastructure data, improving query efficiency and asset reuse rate.

[0021] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.

[0022] Before introducing the data processing method provided in the embodiment of the present application, Figures 1 to 3 , introduces the data platform provided in the embodiments of this application.

[0023] like Figure 1 As shown, the data platform provided in the embodiment of the present application includes basic resources, a data asset center, a data R&D center, and a data operation and maintenance center.

[0024] Basic resources include computing resources and storage resources. Computing resources can include, for example, Spark on YARN, Flink on YARN, and Hive on YARN, ensuring data processing for each center. Storage resources can include, but are not limited to, HBASE, Doris, HDFS, and Kafka, ensuring data storage for each center.

[0025] The Data Asset Center stores various assets, including but not limited to algorithmic assets and data assets. Algorithmic assets include feature datasets, sample assets, long-sequence assets, and feature synchronization assets. Data assets include data maps, infrastructure data, indicator data for real-time BI, and real-time scenario data.

[0026] The Data Asset Center is integrated with the metadata platform, providing operations personnel with metadata configuration capabilities for querying indicator data. Furthermore, the Data Asset Center integrates with the BI system, allowing it to retrieve the required indicator data from existing data assets, such as real-time multi-dimensional analysis indicator data, and transmit this data to the BI system for business analysis, marketing activity analysis, and more.

[0027] The Data R&D Center is responsible for data development, including both basic and applied data. Basic data R&D encompasses data source management, dimensional model management, model management, and OLAP model management. Application data R&D encompasses the production of application data such as real-time monitoring, real-time features and indicators, real-time scenarios, and offline features. The Data R&D Center is integrated with the BI system, providing the system with real-time business data from OLAP models and real-time data warehouses for business performance analysis and real-time monitoring.

[0028] like Figure 2 As shown, the Data R&D Center includes various UI modules and their corresponding configuration backends. These UI modules provide entry points for data producers (such as R&D personnel) to interact with the Data R&D Center. Examples include, but are not limited to, the real-time data warehouse UI module, the real-time OLAP UI module, the real-time BI management UI module, and the metadata management UI module. The configuration backend corresponding to the real-time data warehouse UI module implements functions such as data source management, dimensional model management, and model management. The configuration backend corresponding to the real-time OLAP UI module implements functions such as OLAP source management and model management. The real-time BI management UI module implements BI management functions, and the metadata management UI module implements indicator data management functions.

[0029] The Data Operation and Maintenance Center is responsible for monitoring the timeliness and accuracy of the data produced by the Data R&D Center, and provides operation and maintenance functions such as duty management, review management, knowledge base and QA archiving. The Data Operation and Maintenance Center is connected with the BI system to provide data quality assurance for the BI system to conduct real-time risk control and early warning. Figure 2 As shown, the operation and maintenance center includes the SLA dashboard module, SLA management module, DQC management module, cost management module, and process management module.

[0030] In one implementation, after the data R&D center launches the production of various types of data, they can be synchronized to the data asset center according to the configured asset category; the data asset center abstracts the data asset logic layer, such as writing offline data assets and real-time data assets into the same business category for maintenance, governance, and use; the data operation and maintenance center ensures the timeliness and quality of the data.

[0031] In this way, through the above-mentioned data platform, the originally single-point distributed and disjointed links such as data production (R&D), query, and operation and maintenance are formed into a whole in product form, providing one-stop services that are visible, traceable, configurable, and operable for R&D personnel, operations personnel, and operation and maintenance personnel, thereby improving data delivery efficiency.

[0032] In practical applications, such as Figure 3 As shown in the figure, at the software level, the aforementioned data platform can be divided into the infrastructure layer, the engine layer, and the application layer. The infrastructure layer includes various infrastructure components, such as MySQL, Kafka, HBASE, and Flink clusters. The engine layer provides processing engines such as Flink-based data integration, real-time data warehouse processing, and OLAP modeling. The application layer serves as the entry point for user interaction with the data platform and primarily provides core components such as rtco-web-ui (real-time collaboration frontend), rtcp-web (real-time console), rtcp-model (model management service), rtcp-bi-admin (BI management backend), and BI.

[0033] Based on the above data platform, this application embodiment proposes a data processing method. Figure 4 , is a flow chart of a data processing method provided in one embodiment of the present application. The method can be applied to the above-mentioned data platform, and the method includes the following steps: S402: In response to the query request, query first indicator data that matches the query request.

[0034] The data platform provides a query interface. Data users (such as operators and product managers) can use the query interface to input relevant information about the indicator data to be queried. For example, Figure 5An example of a query interface is shown. Data users can use the input controls of the interface to input relevant information such as the indicator name, type, business manager, technical manager, and operating status of the indicator data to be queried, and click the query button to trigger a query request to be sent to the data platform. The query request contains the relevant information filled in by the data user.

[0035] The data platform responds to the query request and, based on the relevant information contained therein, queries the existing indicator data of the data asset center to see whether there is indicator data that matches the relevant information. If so, the indicator data is determined as the first indicator data, thereby realizing the reuse of the existing indicator data of the data asset center, avoiding the duplication of indicator data, and improving the utilization rate and query efficiency of the indicator data. If there is no indicator data that matches the relevant information in the existing indicator data of the data asset center, the query is determined to have failed.

[0036] For example, the query request is used to request query of indicator data such as the number of withdrawal applications, the number of withdrawal reviews initiated by investors, the number of withdrawal reviews by investors, the number of successful withdrawal reviews by investors, and the investor withdrawal review failure index. If the indicator data already available in the data asset center includes the number of withdrawal applications but does not include other indicator data, then it is determined that the query for the indicator data of the number of withdrawal applications is successful, and the indicator data already available in the data asset center is fed back to the data user, and it is determined that the query for other indicator data has failed, and the following steps S404 and S410 are executed for the indicator data for which the query failed.

[0037] S404 : In response to the query failure, determine first metadata for describing the query request based on the query request.

[0038] The first metadata may include but is not limited to at least one of the index item (also referred to as the first index item) requested by the query request, dimension (also referred to as dimension), caliber, business description information, etc.

[0039] In response to the query failure, the data asset center may extract relevant information of the indicator data to be queried from the query request, and send the relevant information to the data research and development center, which determines the first metadata.

[0040] The data platform can also display a filling interface to the data user, so that the data user can fill in the basic information of the indicator data to be queried, and then the data R&D center can obtain the first metadata based on the query requirements filled in by the data user. Figure 6 An example of a filling interface is shown, where data users can enter basic information such as report name, report description, and requirement wiki through the input controls of the interface.

[0041] The data research and development center queries the existing metadata to see if there is metadata that can describe the query request. If so, the metadata is reused and determined as the first metadata to improve the utilization rate of the metadata and avoid repeated metadata generation. If not, the first metadata is generated based on the received relevant information or query requirements. For example, the query request is used to request the number of withdrawal applications in a certain time period, and the corresponding metadata includes the indicator item "number of withdrawal applications" and the dimension "time". If these metadata exist in the data research and development center, they are directly reused; if these metadata do not exist in the data research and development center, they are regenerated.

[0042] In the application, in response to the failure to find reusable metadata, the data R&D center can also display the query requirements filled in by the data user to the data R&D personnel, and the data user can confirm according to the query requirements. In addition, the data R&D center can also display the metadata configuration interface to the data R&D personnel, and the data R&D personnel can configure the corresponding metadata according to the query requirements, and then determine the metadata filled in by the data R&D personnel as the first metadata. For example, Figure 7 This shows an example of a metadata configuration interface. Data developers can use the input controls on this interface to configure metadata such as dimension name, English name, Chinese name, indicator name, English name, Chinese name, and indicator caliber. This approach connects data users and producers, achieving information alignment between them.

[0043] S406: Query the first data model that matches the first metadata. The infrastructure data of the asset center may include, but is not limited to, at least one of a data model used to process business logs and business data obtained after processing business logs. The data model is a logical model configured based on real-time data streams, an abstraction of real-time data streams. It defines various information such as data sources, filtering conditions, fields, real-time confluence, dimension expansion conditions, data processing methods, and the destination where the processed data is written.

[0044] The first data model refers to a data model used to obtain business data of the first business to which the first metadata belongs. In one implementation, S406 includes the following steps: determining the first business to which the first metadata belongs based on the first indicator item; querying the data model accessed by the first business, and determining the data model as the first data model that matches the first metadata.

[0045] For example, if the first metadata includes the first indicator item "Number of Successfully Approved Investor Withdrawals" and the first dimension "City," then the first business to which it belongs is the withdrawal business. Furthermore, the data model accessed by the withdrawal business is determined to be the first data model. The first data model is used to obtain business data for the withdrawal business.

[0046] In this way, the data model of the first business access to which the first metadata belongs can be reused, ensuring that accurate business data can be subsequently obtained through the data model for use in subsequent production of indicator data requested by query requests.

[0047] S408: Parse the business log generated by the data source based on the first data model to obtain first business data.

[0048] Business logs are files or data collections that record various operations, events, status, data, and other related information during business processing. Business logs can include various types of logs, such as JSON and Binlog.

[0049] In one implementation, the above S408 includes the following steps: determining a first filtering condition based on the first metadata; if the first data model includes the first filtering condition, filtering the business log generated by the data source based on the first data model to obtain a first business log, and parsing the first business log to obtain first business data; if the first data model does not include the first filtering condition, changing the first data model based on the first filtering condition to obtain a second data model, and filtering the business log generated by the data source based on the second data model to obtain a second business log, and parsing the second business log to obtain the first business data.

[0050] For example, if the first metadata includes the first indicator item "number of successful withdrawal reviews by investors" and the first dimension "city", then the first filtering condition is: filter out the business logs with successful withdrawal reviews in City A. If the first data model includes this filtering condition, the first filtering condition is used to filter out the business logs with successful withdrawal reviews in City A from the business logs generated by the data source; then, using the data processing method in the first data model, the filtered business logs are parsed and processed to obtain the order data with successful withdrawal reviews in City A. These order data are the first business data used to generate the indicator data "number of successful withdrawal reviews by investors in City A" requested by the query request.

[0051] If the first data model does not include the above-mentioned first filtering condition, the first filtering condition is added to the first data model to obtain a second data model containing the first filtering condition; then, the first filtering condition is used to filter out the business logs of successful withdrawal review in City A from the business logs generated by the data source; then, the filtered business logs are parsed and processed using the data processing method in the second data model to obtain the order data of successful withdrawal review in City A, and these order data are determined as the first business data.

[0052] Through the above implementation method, the existing data models of the data asset center can be reused, avoiding the repeated creation of data models, and improving the asset utilization rate of the asset center and the query efficiency of indicator data.

[0053] In another embodiment, after querying the first data model that matches the first metadata, in response to the query failure, a third data model is created based on the model template that matches the first metadata; the business log generated by the data source is parsed based on the third data model to obtain the first business data.

[0054] The model template that matches the first metadata means that a data model created based on the model template can be used to process the first business to which the first metadata belongs, such as the business to which the first indicator item in the first metadata belongs.

[0055] For example, the first metadata includes the first indicator item "the number of successful withdrawal reviews by investors" and the first dimension "city". If the first business to which it belongs is the withdrawal business, the model template corresponding to the withdrawal business is used. The model template defines the data source corresponding to the withdrawal business, the format of the filter conditions, the data processing method corresponding to the withdrawal business, the destination to which the processed data is written, etc. Further, based on the first metadata, the first filter condition required to obtain the business data of the business is determined, and the first filter condition is written into the model template to obtain the third data model. Similar to the specific method of parsing the business log generated by the data source based on the first data model mentioned above, the first business data is obtained by parsing the business log generated by the data source based on the third data model.

[0056] Through the above implementation method, a third data model can be quickly and conveniently constructed to obtain the first business data, so as to produce the index data requested by the query request based on the first business data, thereby improving the production and query efficiency of the index data.

[0057] S410: Perform online analysis and processing on the first business data based on the first metadata to obtain second indicator data that matches the query request.

[0058] The second indicator data that matches the query request refers to the indicator data requested by the query request.

[0059] Online Analytical Processing (OLAP) is a technology used to support complex data analysis operations. It allows multidimensional data to be sliced, diced, drilled, rotated, and other operations, thereby performing in-depth analysis of the data from different angles and levels.

[0060] In one implementation, the above S410 includes the following steps: S4102: Create a fourth data model for analyzing the first business data based on the first metadata.

[0061] The fourth data model is also called an OLAP model, which is used to better organize and manage the first business data to meet data analysis needs. The fourth data model may include but is not limited to: a detail model, an aggregation model, a primary key model, etc.

[0062] The detail model directly retains detailed information from the original data. Each row of data corresponds to a specific business event or record. It does not perform any aggregation operations and can provide a complete and detailed data view. The detail model is suitable for scenarios that require in-depth exploration and analysis of business details. For example, in e-commerce, analyzing the purchase behavior of a specific user provides a basis for precision marketing and personalized recommendations. The detail model usually exists as a large wide table, including all relevant business fields. For example, an e-commerce detail table will include fields such as order ID, user ID, product ID, product name, product category, purchase time, purchase quantity, and purchase amount.

[0063] An aggregation model is built by summarizing and analyzing detailed data to a certain degree. Based on business analysis needs, it aggregates detailed data along certain dimensions to generate an aggregated data table. Aggregation models are suitable for scenarios requiring macro-level statistics and analysis of data. For example, they can calculate the total sales of different product categories over different time periods or count the number of successfully approved withdrawals in a specific city during a specific period. The aggregation model creates corresponding aggregate tables based on the analysis dimensions. For example, it aggregates the data of successfully approved withdrawal orders by city to generate an aggregate table containing fields such as city, order ID, withdrawal status, and approval result.

[0064] The primary key model is based on the business primary key. It uses the business primary key as the key field to link related detailed data. The primary key model is suitable for scenarios requiring in-depth analysis of specific business objects. For example, in e-commerce, the order ID is used as the primary key to link order details, user information, product information, and other related data to form a complete order model, allowing for in-depth understanding of each order. The primary key model typically creates multiple tables, linked by primary keys.

[0065] Because the first metadata can reflect the analysis requirements for the first business data, and different OLAP models are applicable to different scenarios and data analysis requirements, a fourth data model can be created based on the first metadata to analyze the first business data. For example, if the first metadata includes the first indicator item "Number of Successfully Approved Investor Withdrawals" and the first dimension "City," this reflects the analysis requirements for the first business data: aggregating the order data for successfully approved withdrawals in a specific city, and then creating an aggregate model as the fourth data model.

[0066] S4104: Perform online analysis and processing on the first business data based on the fourth data model to obtain second indicator data that matches the query request.

[0067] Specifically, the first business data is summarized based on the fourth data model to obtain a first data table; based on the first data table, the indicator value of the first indicator item under the first dimension is counted, and the indicator value is determined as the second indicator data.

[0068] For example, if the fourth data model is an aggregation model, and the first metadata includes the first indicator items "order volume" and "order total" and the first dimensions "city" and "date," then a data table containing fields such as "order ID," "order amount," "city," and "date" is generated, and the field values ​​of these fields are extracted from the first business data. The extracted field values ​​are then populated into the data table to obtain the first data table. Then, the order IDs in the first data table are counted according to the first dimension "city" to obtain the order volume for each city; the order amounts in the first data table are counted according to the first dimension "date" to obtain the order total for each city; the order IDs in the first data table are counted according to the first dimension "date" to obtain the order volume for different dates; and the order amounts in the first data table are counted according to the first dimension "date" to obtain the order total for different dates.

[0069] In an application, the first business data can be stored in a data warehouse. A computing engine such as Flink cleans, filters, and transforms the first business data before sending it to a message-based middleware such as Kafka. By sequentially consuming the first business data in the message-based middleware, an OLAP model is created as the fourth data model. The fourth data model is used to aggregate the first business data to generate a first data table, which is then written to an OLAP database (such as Doris or StarRocks). In the OLAP database, based on the first data table, the indicator values ​​of the first indicator item under the first dimension are calculated to obtain the second indicator data.

[0070] Of course, it should be understood that the above S410 can also analyze the first business data through various OLAP methods in the art to obtain the second indicator data, and this embodiment of the application is not limited to this.

[0071] In another embodiment, after determining the indicator value of the first indicator item in the first dimension as the second indicator data, it also includes: storing the second indicator data in a data asset center for subsequent asset management and reuse.

[0072] In another embodiment, after determining the indicator value of the first indicator item in the first dimension as the second indicator data, it also includes: determining a first field that matches the first indicator item from the fields contained in the first data table, and binding the first field to the first indicator item; determining a second field that matches the first dimension from the fields contained in the first data table, and binding the second field to the first dimension.

[0073] The first field that matches the first indicator item is the field required to determine the indicator value of the first indicator item. For example, if the first data table contains fields such as "Order ID," "Order Amount," "City," and "Date," the first field that matches the first indicator item "Order Amount" includes "Order Amount," and the first field that matches the first indicator item "Order Quantity" includes "Order ID."

[0074] The second field matching the first dimension refers to the field required for determining statistics based on the first dimension. For example, using the first data table above as an example, the second field matching the first dimension "City" includes "City", and the second field matching the first dimension "Date" includes "Date".

[0075] In the above embodiment, by binding the first indicator item to the first field that matches it in the first data table, and binding the first dimension to the second field that matches it in the first data table, real-time indicator data that meets the query requirements can be quickly and conveniently obtained from the first data table as the first business data increases. This meets the real-time BI analysis requirements.

[0076] In another embodiment, during the process of producing the second indicator data, SLA, DQC, etc. can also be enabled to send relevant data of the process to the data operation and maintenance center, which will ensure the quality of the second indicator data.

[0077] Specifically, SLA includes end-to-end SLQ and SLA for operational production efficiency. For example, in the process of producing the second indicator data, the method of tracking and reporting is adopted, and the output of each production link is uniformly reported to the data operation and maintenance center through the same interface, and the data operation and maintenance center calculates the end-to-end SLA. Among them, the tracking specification includes SLA name, level, log generation time, and the time when the log enters the current production link. In this way, different real-time production links can be protected by graded protection, such as protecting the production process of the second indicator data according to the levels of core, non-core, and trial period.

[0078] For DQC, important real-time data from the production process is written to the database through message middleware such as Kafka. Then, by configuring monitoring indicators such as the data null value rate, year-on-year data volume information, and month-on-month data volume information, alarms are issued for abnormal links in the production process.

[0079] The data processing method provided in the embodiment of the present application, in response to a query request, queries the first indicator data that matches the query request, in response to a query failure, determines the first metadata used to describe the query request, such as indicator items and dimensions, and queries the first data model that matches the first metadata, in response to a query success, directly reuses the first data model to obtain the first business data, and performs online analysis and processing on the first business data based on the first metadata to obtain the second indicator data that matches the query request. As an important asset in the data production process, the data model is reused to unify and connect the query and production links of the indicator data, and to achieve information alignment between these links. Compared with implementing different links in different systems, it avoids duplicate production of infrastructure data and improves query efficiency and asset reuse rate.

[0080] To facilitate understanding of the data processing method provided in the above embodiment of the present application, Figure 8 and Figure 9 , described with a specific embodiment.

[0081] like Figure 8 As described above, based on BI analysis requirements, operators send query requests to the data asset center, requesting indicator data that meets the BI analysis requirements. In response to the query request, the data asset center checks whether reusable indicator data (i.e., the first indicator data that matches the query request) exists. If so, the indicator data is reused and sent to the BI system for BI analysis. If not, the real-time BI module displays a fill-in interface to the operator, allowing the operator to fill in the query requirements, such as basic information about the indicator data to be queried.

[0082] The real-time BI module sends the received query requirements to the indicator dimension management module of the data R&D center, which generates the first metadata for describing the query request and delivers it to the data R&D personnel for confirmation.

[0083] Afterwards, the real-time BI module queries the existing infrastructure data of the data asset center to see whether there is a reusable data model (i.e., the first data model that matches the first metadata); if a reusable data model exists, the data model is reused, and the business log generated by the data source is parsed based on the data model to obtain the first business data, and the first business data is written into the real-time data warehouse of the data R&D center; if a reusable data model does not exist, a corresponding data model is created, and the data source is connected to the data model through the data integration module of the data R&D center, and the business log is parsed by the data model to obtain the first business data, and the first business data is written into the real-time data warehouse.

[0084] Furthermore, the OLAP module of the data R&D center performs OLAP analysis on the first business data in the real-time data warehouse to obtain the first data table and the second indicator data. Specifically, Figure 9 As shown, performing OLAP analysis on the first business data includes: creating a fourth data model for analyzing the first business data based on the first metadata, such as a detailed model, an aggregation model, or a primary key model; aggregating the first business data based on the fourth data model to obtain a first data table; and based on the first data table, counting the indicator value of the first indicator item under the first dimension (i.e., the second indicator data that matches the query request).

[0085] Afterwards, the OLAP module not only sends the first data table to the indicator dimension management module, which binds the fields in the first data table with the first metadata, but also sends the second indicator data to the real-time BI module, which then sends it to the data asset center for storage for subsequent asset management and reuse. Figure 9 As shown, the data integration module consumes the first data table in the OLAP database using the Kafka to Doris technology and transmits it to the indicator dimension management module.

[0086] In addition, the data asset center can also send the second indicator data to the BI system for BI analysis to meet the BI analysis needs of operators.

[0087] In addition, the real-time BI module can also enable data assurance functions, such as SLA and DQC functions, to monitor and issue alarms for the production process of the second indicator data.

[0088] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0089] Based on the same inventive concept, the present application also provides a data processing device. Figure 10 , is a structural diagram of a data processing device 1000 provided in an embodiment of the present application, and the device 1000 includes: a first query module 1010, a first determination module 1020, a second query module 1030, an acquisition module 1040 and an analysis module 1050.

[0090] The first query module 1010 is used to query first indicator data matching the query request in response to the query request.

[0091] The first determining module 1020 is configured to determine, in response to a query failure, first metadata for describing the query request based on the query request.

[0092] The second query module 1030 is used to query a first data model that matches the first metadata.

[0093] The acquisition module 1040 is used to parse the business log generated by the data source based on the first data model to obtain first business data.

[0094] The analysis module 1050 is used to perform online analysis and processing on the first business data based on the first metadata to obtain second indicator data that matches the query request.

[0095] In another embodiment, the first metadata includes a first index item; The second query module is used for: Determining, based on the first indicator item, a first business to which the first metadata belongs; A data model of the first service access is queried, and the data model is determined to be a first data model that matches the first metadata.

[0096] In another embodiment, when the acquisition module parses the business log generated by the data source based on the first data model to obtain the first business data, the acquisition module performs the following steps: determining a first filtering condition based on the first metadata; If the first data model includes the first filtering condition, filtering the business log generated by the data source based on the first data model to obtain a first business log, and parsing the first business log to obtain first business data; If the first data model does not include the first filtering condition, the first data model is changed based on the first filtering condition to obtain a second data model, and the business log generated by the data source is filtered based on the second data model to obtain a second business log, and the second business log is parsed to obtain the first business data.

[0097] In another embodiment, the acquisition module is further configured to: In response to a failure in querying the first data model, creating a third data model based on a model template matching the first metadata; The business log generated by the data source is parsed based on the third data model to obtain the first business data.

[0098] In another embodiment, the analysis module is configured to: Creating a fourth data model for analyzing the first business data based on the first metadata; The first business data is analyzed and processed online based on the fourth data model to obtain second indicator data that matches the query request.

[0099] In another embodiment, the first metadata includes a first index item and a first dimension; When the analysis module performs online analysis and processing on the first business data based on the fourth data model and obtains second indicator data that matches the query request, the analysis module executes the following steps: Summarize the first business data based on the fourth data model to obtain a first data table; Based on the first data table, counting the index value of the first index item in the first dimension; The indicator value is determined as the second indicator data.

[0100] In another embodiment, the analysis module is further configured to: Determining a first field that matches the first index item from the fields included in the first data table, and binding the first field to the first index item; A second field matching the first dimension is determined from the fields included in the first data table, and the second field is bound to the first dimension.

[0101] Obviously, the data processing device provided in the embodiment of the present application can be used as the above Figure 4 The execution subject of the data processing method shown in FIG. Figure 4 Since the principle is the same, the functions realized are not described here.

[0102] Figure 11 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 11 At the hardware level, the electronic device includes a processor and, optionally, an internal bus, a network interface, and memory. The memory may include internal memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for its services.

[0103] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 11 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0104] The memory is used to store programs. Specifically, the program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.

[0105] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a data processing device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations: In response to a query request, querying first indicator data matching the query request, and in response to a query failure, determining first metadata for describing the query request based on the query request; Querying a first data model that matches the first metadata; Parsing the business log generated by the data source based on the first data model to obtain first business data; Based on the first metadata, the first business data is analyzed and processed online to obtain second indicator data that matches the query request.

[0106] The above application Figure 4 The methods performed by the data processing devices disclosed in the illustrated embodiments can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be performed by hardware integrated logic circuits within the processor or by software instructions. The processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules within the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0107] The electronic device may also perform Figure 4 Method, and realize the data processing device in Figure 4 、 Figure 8 、 Figure 9 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.

[0108] Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0109] The embodiment of the present application also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by an electronic device including multiple application programs, can enable the electronic device to execute Figure 4 The method of the embodiment shown is specifically used to perform the following operations: In response to a query request, querying first indicator data matching the query request, and in response to a query failure, determining first metadata for describing the query request based on the query request; Querying a first data model that matches the first metadata; Parsing the business log generated by the data source based on the first data model to obtain first business data; Based on the first metadata, the first business data is analyzed and processed online to obtain second indicator data that matches the query request.

[0110] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute some or all of the steps in the data processing method provided in the embodiment of the present application.

[0111] In short, the above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

[0112] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0113] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0114] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0115] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

Claims

1. A data processing method, characterized in that: include: In response to a query request, querying first indicator data matching the query request, and in response to a query failure, determining first metadata describing the query request based on the query request; Querying a first data model that matches the first metadata; Parsing the business log generated by the data source based on the first data model to obtain first business data; Based on the first metadata, the first business data is analyzed and processed online to obtain second indicator data that matches the query request.

2. The method according to claim 1, characterized in that The first metadata includes a first index item; and the querying of a first data model matching the first metadata includes: Determining, based on the first indicator item, a first business to which the first metadata belongs; A data model of the first service access is queried, and the data model is determined to be a first data model that matches the first metadata.

3. The method according to claim 1, characterized in that The parsing of the business log generated by the data source based on the first data model to obtain the first business data includes: determining a first filtering condition based on the first metadata; If the first data model includes the first filtering condition, filtering the business log generated by the data source based on the first data model to obtain a first business log, and parsing the first business log to obtain first business data; If the first data model does not include the first filtering condition, the first data model is changed based on the first filtering condition to obtain a second data model, and the business log generated by the data source is filtered based on the second data model to obtain a second business log, and the second business log is parsed to obtain the first business data.

4. The method according to claim 1, wherein After querying the first data model that matches the first metadata, the method further includes: In response to a failure in querying the first data model, creating a third data model based on a model template matching the first metadata; The business log generated by the data source is parsed based on the third data model to obtain the first business data.

5. The method according to claim 1, wherein The performing online analysis and processing on the first business data to obtain second indicator data matching the query request includes: Creating a fourth data model for analyzing the first business data based on the first metadata; The first business data is analyzed and processed online based on the fourth data model to obtain second indicator data that matches the query request.

6. The method according to claim 5, characterized in that The first metadata includes a first indicator item and a first dimension; The performing online analysis and processing on the first business data based on the fourth data model to obtain second indicator data matching the query request includes: Summarize the first business data based on the fourth data model to obtain a first data table; Based on the first data table, counting the index value of the first index item in the first dimension; The indicator value is determined as the second indicator data.

7. The method according to claim 6, characterized in that After determining the indicator value as the second indicator data, the method further includes: Determining a first field that matches the first index item from the fields included in the first data table, and binding the first field to the first index item; A second field matching the first dimension is determined from the fields included in the first data table, and the second field is bound to the first dimension.

8. A data processing device, characterized in that: include: A first query module, configured to query first indicator data matching the query request in response to the query request; a first determining module, configured to determine, in response to a query failure, first metadata for describing the query request based on the query request; A second query module, configured to query a first data model that matches the first metadata; an acquisition module, configured to parse the business log generated by the data source based on the first data model to obtain first business data; An analysis module is used to perform online analysis and processing on the first business data based on the first metadata to obtain second indicator data that matches the query request.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the data processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the data processing method according to any one of claims 1 to 7.

11. A computer program product, characterized in that The computer program product includes a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to execute part or all of the steps of the data processing method according to any one of claims 1 to 7.