Self-service BI integration method based on index medium station and data medium station system

By building a self-service BI integration method for the indicator middle platform, the problems of weak data integration capabilities, low query efficiency and inconsistency of indicators are solved, efficient data integration and rapid query are achieved, and the efficiency and reliability of data analysis are improved.

CN120277137APending Publication Date: 2025-07-08BEIJING BAIJU YIXING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510319890.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing BI tools have shortcomings in terms of weak data integration capabilities, inefficient query efficiency and inconsistency in data indicators, resulting in inefficient data analysis and poor reliability.

Method used

Build a self-service BI integration method based on the indicator middle platform, and generate standardized data indicators through unified data source configuration, access, cleaning, conversion and aggregation, and provide a flexible data service configuration interface to achieve efficient data integration and rapid query.

Benefits of technology

It improves data integration efficiency, enhances data service quality, improves data query response speed and indicator consistency, and ensures the continuity and reliability of data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277137A_ABST
    Figure CN120277137A_ABST
Patent Text Reader

Abstract

The invention discloses a self-service BI integration method based on an index middle station and a data middle station system. The invention relates to the technical field of BI visualization tools. Receiving data source configuration information submitted by a user through an interface or an API (Application Program Interface), wherein the data source configuration information comprises a data source type, connection parameters and a data format; establishing connection with a data source according to the configuration information; and accessing data from the data source, and performing data extraction according to a preset rule and a data format. Through data source configuration and access in the step S1, the system can quickly receive and connect various types of data sources, data is efficiently extracted according to the preset rule, and the complexity and time cost of data integration are reduced. According to the data integration and processing in the step S2, the data is cleaned, converted and aggregated by using the strong capability of a data index middle table, and a standardized data index is generated, so that the efficiency and accuracy of data integration are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of BI visualization tools, specifically to a self-service BI solution based on an indicator middle platform, and particularly to a self-service BI integration method and a data middle platform system based on an indicator middle platform. Background Art

[0002] Traditional BI visualization tools in the field of data analysis focus on the construction of charts, the fine configuration of chart attributes, and the efficient management of reports and data portals. As Figure 2 shown, these tools enable non-technical business users to participate in the data visualization process by providing a user-friendly interface, thus promoting data-driven decision-making. However, in terms of the data indicators that support these front-end visualization components, the current design of BI tools still has certain limitations. Specifically, users need to first configure the data source, and then select appropriate data processing means (such as SQL queries, Python scripts, or R language) according to their own business needs to create data sets and data indicators. Although this process is flexible, it also exposes the deficiency of the BI system in "data integration" ability. These include:

[0003] (1) Weak data integration ability: BI tools on the current market generally focus on building the ability of "data visualization", but are inadequate in the equally crucial ability of "data integration". This imbalance makes it difficult for BI systems to provide a general data integration solution in the face of complex and changing data environments. Most BI tools only provide the basic ability to deploy data integration logic and directly access the original database tables through a query engine. This approach ignores the complexity and diversity of business data integration and lacks comprehensive support for different data sources, different data formats, and different data processing requirements. In practical applications, data acquisition and data integration often occupy most of the data analysis process (up to 80%), which seriously reduces the efficiency and timeliness of data analysis.

[0004] (2) Low query efficiency: In the real-time connection mode, each query triggers data acquisition and processing logic, which is particularly disadvantageous for large amounts of data and complex processing logic. This not only results in slow report configuration and query speeds, but also places extremely high demands on database performance and consumes a large amount of computing resources. Although the data loading mode improves query efficiency by executing logic at regular intervals and storing the results in memory in advance, when the data volume is huge and the computational complexity is high, the data loading process is prone to failure, resulting in users being unable to query the required data, seriously affecting the continuity and reliability of data analysis.

[0005] (3) Inconsistency of data indicators: Due to the lack of a universal data integration solution, the data integration process often relies on user-written code and display specifications. This decentralized management method easily leads to confusion in the naming and definition of data indicators, such as "same name but different meaning" and "same meaning but different name", which seriously affects the consistency and accuracy of the data.

[0006] To this end, the present invention proposes a self-service BI integration method and a data middle platform system based on an indicator middle platform. Summary of the invention

[0007] In view of this, the present invention hopes to provide a self-service BI integration method based on the indicator middle platform and a data middle platform system to solve or alleviate the technical problems existing in the prior art, namely, weak data integration capability, low query efficiency and data indicator inconsistency, and at least provide a beneficial option for this; the technical solution of the present invention is implemented as follows:

[0008] First, the self-service BI integration method based on the indicator platform:

[0009] 1. Overview:

[0010] The present invention aims to achieve comprehensive data integration, standardized processing and flexible services by building an efficient data indicator middle platform. First, the solution receives the data source information configured by the user, establishes a connection and accesses the data from the source, and extracts it according to preset rules. Then, the data is deeply integrated in the middle platform, including cleaning, conversion, aggregation, and generating standardized data indicators. These indicators are uniformly managed and stored in the indicator library. Subsequently, the solution provides an intuitive data service configuration interface, allowing users to customize service parameters as needed, and automatically generate service interfaces for use by applications such as query engines. When a user initiates a query request, the system responds quickly, calls the corresponding interface to obtain data from the indicator library and returns it. In addition, the solution also has a loop and monitoring mechanism to ensure that the data is updated in a timely manner when the data source changes, and quickly responds to new query requirements to achieve continuous and efficient use of data.

[0011] (II) Technical solution:

[0012] In order to achieve the above technical objectives, the present invention selects to execute the following operation steps.

[0013] 2.1 Step S1, data source configuration and access:

[0014] Receive data source configuration information submitted by the user through the interface or API, including data source type, connection parameters, and data format; establish a connection with the data source based on the configuration information; access data from the data source and extract data based on preset rules and data formats.

[0015] 2.1.1 Step S100, Receive data source configuration information:

[0016] The system receives the data source configuration information submitted by the user through the user interface (UI) or application programming interface (API). It includes:

[0017] S1000, Data source type, including database, file system, or / and API.

[0018] S1001, Connection parameters, including database address, port number, username, password, file path, URL of the API, or / and authentication information.

[0019] S1002, Transformed data format, including CSV, JSON, XML, or / and SQL query statements.

[0020] 2.1.2 Step S101, Verify data source configuration information:

[0021] The system verifies the received data source configuration information to ensure its correctness and integrity. It includes:

[0022] S1010, Whether the data source type is supported.

[0023] S1011, Whether the connection parameters are valid, including whether the database connection is successful, whether the file path exists, and whether the API is accessible.

[0024] S1012, Whether the data format is consistent with the expectation.

[0025] The system adopts an automated verification mechanism to quickly feedback the verification result. If the configuration information is incorrect, it prompts the user to make modifications.

[0026] 2.1.3 Step S102, Establish a connection with the data source:

[0027] Based on the verified data source configuration information, the system establishes a connection with the data source. It includes:

[0028] S1020, For database data sources, the system uses the provided connection parameters to establish a database connection.

[0029] S1021, For file system data sources, the system accesses the specified file path.

[0030] S1022, For API data sources, the system initiates an HTTP request to establish a connection.

[0031] S1023, If the connection fails, the system records an error log and prompts the user to check the configuration information or the data source status.

[0032] 2.1.4 Step S103, Access Data:

[0033] The system accesses data from the data source through the established connection. It includes:

[0034] Execute SQL queries, read file contents, and / or parse API responses according to the data source type and data format.

[0035] 2.1.5 Step S104, Data Extraction and Formatting:

[0036] The system extracts and formats the accessed data according to the preset rules and data format. It includes:

[0037] S1040, The system extracts the required data fields according to the definition of data metrics.

[0038] S1041, The system converts the data into a unified format, including converting the date format in the database to the standard date and time format).

[0039] 2.1.6 Step S105, Store or Transmit the Extracted Data:

[0040] The system stores the extracted and formatted data in the specified location or transmits it to the next processing link. It includes:

[0041] S1050, The data can be stored in a database, file system, memory cache, or / and location.

[0042] S1051, The data can be directly transmitted to the data integration and processing module (Step S2) for further processing.

[0043] 2.2 Step S2, Data Integration and Processing:

[0044] The accessed data is sent to the data metrics middleware for integration processing. The middleware cleans, transforms, and aggregates the data according to the preset data integration rules. During the integration process, standardized data metrics are generated and the data metrics are uniformly managed according to the definition and production rules of the data metrics, including the naming, definition, classification, and version control of the metrics, and are stored in the data metrics library.

[0045] 2.2.1 Step S200, Data Reception and Preliminary Check:

[0046] The system receives the data transmitted from the data source configuration and access module (Step S1) and conducts a preliminary data quality check. It includes:

[0047] S2000, Check the integrity of the data to ensure that there is no lost or truncated data.

[0048] S2001, Check the data consistency and confirm that the data format conforms to the expectation.

[0049] S2002, Record the data reception time and source to provide a basis for data traceability.

[0050] 2.2.2 Step S201, Data cleaning:

[0051] Clean the received data according to the preset cleaning rules to remove or correct incorrect, abnormal or unnecessary data. Include:

[0052] S2010, Remove null values, duplicate values or invalid values.

[0053] S2011, Correct data format errors, including date format or / and numerical format.

[0054] S2012, Process outliers, including by replacement, interpolation or / and deletion methods.

[0055] S2013, Filter and screen the data according to business requirements.

[0056] 2.2.3 Step S202, Data conversion:

[0057] Convert the cleaned data into a format or structure suitable for subsequent processing and analysis. Include:

[0058] S2020, Data type conversion, including string to date or / and integer to float.

[0059] S2021, Data structure conversion, including converting nested JSON to a flat table structure.

[0060] S2022, Data splitting and merging, including splitting one field into multiple fields or merging multiple fields into one field.

[0061] S2023, Apply business rules, such as calculating the value of a new field according to specific conditions.

[0062] 2.2.4 Step S203, Data aggregation:

[0063] Perform data aggregation operations according to business requirements and the definition of data metrics, including summation, averaging or / and counting. Include:

[0064] S2030, Determine the aggregation dimensions and metrics, including aggregating sales by time dimension.

[0065] S2031, Apply aggregation functions to calculate the aggregation results.

[0066] S2032, Handle exceptions during the aggregation process, including null value handling or / and division-by-zero protection.

[0067] 2.2.5 Step S204, Data metric generation and management:

[0068] Generate standardized data metrics according to the definitions and production rules of data metrics, and uniformly manage the data metrics. Include:

[0069] S2040, Define the naming rules, data types or / and calculation logics of data metrics.

[0070] S2041, Generate data metrics and assign unique identifiers to them.

[0071] S2042, Store the generated data metrics in the data metric library for subsequent use.

[0072] 2.3 Step S3, Data service configuration and management:

[0073] Provide a data service configuration interface for users or query engines to configure the required data service parameters through the interface according to business requirements. Generate data service interfaces based on the configuration parameters and provide them to the query engine or other data applications for use.

[0074] 2.4 Step S4, Data query and response:

[0075] When a user or query engine initiates a data query request, receive the request and parse the query parameters. At the same time, based on the data service configuration, call the corresponding data service interface, obtain the required data from the data metric library, and then return it to the user or query engine.

[0076] 2.4.1 Step S400, Receive query request:

[0077] Receive the data query request initiated by the user or query engine. Based on HTTP requests, WebSocket or API calls, record metadata such as the timestamp and source IP of the request for subsequent tracking and analysis.

[0078] 2.4.2 Step S401, Parse query parameters

[0079] Parse according to the format of the query request (such as JSON, XML, query string, etc.);

[0080] Check the parameter types, ranges and required fields;

[0081] Convert the parsed query parameters into a format recognizable by the system internally;

[0082] 2.4.3 Step S402, Query data service configuration:

[0083] Access the data service configuration storage to find the data service configuration that matches the query parameters;

[0084] Determine the data service interface, data source, and / or data metrics to be invoked according to the configuration;

[0085] 2.4.4 Step S403, Invoke the data service interface:

[0086] Construct the request parameters for the data service interface, including query conditions, paging information, and / or sorting rules, etc.;

[0087] Send a request to the data service interface and wait for a response;

[0088] Process the interface response, including checking the response status and parsing the response data;

[0089] 2.4.5 Step S404, Obtain data from the data metrics library

[0090] Access the data metrics library according to the data indication in the interface response;

[0091] Execute a data retrieval operation, such as obtaining data according to the metric ID, time range, and filtering conditions;

[0092] Perform necessary processing on the obtained data, such as format conversion and data merging;

[0093] 2.4.6 Step S405, Assemble the query result

[0094] Sort, page, and aggregate the data according to the requirements of the query request; add the query time consumption and data source;

[0095] 2.4.7 Step S406, Return the query result

[0096] Send the query result to the specified receiving address of the user or the query engine; record the sending status of the query result for subsequent tracking and monitoring.

[0097] 2.5 Step S5, Loop and monitor:

[0098] When the data source changes, repeat Step S2; when a new data query request is received, repeat Step S4.

[0099] (III) Mechanism for solving technical problems:

[0100] 3.1 Mechanism and principle for solving the problem of weak data integration ability:

[0101] Through unified configuration and access interfaces, the system can flexibly handle data sources of different types and sources, breaking down data silos and achieving preliminary data integration. Through preset data integration rules and data indicator definitions, the system can deeply process and standardize data to ensure data consistency and availability, thereby enhancing data integration capabilities.

[0102] Through flexible data service configuration, the system can provide users with customized data query services, avoiding the tedious configuration and debugging process in traditional query methods and improving query efficiency. The system adopts an efficient query processing and response mechanism, which can quickly process query requests and return results, and at the same time use data caching, distributed computing and other technical means to further improve query efficiency.

[0103] 3.2 Mechanism and principle for solving the problem of data indicator inconsistency:

[0104] Through unified indicator definitions and production rules, the system can ensure the consistency and accuracy of data indicators, avoiding the problem of inconsistent data indicators between different departments or systems. The system names, defines, classifies and version controls data indicators to ensure the uniqueness and traceability of indicators. The system can clearly record and track changes in data indicators, providing a reliable basis for data analysis and use.

[0105] Second, the data middle platform system based on the self-service BI integration of the indicator middle platform:

[0106] like Figure 3 As shown, the system is used to implement the self-service BI integration method based on the indicator middle platform mentioned above, which includes:

[0107] (1) Data source configuration module: Receives data source configuration information submitted by users through the interface or API, including data source type, connection parameters, and data format. Based on the configuration information, the system establishes a connection with the data source and accesses data from the data source. This process extracts data based on preset rules and data formats to ensure data accuracy and completeness.

[0108] (2) Data indicator middle platform: Clean, transform and aggregate data according to preset data integration rules. During the integration process, standardized data indicators are generated according to the definition and production rules of data indicators, and unified management is carried out, including the naming, definition, classification and version control of indicators, and finally stored in the data indicator library.

[0109] (3) Data service configuration module: According to business requirements, it provides a data service configuration interface through which users or query engines can configure the required data service parameters. The system generates data service interfaces based on the configured parameters and provides them to the query engine or other data applications, ensuring the flexibility and scalability of data services.

[0110] (4) Data query response module: When a user or query engine initiates a data query request, the system receives and parses the query parameters. Based on the data service configuration, the system calls the corresponding data service interface, retrieves the required data from the data metric library, and returns it to the user or query engine, achieving fast response and efficient utilization of data.

[0111] (5) Monitoring module: When the data source changes or a new data query request is received, the system can automatically trigger the corresponding processing flow. When the data source changes, the system repeats the data integration and processing steps (S2); when a new data query request is received, it repeats the data query and response steps (S4). Meanwhile, the system has a monitoring function to ensure the stability and reliability of the data processing and analysis process.

[0112] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0113] First, improve data integration efficiency: Through the data source configuration and access in step S1 of the present invention, the system can quickly receive and connect various types of data sources, efficiently extract data according to preset rules, reducing the complexity and time cost of data integration. In step S2 of data integration and processing, by utilizing the powerful capabilities of the data metric middle platform, the data is cleaned, transformed, and aggregated to generate standardized data metrics, further enhancing the efficiency and accuracy of data integration.

[0114] Second, enhance data service quality: In step S3 of data service configuration and management of the present invention, an intuitive and easy-to-use data service configuration interface is provided for users, enabling them to flexibly configure data service parameters according to business requirements and generate data service interfaces that meet the requirements. Through standardized data metrics and unified management, the consistency and reliability of data services are ensured, improving the quality of data services.

[0115] Third, improve data query response speed: In step S4 of data query and response of the present invention, the system can quickly receive and parse query requests, call the corresponding data service interfaces based on the data service configuration, quickly retrieve the required data from the data metric library, and return it to the user or query engine, significantly improving the response speed of data queries. Description of the Drawings

[0116] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0117] Figure 1 It is a schematic flowchart of the method of the present invention;

[0118] Figure 2 It is a schematic flowchart of the execution process of the traditional BI solution;

[0119] Figure 3 It is a schematic diagram of the system composition of the present invention. Specific Embodiments

[0120] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will give a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings. Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below;

[0121] It should be noted that the various embodiments in this specification are described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0122] Explanation of related terms:

[0123] (1) Data source type: Refers to the type or nature of the data source, such as database, file, API, etc.

[0124] (2) Connection parameters: Information required to establish a connection with the data source, such as address, port, username, password, etc.

[0125] (3) Data format: The format adopted by the data during storage or transmission, such as JSON, CSV, XML, etc.

[0126] (4) Data source: The original source or system that provides data, such as a database server, a file storage location, or a third-party data service.

[0127] (5) Preset rules and data format: A series of rules set in advance for processing data and the format standard that the data should follow.

[0128] (6) Integration processing: Processing such as merging, cleaning, and transforming data from different data sources to form a unified and usable dataset.

[0129] (7) Cleaning, transformation, and aggregation: Specific steps of data integration processing. Cleaning is to remove invalid or incorrect data, transformation is to change the data format or type, and aggregation is to merge multiple pieces of data into one.

[0130] (8) Definition and production rules of data metrics: Rules for defining, calculating, and managing data metrics to ensure the consistency and accuracy of data metrics.

[0131] (9) Standardized data metrics: Data metrics defined and calculated according to a unified standard, which is convenient for comparison and analysis.

[0132] (10) Naming, definition, classification, and version control of metrics: A series of measures for naming data metrics, clarifying their meanings, classifying and managing them, and controlling their version updates.

[0133] (11) Configuration parameters: Parameters related to data services set by users or system administrators according to requirements, such as query conditions, data return formats, etc.

[0134] (12) Query engine: A software component used to execute data query requests, which can parse query statements and retrieve data from data sources.

[0135] (13) Data service interface: An entry or channel for providing access to data services, allowing users or application programs to obtain data services through this interface.

[0136] Example 1: This example discloses the application of the self-service BI integration method based on the metric middle platform in the online car-hailing platform.

[0137] In the online car-hailing business, it is necessary to collect, integrate, and analyze data from multiple data sources (such as vehicle GPS data, passenger order data, payment data, etc.) to support business decision-making and operation optimization. The method for implementing this set of solutions is as follows, please refer to Figure 1 :

[0138] In this example, regarding step S1: Data source configuration and access:

[0139] Specifically, S100: Receive data source configuration information: Users submit data source information through the UI interface or API:

[0140] (1) Data source types: Databases (such as MySQL), file systems (such as CSV files), APIs (such as third-party payment APIs).

[0141] (2) Connection parameters: database address (such as 192.168.1.100), port number (such as

[0142] 3306), username (such as root), password (such as password123); file path (such as

[0143] / data / gps_data.csv); URL of the API (such as

[0144] https: / / api.payment.com / transactions) and authentication information (such as API Key).

[0145] (3) Data formats: CSV, JSON, XML or SQL query statements.

[0146] Specifically, S101: Verify the data source configuration information: whether the data source type is supported (check whether the system supports MySQL, CSV, etc.);

[0147] whether the connection parameters are valid (attempt to connect to the database, access the file path, call the API);

[0148] whether the data format is consistent with the expectation (check the file content, API response format);

[0149] Specifically, S102: Establish a connection with the data source: According to the verified information: For the database, establish a connection using the connection parameters. For the file system, access the specified path. For the API, initiate an HTTP request to establish a connection.

[0150] Specifically, S103: Access the data: Execute an SQL query to obtain database data. Read the file content to obtain file system data. Parse the API response to obtain API data.

[0151] Specifically, S104: Data extraction and formatting: Extract the required fields (such as vehicle ID, longitude, latitude, order ID, payment amount). Convert the data into a unified format (such as convert all dates to the YYYY-MM-DD HH:MM:SS format).

[0152] Specifically, S105: Store or transmit the extracted data: Store the data in the database or file system. Or directly transmit it to the data integration and processing module.

[0153] In this embodiment, regarding step S2: Data integration and processing:

[0154] Specifically, S200: Data reception and preliminary check: Receive the data transmitted from S1, check the integrity and consistency of the data. Record the data reception time and source.

[0155] Specifically, S201: Data cleaning: Remove null values and duplicate values (such as duplicate order records). Correct format errors (such as correcting the wrong date format 2023 / 01 / 01 to 2023-01-01). Process outliers (such as marking or deleting abnormally large payment amounts). Filter data according to business requirements (such as only retaining order data for the most recent month).

[0156] Specifically, S202: Data transformation: Convert date strings to date types. Flatten nested JSON data into a flat table structure. Split or merge fields (such as splitting the address field into province, city, and district fields). Calculate new fields (such as calculating the driving distance based on longitude and latitude).

[0157] Specifically, S203: Data aggregation: Aggregate the order quantity, payment amount, etc. by time dimension. Calculate the average driving distance, average payment amount, etc.

[0158] Specifically, S204: Data metric generation and management: Define metric naming rules (such as order_count_daily represents the daily order quantity). Generate data metrics and assign unique identifiers. Store the metrics in a data metric library for subsequent queries.

[0159] Through the steps of S200 - S203 in this embodiment, the system can automatically clean, transform, and aggregate data, solving the problem of weak data integration ability. The principle is to perform standardization and normalization processing on the data through preset rules and algorithms, enabling data from different sources and in different formats to be unified and integrated together.

[0160] By generating standardized data metrics through S204 and storing them in the data metric library, the query efficiency is improved. The principle is that the metrics in the data metric library have been preprocessed and optimized, so there is no need to clean and transform the data again during querying, and the required data can be directly obtained. Through S204, unified management of data metrics is carried out, including naming, definition, classification, and version control, solving the problem of data metric inconsistency. The principle is that all data metrics follow unified naming rules and definition standards, ensuring the consistency and comparability of data metrics.

[0161] In this embodiment, regarding step S3: Data service configuration and management: Provide an interface for users to configure the required data service parameters. Generate data service interfaces according to the configuration parameters. Provide the interfaces for the query engine or other data applications to use.

[0162] In this embodiment, regarding step S4: Data query and response:

[0163] Specifically, S400: Receive a query request: Receive a query request from a user or a query engine. Record the request timestamp and the source IP.

[0164] Specifically, S401: Parse query parameters: Parse the format and content of the query request. Check the parameter types, ranges, and required fields.

[0165] Specifically, S402: Query data service configuration: Find the data service configuration that matches the query parameters. Determine the data service interface and data metrics to be called.

[0166] Specifically, S403: Invoke the data service interface: Construct the request parameters and send a request to the data service interface. Wait for and process the interface response.

[0167] Specifically, S404: Retrieve data from the data metrics library: Access the data metrics library based on the interface response. Retrieve the required data and perform necessary processing.

[0168] Specifically, S405: Assemble the query result: Sort, paginate, and aggregate the data. Add metadata such as query elapsed time and data source.

[0169] Specifically, S406: Return the query result: Send the query result to the user or the query engine. Record the sending status of the query result.

[0170] In this embodiment, regarding step S5: Loop and monitoring: When the data source changes (such as adding a new data source, data format change), repeat step S2 for data integration and processing. When a new data query request is received, repeat step S4 for data query and response. Monitor and log the entire process to promptly detect and handle abnormal situations.

[0171] It can be understood that through the automated data integration and processing process, manual intervention and errors are reduced, and the efficiency and quality of data integration are improved. Through the standardized data metrics library and optimized query interface, the response speed and accuracy of data query are improved. Through unified data metrics management and version control, the consistency and comparability of data metrics are ensured, providing reliable data support for business decisions.

[0172] Embodiment Two: This embodiment will further provide the specific Python execution program of Embodiment One:

[0173]

[0174]

[0175]

[0176]

[0177] In the above program, the receive_config function is used to simulate receiving the data source configuration information submitted by the user, and the validate_config function validates the received configuration information to ensure that the data source type is supported, the connection parameters are valid, and the data format meets the expectations. The establish_connection function establishes a connection to the data source. For data sources of the database type, the pymysql library is used to establish the connection. The fetch_data function performs corresponding data access operations according to the data source type. For databases, SQL query statements are executed to obtain data and converted into a Pandas DataFrame.

[0178] The extract_and_format_data function extracts and formats the accessed data, such as extracting required fields, converting data types, etc. The store_or_transmit_data function stores the formatted data into a CSV file at a specified path or transmits it to the data integration and processing module for subsequent processing.

[0179] All of the above embodiments only represent the implementation manners of the relevant actual applications of the present invention. The descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

[0180] For those skilled in the art, it can be further realized that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered as exceeding the scope of the present invention.

[0181] Meanwhile, those skilled in the art can understand that all or part of the processes in the methods of implementing the above-mentioned all embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned various methods. Among them, any reference to a memory, storage, database, or other medium provided in the present application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

Claims

1. A self-service BI integration method based on an indicator middle platform, characterized in that It includes the following execution steps: S1. Receive the data source configuration information submitted by the user through the interface or API; establish a connection with the data source according to the configuration information, and perform data extraction according to the preset rules and data formats; S2. The accessed data is sent to the data metrics middleware for integration processing, and standardized data metrics are generated according to the definitions and production rules of the data metrics; S3. Provide a data service configuration interface for the user or query engine that can configure the required data service parameters through the interface; generate a data service interface according to the configuration parameters; S4. When the user or query engine initiates a data query request, receive the request and parse the query parameters. At the same time, based on the data service configuration, call the corresponding data service interface, obtain the required data from the data metrics library, and then return it to the user or query engine.

2. The integration method according to claim 1, wherein: The execution process of S1 includes: S100. The system receives the data source configuration information submitted by the user through the user interface or application programming interface; S101. Verify the received data source configuration information; S102. According to the verified data source configuration information, the system establishes a connection with the data source; S103. Access data from the data source; S104. According to the preset rules and data formats, extract and format the accessed data; S105. Store the extracted and formatted data in a specified location, or transmit it to the next processing link.

3. The integration method according to claim 1, characterized in that: The execution process of S100 includes: S1000. Data source types, including databases, file systems, or / and APIs; S1001. Connection parameters, including database addresses, port numbers, usernames, passwords, file paths, API URLs, or / and authentication information; S1002. Convert data formats, including CSV, JSON, XML, or / and SQL query statements; The execution process of S101 includes: S1010. Whether the data source type is supported; S1011. Whether the connection parameters are valid, including whether the database connection is successful, whether the file path exists, and whether the API is accessible.

4. The integration method according to claim 3, wherein: The execution process of S102 includes: S1020. For database data sources, establish a database connection using the provided connection parameters; S1021. For file system data sources, access the specified file path; S1022. For API data sources, initiate an HTTP request to establish a connection; S1023. If the connection fails, record the error log and prompt the user to check the configuration information or the data source status.

5. The integration method according to claim 1, wherein: S2 cleans the received data according to the preset cleaning rules, removing or correcting errors, anomalies, or unnecessary data, including: S2010. Remove null values, duplicate values, or invalid values; S2011. Correct data format errors, including date formats or / and numerical formats; S2012. Process outliers, including by replacement, interpolation, or / and deletion methods; S2013. Filter and screen the data according to business requirements; Convert the cleaned data into a format or structure suitable for subsequent processing and analysis; including: S2020, Data type conversion, including string to date and / or integer to floating point number; S2021, Data structure conversion, including converting nested JSON to a flat table structure; S2022, Data splitting and merging, including splitting one field into multiple fields or merging multiple fields into one field.

6. The integration method according to claim 5, wherein: The aggregation method of S2 includes: S2030, Determine the dimensions and metrics for aggregation, including aggregating sales by time dimension; S2031, Apply aggregation functions to calculate the aggregation results; S2032, Handle exceptions during the aggregation process, including null value handling and / or division by zero protection; Generate standardized data metrics according to the definitions and production rules of data metrics, including: S2040, Define the naming rules, data types and / or calculation logics of data metrics; S2041, Generate data metrics and assign unique identifiers to them; S2042, Store the generated data metrics in the data metric library.

7. The integration method according to claim 1, wherein: The execution process of S4 includes: S400, Receive a data query request initiated by a user or a query engine; S401, Parse according to the format of the query request; check the parameter types, ranges and required fields; convert the parsed query parameters into a format recognizable by the system internally; S402, Access the data service configuration storage to find the data service configuration that matches the query parameters; S403, Construct the request parameters for the data service interface, including query conditions, pagination information and / or sorting rules; S404, According to the data indication in the interface response, access the data metric library; perform data retrieval operations, such as obtaining data according to the metric ID, time range and filtering conditions; S405, Sort, paginate and aggregate the data according to the requirements of the query request; add the query time consumption and data sources; S406, Send the query results to the specified receiving address of the user or the query engine; record the sending status of the query results for subsequent tracking and monitoring.

8. The integration method according to any one of claims 1 to 7, characterized in that: It also includes S5. When the data source changes, repeat step S2; when a new data query request is received, repeat step S4.

9. A data middle platform system for implementing the integration method according to any one of claims 1 to 8, characterized in that The system includes: Data source configuration module: Receive the data source configuration information submitted by the user through the interface or API; Data metric middle platform: Perform integration processing according to the preset data integration rules; Data service configuration module; Provide a data service configuration interface according to business requirements; Data query response module: When the user or the query engine initiates a data query request, the system receives and parses the query parameters.

10. The system according to claim 9, wherein: According to the configuration information, the data source configuration module establishes a connection with the data source and accesses data from the data source; The data metric middle platform cleans, converts and aggregates the data according to the preset data integration rules; during the integration process, standardized data metrics are generated according to the definitions and production rules of data metrics.