Cloud customer service system report generation method and device based on Apache Dores, equipment, medium and product

By leveraging Apache Doris's data catalog functionality and visual interface, the cloud customer service system's reports can be configured quickly, solving the problem of long development cycles in traditional reports and improving query efficiency.

CN121960416APending Publication Date: 2026-05-01BEIJING WISDOM TOOTH TECH CONSULTING CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING WISDOM TOOTH TECH CONSULTING CO LTD
Filing Date
2025-12-25
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional cloud customer service systems have long report development response cycles, which cannot meet the high-frequency demand for real-time strategy adjustments, and traditional databases are slow when performing large-scale data statistics.

Method used

Metadata is obtained through Apache Doris's data catalog feature, and a visual interface is used to output heterogeneous data sources, statistical indicator fields, dimension fields, and a list of Doris functions. Business users can independently select and configure reports, reducing reliance on technical personnel.

Benefits of technology

It shortens the report development response cycle and improves report query efficiency, allowing business personnel to complete report configuration without writing complex SQL statements or configuring data extraction scripts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121960416A_ABST
    Figure CN121960416A_ABST
Patent Text Reader

Abstract

The invention discloses a cloud customer service system report generation method and device based on Apache Dores, equipment, a medium and a product, and relates to the technical field of big data process.The method comprises the steps that a selectable heterogeneous data source list corresponding to a cloud customer service system is output through a visual interface; after receiving a target data source selection instruction submitted by a user, accessing metadata of a target data source by using a data directory function of Apache Dores; outputting a selectable statistical index field list, a selectable statistical dimension field list and a selectable Dores function list through a visual interface, and processing data of a target data source by calling a target Dores function after receiving a target statistical index field, a target statistical dimension field and a target Dores function corresponding to a statistical report selected by a user. Obtaining statistical data corresponding to the statistical report; and outputting the statistical data through a visual interface. The report development response period can be shortened, and the report query efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data processing technology, and in particular to a method, apparatus, device, medium and product for generating reports for a cloud customer service system based on Apache Doris. Background Technology

[0002] As enterprise customer service levels continue to improve, cloud-based customer service systems have become an indispensable tool for businesses. These systems generate massive amounts of interactive data (such as work orders, conversations, and satisfaction ratings) in their daily operations. Extracting valuable business insights from this massive amount of interactive data relies on flexible, powerful, and efficient reporting functions.

[0003] In traditional reporting systems, after business personnel submit report requests based on cloud customer service operations needs, technical staff will write complex SQL statements and configure data extraction scripts accordingly. This results in a long report development response cycle and cannot meet the high-frequency needs of real-time strategy adjustments in customer service operations. Traditional databases such as MySQL are relatively slow when handling large amounts of data. Using different types of tables in Doris for different scenarios can significantly improve report query performance. Summary of the Invention

[0004] The purpose of this application is to provide a method, apparatus, device, medium and product for generating reports for a cloud customer service system based on Apache Doris, which can shorten the response cycle of report development and improve the efficiency of report query.

[0005] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a method for generating reports for a cloud customer service system based on Apache Doris, the method comprising: A list of heterogeneous data sources corresponding to selectable cloud customer service systems is displayed through a visual interface; After receiving the target data source selection instruction submitted by the user based on the list of heterogeneous data sources, the metadata of the target data source is obtained using the data catalog function of Apache Doris, and the metadata includes field names; The system outputs a list of selectable statistical indicator fields, a list of selectable statistical dimension fields, and a list of selectable Doris functions through a visual interface. The list of statistical indicators includes the field names in the metadata. After receiving the target statistical indicator field, target statistical dimension field, and target Doris function corresponding to the statistical report selected by the user, the data of the target data source is processed by calling the target Doris function according to the target statistical dimension field and the target statistical indicator field to obtain the statistical data corresponding to the statistical report. The statistical data is output through a visual interface.

[0006] Secondly, this application provides a report generation device for a cloud customer service system based on Apache Doris, the device comprising: The first output module is used to output a list of heterogeneous data sources corresponding to selectable cloud customer service systems through a visual interface. The first acquisition module is used to acquire the metadata of the target data source after receiving the target data source selection instruction submitted by the user based on the list of heterogeneous data sources, using the data catalog function of Apache Doris. The metadata includes field names. The second output module is used to output a list of selectable statistical indicator fields, a list of selectable statistical dimension fields, and a list of selectable Doris functions through a visual interface. The list of statistical indicators includes the field names in the metadata. The second acquisition module is used to, after receiving the target statistical indicator field, target statistical dimension field, and target Doris function corresponding to the statistical report selected by the user, process the data of the target data source by calling the target Doris function according to the target statistical dimension field and the target statistical indicator field to obtain the statistical data corresponding to the statistical report. The third output module is used to output the statistical data through a visual interface.

[0007] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the Apache Doris-based cloud customer service system report generation method described above.

[0008] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the Apache Doris-based cloud customer service system report generation method described above.

[0009] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the Apache Doris-based cloud customer service system report generation method described above.

[0010] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a method, apparatus, device, medium, and product for generating reports for a cloud customer service system based on Apache Doris. It outputs a list of selectable heterogeneous data sources, a list of statistical indicator fields, a list of statistical dimension fields, and a list of Doris functions through a visual interface. Business personnel (such as cloud customer service operators) do not need to rely on technical personnel to write complex SQL statements or configure data extraction scripts. They can complete the report configuration simply by selecting the target data source, the target statistical indicator / dimension, and the Doris function through a visual operation, thus shortening the report response cycle. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart illustrating a report generation method for a cloud customer service system based on Apache Doris, according to an exemplary embodiment. Figure 2 This is a schematic diagram of a display interface according to an exemplary embodiment. Figure 1 ; Figure 3 This is a schematic diagram of a display interface according to an exemplary embodiment. Figure 2 ; Figure 4 A schematic diagram of the functional modules of a cloud customer service system report generation device based on Apache Doris, provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0014] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0015] Figure 1This is a flowchart illustrating a report generation method for a cloud customer service system based on Apache Doris, according to an exemplary embodiment. Figure 1 As shown, the method includes the following steps S101-S105: In step S101, a list of selectable heterogeneous data sources is output through a visual interface.

[0016] The list of heterogeneous data sources must include at least Hive, Iceberg, and MySQL data sources.

[0017] Hive data sources are based on Hadoop's distributed data warehouse and are used in cloud customer service scenarios to store historical data that does not require real-time access but needs to be retained long-term for retrospective analysis. For example: Historical conversation details for the past 1 / 3 years (used for annual customer service performance review and long-term customer demand trend analysis). Archived work order data (used for auditing historical work order processing efficiency and analyzing the repetition rate of old customer issues). Import customer profile data in batches (for analyzing the consultation preferences of customers with different profiles).

[0018] Iceberg-type data sources are used in cloud customer service scenarios to store detailed data that needs to be updated frequently and has a large volume, such as: Customer inquiry question tag data (tags may be manually corrected later, and single / batch updates must be supported); Cross-channel integrated customer interaction data (interaction records from APP, web, and telephone, which must be combined and support updates); Work order details with status changes (work order status changes from pending → processing → completed, status needs to be updated in real time and historical versions need to be retained).

[0019] MySQL data sources are used in cloud customer service scenarios to store core business data that needs to be updated in real time and accessed frequently, such as: Real-time conversation records between customer service representatives and customers (including conversation ID, customer service representative ID, conversation duration, satisfaction rating, etc.); Unclosed-loop work order data (including work order ID, customer request, processing status, assigned customer service representative, etc.); Basic customer service information (including customer service ID, department, and on-duty status).

[0020] When users need real-time, high-frequency reports (such as the number of orders received today, the current work order status), choose the MySQL data source; when users need massive offline reports (such as annual performance, long-term trends), choose the Hive data source; when users need dynamically updated reports (such as work order status progress, tag correction statistics), choose the Iceberg data source.

[0021] The interface presentation format can be as follows: Figure 2 As shown, the design employs a visual approach combining category navigation and list display. The list below displays selectable data sources, which users can click upon. Figure 2 After selecting a button, the data source corresponding to that button becomes the target data source selected by the user.

[0022] Accessibility design, such as Figure 2 As shown, if there are many data source types, keyword search can also be supported (e.g., if business personnel enter a conversation, data sources containing conversation keywords will be automatically filtered).

[0023] In this disclosure, the external table function of Apache Doris can be used to pre-register various heterogeneous data sources for cloud customer service, without requiring business personnel to manually configure connection parameters (such as database IP, port, account password); it can also perform validity verification on each data source and periodically check the connection status of the data source (such as whether the MySQL service is normal and whether the Hive table exists). If the data source is unavailable, the interface will mark the abnormal status and indicate the reason (such as connection timeout table non-existence).

[0024] In step S102, after receiving the target data source selection instruction submitted by the user based on the heterogeneous data source list, the data catalog function of Apache Doris is invoked to obtain the metadata of the target data source, which includes field names.

[0025] In one embodiment, the metadata also includes the database name, table name, and the storage location of the data corresponding to each field.

[0026] Apache Doris is a high-performance, real-time analytical database based on an MPP architecture, renowned for its speed and ease of use. It can return query results for massive datasets in sub-second response times, supporting both high-concurrency point queries and high-throughput complex analytical scenarios. Based on this, Apache Doris effectively meets the needs of various use cases, including report analysis, ad-hoc queries, unified data warehouse construction, and accelerated federated queries in data lakes. Users can build applications on top of it, such as user behavior analysis, A / B testing platforms, log retrieval and analysis, user profiling analysis, and order analysis.

[0027] Doris is internally layered into three layers: ODS (Detailed Data Layer), which stores the raw detailed data, sourced from the business database; DWS (Data Service Layer), which contains multiple dimension fields for reuse across different reports, with data sourced from the ODS layer; and ADS (Application Layer), which performs the necessary statistical calculations based on the statistical requirements of different reports, grouping the data from the DWS layer by dimension fields to obtain the statistical results. Ultimately, users can view the reports by querying the data in the ADS layer.

[0028] After the business personnel select the target data source (multiple selections are allowed, such as selecting both session data and agent basic data) on the S101 interface and click confirm selection, the system will automatically trigger the following actions: Call the Apache Doris Data Catalog feature, which is equivalent to a structured dictionary of data sources and can collect metadata across heterogeneous data sources; By using pre-configured data source connection information (encrypted and stored, such as MySQL's IP address and Hive's Metastore address), a seamless connection with the target data source can be established without requiring business personnel to manually enter any connection parameters. If multiple data sources are selected, the system automatically detects the relationships between the data sources (such as linking session data and agent basic data through agent ID) and marks the related fields in the subsequent metadata display to help the business understand the logic between the tables.

[0029] The Apache Doris data catalog collects full metadata from the target data source and is specifically optimized for cloud customer service scenarios. The collected metadata includes six key types of information: Basic structure information: database name, table name, field name, field data type (e.g., agent ID is an integer, session duration is a long integer, and work order type is a string); Business Meaning Notes: The system adds business notes to the core fields of cloud customer service in advance (such as duration_sec as the session duration (unit: seconds), first_respond_sec as the first response time (unit: seconds), ticket_status as the ticket status (0=pending, 1=processing, 2=resolved, 3=closed)) to avoid misunderstandings caused by obscure field names by business personnel; Data storage location: Specify the physical storage path of the data corresponding to the field (such as the HDFS path of Hive tables, the database instance name of MySQL tables) to provide a basis for subsequent data processing; Partitioning rule information: If the data source is partitioned according to business rules (such as Hive session tables partitioned by date (event_date) and MySQL work order tables partitioned by work order type), the partitioning field and partitioning range will be collected (such as the partitioning range of the session table being 2024-01-01 to present), providing a basis for partition pruning in subsequent data processing (processing only the target partition data to improve efficiency); Data quality information: Annotate fields for non-null constraints (e.g., agent ID is a non-null field) and data redundancy (e.g., the customer ID field in the session table contains some null values) to help business personnel select reliable fields as statistical dimensions / indicators; Data update time: Collect the latest update time of each data table (e.g., the last update time of the session table: 2024-05-20 14:30:00) to help the business judge the timeliness of the data (e.g., when making real-time reports, prioritize data sources with the latest update time ≤ 5 minutes).

[0030] In one embodiment, the system performs business-specific filtering on the collected metadata, automatically removing invalid information that is irrelevant to the cloud customer service reports: Remove test tables and log tables (such as test_session_table test session table and server_log server log table); Remove redundant fields (such as duplicate creation time fields and primary key ID fields with no business meaning); Remove invalid fields (such as spare fields with all empty values); after filtering, the metadata is displayed in the backend in a tree structure (not visible to business personnel, only used to support subsequent steps), organized in the hierarchy of database → data table → field, and the relationship between each field is marked (such as the agent ID in the session table being associated with the agent ID in the agent base table).

[0031] In step S103, a list of selectable statistical indicator fields, a list of selectable statistical dimension fields, and a list of selectable Doris functions are output through a visual interface. The list of statistical indicators includes the field names in the metadata.

[0032] This step allows business personnel to independently define what the report statistics include (indicators), at what granularity (dimensions), and what technical solutions (Doris functions) to use to process the data.

[0033] (1) The list of statistical indicator fields may include: Original business fields: directly reuse basic fields from metadata, such as work order ID, session ID, customer satisfaction rating, and agent name. Each field can also be labeled with the associated data source and field type (e.g., customer satisfaction rating: from Iceberg - customer feedback data, numeric). Derivative Aggregated Metrics: Combining Apache Doris' built-in aggregation capabilities (summation, counting, averaging, deduplication counting, etc.) or custom functions specific to cloud customer service, we generate computational metrics. Each metric is labeled with its calculation logic, associated basic fields, and applicable scenarios. For example: Conversation-related metrics: Total conversation duration (calculation logic: sum of conversation durations; related field: duration_sec; applicable scenario: agent performance statistics), Average conversation duration (calculation logic: average conversation duration; related field: duration_sec; applicable scenario: agent service efficiency analysis), First response rate (calculation logic: number of conversations with first response time ≤ 120 seconds / total number of conversations; related field: first_respond_sec; applicable scenario: customer service quality monitoring). Work order metrics: Total number of work orders (calculation logic: work order ID count; related field: ticket_id; applicable scenario: work order processing volume statistics), First-time resolution rate of work orders (calculation logic: number of work orders that are closed for the first time without being reopened / total number of work orders; related fields: ticket_status, reopen_flag; applicable scenario: work order processing efficiency analysis), Average resolution time (calculation logic: average of resolution time and creation time; related fields: create_time, solve_time; applicable scenario: work order processing timeliness monitoring). This disclosure also provides a custom indicator entry point: business personnel can independently configure new indicators (such as cross-departmental work order transfer rate). They only need to select basic fields and set calculation logic (such as the number of transferred work orders / total number of work orders), and the system will automatically match the calculation capabilities of Apache Doris without the need for technical development. If multiple indicators are selected, the system will automatically detect the compatibility of the indicators (e.g., detailed fields and aggregate indicators cannot be selected at the same time) and indicate the reason (e.g., detailed fields are suitable for displaying raw data, while aggregate indicators are suitable for statistical analysis and cannot be used in the same report at the same time).

[0034] (2) The list of statistical dimension fields can cover the granularity of statistics for the entire cloud customer service scenario, including filtering fields suitable for grouped statistics from metadata fields, prioritizing high-frequency dimensions of cloud customer service, and classifying them by dimension type to match the statistical habits of business personnel, for example: Time dimension: Includes year, quarter, month, day, hour and other levels (derived from date / time fields in metadata, such as statistical year, statistical month and statistical day derived from event_date), suitable for trend analysis scenarios (such as daily session volume trend, monthly work order resolution rate change); Business dimensions: Agent ID, Agent Name, Department (Agent-related), Work Order Type (Consultation, After-sales, Complaint, etc.), Conversation Channel (APP, Mini Program, Webpage, Telephone) (Business attribute related), suitable for classification and comparison scenarios (such as comparing agent performance across departments, comparing conversation volume across channels); Customer dimensions: Customer ID (agent_id), customer level (normal, VIP), region (customer-related), suitable for customer segmentation analysis scenarios (such as VIP customer satisfaction analysis, analysis of customer consultation hotspots in various regions). Tenant dimension: Tenant ID, Tenant name (for cloud customer service systems in SaaS mode), suitable for multi-tenant data isolation scenarios (such as exclusive reports for different enterprises, only displaying customer service data of this enterprise).

[0035] In one embodiment, dimension association prompts can also be provided. For example, if the agent ID dimension is selected, the system will automatically prompt whether to associate the agent name and department fields (no additional selection is required, the system will automatically associate and display them) to ensure the completeness of the report information.

[0036] (3) The Doris function list is used to demonstrate the three core table models of Apache Doris: Aggregate Model, Unique Key Model, and Duplicate Model. Among them, the Duplicate Key Model allows duplicates of specified key columns, and the Doris storage layer retains all written data, which is suitable for situations where all original data records must be retained; the Unique Key Model ensures that there are no duplicate rows for a given key column, and the Doris storage layer retains only the most recently written data for each key, which is suitable for data updates; the Aggregate Key Model can aggregate data based on the key column, and the Doris storage layer retains the aggregated data, thereby reducing storage space and improving query performance; it is usually used when information needs to be summarized or aggregated (such as totals or averages).

[0037] Once the table is created, its attributes are fixed and cannot be modified. Therefore, choosing the appropriate model for the business requirements is crucial. Duplicate Key: Suitable for ad-hoc queries of any dimension. Although it also cannot take advantage of pre-aggregation features, it is not constrained by the aggregation model and can take advantage of the columnar storage model (only reading relevant columns, without needing to read all key columns).

[0038] Unique Key: For scenarios requiring a unique primary key constraint, it can guarantee the uniqueness of the primary key. However, it cannot take advantage of the query benefits brought by pre-aggregation such as ROLLUP.

[0039] Aggregate Key: By pre-aggregating, it can significantly reduce the amount of data scanned and the computational load of aggregate queries, making it ideal for report-type query scenarios with fixed patterns. However, this model is not suitable for... The query is not user-friendly. Furthermore, because the aggregation method on the Value column is fixed, semantic correctness needs to be considered when performing other types of aggregation queries.

[0040] In Doris, data is stored in columns, and a table can be divided into key columns and value columns. Key columns are used for grouping and sorting, while value columns are used for aggregation. A key column can be one or more fields. When creating a table, data is sorted and stored according to the Aggregate Key, Unique Key, and Duplicate Key columns used in various table models.

[0041] Different table models require specifying a Key column when creating the table, each with different meanings: For the DuplicateKey model, the Key column indicates sorting and has no unique key constraint. In the Aggregate Key and Unique Key models, aggregation is based on the Key column, which has both sorting capabilities and unique key constraints.

[0042] Using sorting keys appropriately can bring the following benefits: Accelerate query performance: Sort keys help reduce data scanning. For range or filtered queries, sort keys can be used to directly locate data positions. For queries that require sorting, sort keys can also be used to speed up the process. Data compression optimization: Storing data in an ordered manner according to the sort key will improve compression efficiency. Similar data will be grouped together, which will greatly improve the compression ratio and reduce the data storage space.

[0043] Reduce deduplication costs: When using a Unique Key table, Doris can perform deduplication more efficiently by using the sort key, ensuring data uniqueness.

[0044] When choosing a sort key, you can follow these suggestions: The key column must precede all value columns.

[0045] Choose integer types whenever possible, as integers are far more efficient for calculations and searches than strings.

[0046] When choosing integer types of different lengths, the principle should be "enough is enough".

[0047] For the length of VARCHAR and STRING types, follow the principle of "sufficient is sufficient".

[0048] In one embodiment, a model recommendation function is also provided to automatically recommend models based on the type of metrics selected by the business (e.g., if aggregated metrics such as total session duration and response rate are selected, an aggregated model is recommended; if raw fields such as session details are selected, a unique model is recommended; if metrics such as work order status that need to be updated in real time are selected, an updated model is recommended).

[0049] like Figure 3 As shown, users select the desired statistical indicators in the "Statistical Indicator Fields List", the desired statistical dimensions in the "Statistical Dimension Fields List", and the desired Doris functions in the "Doris Functions List".

[0050] In step S104, after receiving the target statistical indicator field, target statistical dimension field, and target Doris function corresponding to the statistical report selected by the user, the data from the target data source is processed by calling the target Doris function based on the target statistical dimension field and the target statistical indicator field to obtain the statistical data corresponding to the statistical report.

[0051] After business personnel complete the selection of target statistical indicator fields, target statistical dimension fields, and target Doris functions, the system automatically generates standardized SQL commands that can be directly executed by Apache Doris. The generated SQL commands are sent to the Apache Doris cluster to trigger the corresponding calculation process of the model (completed by the collaboration of Doris FE / BE nodes). After the Doris cluster completes the calculation, it obtains the statistical data.

[0052] In one embodiment, the above steps S103-S104 can be implemented using ChatBI, so that the user only needs to state their needs.

[0053] In step S105, the statistical data is output through a visual interface.

[0054] In this disclosure, a list of selectable heterogeneous data sources, statistical indicator fields, statistical dimension fields, and Doris functions are output through a visual interface. Business personnel (such as cloud customer service operators) do not need to rely on technical personnel to write complex SQL statements or configure data extraction scripts. They can complete the report configuration simply by selecting the target data source, the target statistical indicator / dimension, and the Doris function through a visual operation, thus shortening the report response cycle.

[0055] In one embodiment, outputting the statistical data through a visual interface includes the following sub-steps S1051-S1052: S1051. Through the visual interface, output a list of selectable report display types, including: display shape, display index, and display dimension.

[0056] The display shape is the visual presentation of data, which is the "appearance" of the report. The computer has preset options adapted to cloud customer service reports, such as bar charts (suitable for comparing the number of orders received by different customer service representatives), line charts (suitable for viewing the time trend of the number of orders received), tables (suitable for viewing detailed data), pie charts (suitable for viewing the percentage of satisfaction ratings), radar charts (suitable for comparing customer service performance in multiple dimensions), etc. The displayed metrics are specific statistical metrics that users can choose to display in the report. The options come from the metric fields in the statistical data generated in step S104 (such as the number of orders received, average session duration, satisfaction score, work order resolution rate, etc.), and single / multiple selection is supported. The display dimension is a statistical dimension that users can choose to display in the report. The options come from the statistical dimension fields in step S104 (such as customer service ID, statistical date, tenant identifier, work order type, etc.). It supports adjusting the dimension level (such as displaying by "statistical date + customer service ID" as a dual dimension).

[0057] S1052. After receiving the user's instruction to select a target display type based on the report display type list, the statistical data is displayed through a visual interface according to the target display type.

[0058] Suppose the user selects "Display Shape: Bar Chart, Display Metric: Number of Orders Received, Display Dimension: Customer Service ID + October 2024 Date": Step 1: Extract the three types of parameters selected by the user, and specify that the number of orders received by each customer service representative from October 2024 should be displayed in the form of a bar chart. The X-axis represents the customer service representative ID, the Y-axis represents the number of orders received, and different colored bars are used to distinguish different dates. Step 2: Call the visualization rendering engine to map the structured statistical data generated in step S104 (such as Customer Service 101's order count 25 from 10:01 to 10:01 and Customer Service 102's order count 22 from 10:01 to 10:01) into the visual parameters of the bar chart (X-axis coordinates, Y-axis values, bar colors, legend, etc.). Step 3: Render a bar chart that meets the requirements on the interface, supporting interactive operations (such as hovering the mouse over the bar to display Customer Service 101's order count from October 1st to October 1st as 25, and clicking on the dimension to switch the display level, such as displaying the total number of orders by date only).

[0059] Users can customize the visual appearance of reports, the metrics and dimensions displayed, adapt to different data analysis scenarios, and lower the barrier to report interpretation for non-technical users of cloud customer service.

[0060] This embodiment transforms structured statistical data into a user-customizable visualization format, meeting the personalized needs of cloud customer service users for report display formats and making statistical data easier to interpret.

[0061] In one embodiment, if the target statistical dimension field includes a statistical date, the step of processing the data from the target data source by calling the target Doris function to obtain the statistical data corresponding to the statistical report includes the following sub-steps A1-A2: A1. Using the statistical date as the dynamic partition key in Apache Doris, trigger a partition pruning operation to filter partition data that matches the statistical date range from the data in the target data source.

[0062] By setting the user-selected "statistical date" (such as the event_date field in the metadata) as the basis for dynamic partitioning of the Apache Doris table, Doris will automatically manage partitions based on this field (such as creating new partitions daily and automatically cleaning up historical partitions older than 90 days). Before performing data processing, Doris will first parse the user-selected statistical date range (such as "2024-10-01 to 2024-10-31"), and only scan the partition data that matches this date range, directly skipping irrelevant partitions (such as partitions from 2024-09 and 2024-11). Essentially, it coarsely filters the data at the partition level to avoid full table scans.

[0063] A2. Using the statistical date as a filter condition, perform the processing logic corresponding to the target Doris model on the cropped partition data to obtain the statistical data corresponding to the statistical report.

[0064] Partition pruning is a coarse screening at the partition level (only determining which partitions to scan for specific dates), while filtering by statistical date is a precise fine screening within a partition, ensuring that only data within the partition that meets the user's precise date requirements is processed. This is then combined with the selected Doris model to perform calculations and ultimately produce statistical data.

[0065] For example, users need to generate a "Daily Customer Service Order Count Report from October 1, 2024 to October 7, 2024" (with the statistical date as the dimension and the aggregation model as the objective Doris function): Set event_date (statistical date) as the Doris dynamic partition key to trigger partition pruning, and only filter out the 7 daily partition data from 2024-10-01 to 2024-10-07; Add an event_date filter to precisely target session data for these 7 days; The aggregation model is invoked to execute the logic of "grouping and counting session IDs by statistical date + customer service ID", which ultimately yields the statistical data of "the number of orders received by each customer service representative per day from October 1st to 7th".

[0066] This embodiment utilizes Doris's partitioning feature to adapt to the statistical date dimension, enabling efficient and accurate calculation of cloud customer service report data.

[0067] In one embodiment, if the target statistical dimension field includes a tenant identifier, the step of processing the data from the target data source by calling the target Doris function to obtain the statistical data corresponding to the statistical report includes: using the tenant identifier as a filtering condition, processing the filtered data from the target data source by calling the target Doris function to obtain the statistical data corresponding to the statistical report.

[0068] Cloud customer service systems typically employ a multi-tenant architecture, where different enterprises / vendors (tenants) share a single system. However, their respective customer service data (such as session logs, work orders, and order counts) must be strictly isolated (for example, tenant A's customer service data cannot be accessed by tenant B). The tenant identifier (tenant_id) is the unique core field that distinguishes data from different tenants (e.g., tenant A's identifier is T001, and tenant B's is T002). The target data source (such as MySQL / Hive / Iceberg) stores customer service data from all tenants in a mixed manner. Therefore, the core of this step is to first isolate the data of the current tenant using the tenant identifier as a filtering condition, and then call the Doris function for processing. This ensures data security while meeting reporting and statistical needs.

[0069] If a user actively selects a tenant identifier as a statistical dimension (e.g., "counting order receipts by tenant T001"), then that tenant identifier (e.g., T001) is directly extracted. If the user does not actively select one, but the system is a multi-tenant architecture, the computer automatically extracts the tenant identifier from the current user's login context (e.g., if the currently logged-in user is an operations staff member of tenant T001, T001 is extracted by default). The tenant identifier is then converted into a standardized filtering condition that Apache Doris can recognize: generating an SQL filtering clause WHERE tenant_id='T001' (T001 is the target tenant identifier). This condition is used as a "forced pre-filtering rule" and embedded into the subsequent calculation logic of the Doris function. Specifically, it first filters out the data source data that belongs only to the target tenant, and then calculates according to the Doris function rules, rather than calculating all the data first and then filtering. This minimizes invalid calculations and ensures isolation. After the Doris cluster completes the calculation, the returned statistical data only includes the customer service data of the target tenant (T001), without any data from other tenants mixed in.

[0070] For example, the user (the operations staff of tenant T001) needs to generate a "Customer Service Order Count Report for this Tenant from October 2024" (the statistical dimensions include tenant identifier and customer service ID, and the target Doris function is an aggregation model): Extract tenant identifier T001 from user login context and generate filter condition WHERE tenant_id='T001'; The filter condition is integrated with the aggregation calculation logic of "group by customer service ID and count session IDs" and sent to the Doris cluster; The Doris cluster only scans the session data of tenant T001 from October 2024 to perform pre-aggregation calculations; Returned statistics: Only includes the number of orders received by each customer service representative under tenant T001 in October, with no data from other tenants included.

[0071] In this disclosure, the tenant identifier (such as tenant_id) is used as a mandatory filtering condition and embedded into the full data processing flow of the Apache Doris function to ensure that each tenant can only calculate and obtain customer service data belonging to its own entity, thereby avoiding cross-tenant data leakage.

[0072] In one embodiment, the above method further includes the following steps B1-B4: B1. The upload entry for custom Doris functions is output through the visual interface.

[0073] Add a custom model upload function entry (in the form of a button / pop-up window) to the visual interface area that displays the list of selectable Doris models (such as next to the model list or in the sidebar of the page).

[0074] B2. Receive user-submitted custom Doris function files.

[0075] The system receives user-submitted custom Doris function files via a visual upload interface, temporarily stores the files in the system's temporary storage directory (such as / tmp / doris_custom_model / on the server), and generates a unique file identifier (such as the file's MD5 value plus the upload timestamp).

[0076] It can also call a preset file verification engine to perform multi-dimensional verification on uploaded files: Format validation: Verify that the file extension and encoding format meet the requirements of the Doris cluster (e.g., Jar files must be executable Java bytecode files, and JSON configuration files must be grammatically correct). Version compatibility check: Compare the Doris version dependency information built into the file with the current cluster version (e.g., Doris 2.1.x). If incompatible, return a prompt (e.g., "This model depends on Doris 1.2.x, the current cluster is 2.1.x, it is recommended to upgrade the model and re-upload"). Security verification: Scan files for malicious code and illegal dependency libraries to avoid threats to Doris cluster security.

[0077] If the verification passes, the computer will display "File verification successful, registration will be performed soon"; if it fails, it will return the specific error reason (such as "File format error - only Jar / so / JSON formats are supported") and allow the user to upload again.

[0078] B3. Call the Doris cluster model registration interface to perform the registration operation of the custom Doris function.

[0079] Parse the metadata of the custom model file (such as model name, input and output parameters, and applicable data processing scenarios), assemble it into registration request parameters that the Doris cluster can recognize, and call the Doris FE (front-end node) model registration interface (such as CREATE FUNCTION / REGISTER MODEL REST API). After the Doris cluster executes the registration operation, it returns a success / failure response. If successful: The computer records the registration information (model ID, name, applicable scenario, registration time) to the system configuration library; If it fails: The computer analyzes the reason for the failure (such as "model function entry point is missing" or "BE node file synchronization failed") and provides feedback to the user through a visual interface, supporting retry.

[0080] B4. After registration, the list of selectable Doris functions will be automatically updated, and custom Doris functions will be added to the list as independent options.

[0081] The system reads information about registered custom Doris functions from the system configuration library and automatically updates the data source of the "selectable Doris function list" to ensure that the list data is synchronized in real time.

[0082] By providing a custom Doris function upload entry through a visual interface, it automatically receives files, calls the Doris cluster interface to complete registration and update the model list, allowing non-technical users to access custom models without operating the cluster backend, improving the efficiency of custom model access and the ease of use of subsequent report configuration.

[0083] Based on the same inventive concept, this application also provides an Apache Doris-based cloud customer service system report generation apparatus for implementing the aforementioned Apache Doris-based cloud customer service system report generation method. The solution provided by this apparatus is similar to the implementation described in the above method. Therefore, the specific limitations of one or more Apache Doris-based cloud customer service system report generation apparatus embodiments provided below can be found in the limitations of the Apache Doris-based cloud customer service system report generation method described above, and will not be repeated here.

[0084] In one exemplary embodiment, such as Figure 4 As shown, a cloud customer service system report generation device based on Apache Doris is provided, comprising: The first output module is used to output a list of heterogeneous data sources corresponding to selectable cloud customer service systems through a visual interface. The first acquisition module is used to acquire the metadata of the target data source after receiving the target data source selection instruction submitted by the user based on the list of heterogeneous data sources, using the data catalog function of Apache Doris. The metadata includes the database name, table name, field name, and the storage location of the data corresponding to each field. The second output module is used to output a list of selectable statistical indicator fields, a list of selectable statistical dimension fields, and a list of selectable Doris functions through a visual interface. The list of statistical indicators includes the field names in the metadata. The second acquisition module is used to, after receiving the target statistical indicator field, target statistical dimension field, and target Doris function corresponding to the statistical report selected by the user, process the data of the target data source by calling the target Doris function according to the target statistical dimension field and the target statistical indicator field to obtain the statistical data corresponding to the statistical report. The third output module is used to output the statistical data through a visual interface.

[0085] In one embodiment, if the target statistical dimension field includes a statistical date, the second acquisition module is specifically used for: Using the statistical date as the dynamic partition key in Apache Doris, a partition pruning operation is triggered to filter partition data that matches the statistical date range from the data in the target data source. Using the statistical date as a filter condition, the processing logic corresponding to the target Doris function is executed on the cropped partition data to obtain the statistical data corresponding to the statistical report.

[0086] In one embodiment, if the target statistical dimension field includes a tenant identifier; the second acquisition module is specifically used for: Using the tenant identifier as a filtering condition, the data from the filtered target data source is processed by calling the target Doris function to obtain the statistical data corresponding to the statistical report.

[0087] In one embodiment, the third output module is specifically used for: The visual interface outputs a list of selectable report display types, including: display shape, display metrics, and display dimensions. After receiving the user's instruction to select a target display type based on the report display type list, the statistical data is displayed through a visual interface according to the target display type.

[0088] In one embodiment, the apparatus further includes: The fourth output module is used to output the upload entry point for custom Doris functions through the visual interface; The receiving module is used to receive user-submitted custom Doris function files; The execution module is used to call the Doris cluster's model registration interface and perform the registration operation of custom Doris functions; The update module automatically updates the list of selectable Doris functions after registration, adding custom Doris functions as independent options to the list.

[0089] In one embodiment, the list of heterogeneous data sources includes at least Hive, Iceberg, and MySQL data sources.

[0090] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 5As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media to run. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a report generation method for a cloud customer service system based on Apache Doris.

[0091] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0092] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0093] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0094] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0095] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0096] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0097] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0098] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0099] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for generating reports in a cloud customer service system based on Apache Doris, characterized in that, The method includes: A list of heterogeneous data sources corresponding to selectable cloud customer service systems is displayed through a visual interface; After receiving the target data source selection instruction submitted by the user based on the list of heterogeneous data sources, the metadata of the target data source is obtained using the data catalog function of Apache Doris, and the metadata includes field names; The system outputs a list of selectable statistical indicator fields, a list of selectable statistical dimension fields, and a list of selectable Doris functions through a visual interface. The list of statistical indicators includes the field names in the metadata. After receiving the target statistical indicator field, target statistical dimension field, and target Doris function corresponding to the statistical report selected by the user, the data of the target data source is processed by calling the target Doris function according to the target statistical dimension field and the target statistical indicator field to obtain the statistical data corresponding to the statistical report. The statistical data is output through a visual interface.

2. The method according to claim 1, characterized in that, If the target statistical dimension field includes a statistical date; the step of processing the data from the target data source by calling the target Doris function to obtain the statistical data corresponding to the statistical report includes: Using the statistical date as the dynamic partition key in Apache Doris, a partition pruning operation is triggered to filter partition data that matches the statistical date range from the data in the target data source. Using the statistical date as a filter condition, the processing logic corresponding to the target Doris function is executed on the cropped partition data to obtain the statistical data corresponding to the statistical report.

3. The method according to claim 1, characterized in that, If the target statistical dimension field includes a tenant identifier; the step of processing the data from the target data source by calling the target Doris function to obtain the statistical data corresponding to the statistical report includes: Using the tenant identifier as a filtering condition, the data from the filtered target data source is processed by calling the target Doris function to obtain the statistical data corresponding to the statistical report.

4. The method according to claim 1, characterized in that, The step of outputting the statistical data through a visual interface includes: The visual interface outputs a list of selectable report display types, including: display shape, display metrics, and display dimensions. After receiving the user's instruction to select a target display type based on the report display type list, the statistical data is displayed through a visual interface according to the target display type.

5. The method according to claim 1, characterized in that, The method further includes: The visual interface provides an entry point for uploading custom Doris functions; Receive user-submitted custom Doris function files; Call the Doris function registration interface to perform the registration operation of the custom Doris function; After registration, the list of selectable Doris functions will be automatically updated, and custom Doris functions will be added to the list as independent options.

6. The method according to claim 1, characterized in that, The list of heterogeneous data sources must include at least Hive, Iceberg, and MySQL data sources.

7. A report generation device for a cloud customer service system based on Apache Doris, characterized in that, The device includes: The first output module is used to output a list of heterogeneous data sources corresponding to selectable cloud customer service systems through a visual interface. The first acquisition module is used to acquire the metadata of the target data source after receiving the target data source selection instruction submitted by the user based on the list of heterogeneous data sources, using the data catalog function of Apache Doris. The metadata includes field names. The second output module is used to output a list of selectable statistical indicator fields, a list of selectable statistical dimension fields, and a list of selectable Doris functions through a visual interface. The list of statistical indicators includes the field names in the metadata. The second acquisition module is used to, after receiving the target statistical indicator field, target statistical dimension field, and target Doris function corresponding to the statistical report selected by the user, process the data of the target data source by calling the target Doris function according to the target statistical dimension field and the target statistical indicator field to obtain the statistical data corresponding to the statistical report. The third output module is used to output the statistical data through a visual interface.

8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that the processor executes the computer program to implement the steps of the Apache Doris-based cloud customer service system report generation method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the report generation method for a cloud customer service system based on Apache Doris as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the report generation method for a cloud customer service system based on Apache Doris as described in any one of claims 1-6.