A data processing method, device, and storage medium
Through cloud technology combined with ClickHouse database management system, the problem that traditional data query methods cannot meet large-scale real-time query, and realize sub-second response and efficient query analysis of massive data.
Patent Information
- Application Number
- CN202011375953.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-30
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2040-11-30
AI Technical Summary
Traditional data query methods cannot meet the needs of large-scale real-time data query. The existing technology responds to long response times and high loads when processing massive data, and cannot achieve real-time, efficiency and stability.
The cloud technology is combined with the database management system, and ClickHouse is used as a distributed column database. Through the association of content identification and indicator data, data query and analysis is realized, the structured query language is supported, and the data query process is optimized.
It realizes sub-second response of massive data, improves the efficiency and real-timeness of data query and analysis, reduces system response delay, and supports the scalability and operability of multiple data sources.
Smart Images

Figure CN112416991B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a data processing method, device and storage medium. Background Art
[0002] With the large-scale popularization and application of the Internet, the explosive growth of data volume marks the advent of the big data era. The development of massive data has brought many conveniences to people's lives, such as cloud storage, electronic payment, online shopping, etc., but it also brings with it severe challenges in processing massive data.
[0003] The terabyte-level growth of data means that traditional data query methods can no longer meet the needs of real-time query of such large-scale data. For example, the aggregate analysis of video clicks or text browsing data according to various dimensions has reached the billion level. The method of combining hive tables based on the Hadoop ecosystem with Spark application offline analysis takes a long time to execute, usually taking tens of minutes or even hours, and cannot meet the real-time requirements; the response time of MySQL aggregate analysis is also in the hour level, and the massive data will cause it to be overloaded and unavailable. Therefore, the real-time, efficiency and stability of data query analysis are all issues that need to be solved urgently. Summary of the invention
[0004] In view of the above problems, the embodiments of the present invention provide a data processing method, device and storage medium, which can improve the query and analysis speed of massive data, reduce the system response delay, and improve the efficiency of data query and analysis.
[0005] An embodiment of the present invention provides a data processing method, including:
[0006] Obtaining a data query and analysis request submitted by a client, wherein the data query and analysis request includes a data processing rule;
[0007] Acquiring target data from a database management system using the data processing rules, and analyzing and processing the target data to generate a processing result matching the data processing rules, the target data including a content identifier and corresponding indicator data;
[0008] The processing result is sent to the client, wherein the processing result includes one or both of a detailed data query result and an aggregated data analysis result.
[0009] An embodiment of the present invention provides a data processing device, including:
[0010] An acquisition module, used to acquire a data query and analysis request submitted by a client, wherein the data query and analysis request includes a data processing rule;
[0011] A processing module, configured to obtain target data from a database management system by using the data processing rule, and perform analysis and processing on the target data to generate a processing result that matches the data processing rule, where the target data includes a content identifier and corresponding metric data;
[0012] A sending module, configured to send the processing result to the client, where the processing result includes one or both of a detailed data query result and an aggregated data analysis result.
[0013] An embodiment of the present invention provides a computer device, including: a network interface, a processor, and a memory. The network interface, the processor, and the memory are connected to each other. The network interface is configured to provide a data communication function. The memory is configured to store a computer program. The processor is configured to call the computer program to execute some or all of the steps described in one aspect of the embodiments of the present invention.
[0014] An embodiment of the present invention provides a storage medium storing a computer program, where the computer program includes program instructions, and one or more processors load and execute the program instructions to perform the data processing method described in one aspect of the embodiments of the present invention.
[0015] An embodiment of the present invention provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data processing method described in one aspect.
[0016] It can be seen that in the embodiments of the present invention, by importing relevant data into a database management system, then querying and analyzing the massive data stored in the database management system according to the data processing rules set by the client, and obtaining target data in different dimensions, the characteristics of the database management system and the combination of data synchronization services can be used to define query data by oneself, and the response time of aggregated data query analysis can be shortened, the efficiency of online real-time analysis query can be improved, and what you see is what you get for the query result. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1a It is a schematic structural diagram of a data processing system provided by an embodiment of the present invention;
[0019] Figure 1b It is a schematic service configuration diagram in a server provided by an embodiment of the present invention;
[0020] Figure 2 It is a schematic step - flow diagram of a data processing method provided by an embodiment of the present invention;
[0021] Figure 3 It is a schematic step - flow diagram of a data processing method provided by an embodiment of the present invention;
[0022] Figure 4 It is a schematic step - flow diagram of a data processing method provided by an embodiment of the present invention;
[0023] Figure 5 It is a schematic structural diagram of a system architecture for index data synchronization provided by an embodiment of the present invention;
[0024] Figure 6 It is a schematic structural diagram of a data processing device provided by an embodiment of the present invention;
[0025] Figure 7 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners
[0026] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0027] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide - area network or a local - area network to achieve data computing, storage, processing, and sharing.
[0028] Cloud technology refers to the general term of network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The back-end services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the highly developed application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data requires the support of a powerful system. This can only be achieved through cloud computing.
[0029] A database, in short, can be regarded as an electronic filing cabinet - a place to store electronic files. Users can perform operations such as adding, querying, updating, and deleting data in the files. A so-called "database" is a data set stored together in a certain way, shared by multiple users, with as little redundancy as possible, and independent of application programs.
[0030] A database management system (English: Database Management System, abbreviated as DBMS) is a computer software system designed to manage databases. It generally has basic functions such as storage, interception, security protection, and backup. Database management systems can be classified according to the database models they support, such as relational, XML (Extensible Markup Language); or according to the types of computers they support, such as server clusters, mobile phones; or according to the query languages they use, such as SQL (Structured Query Language), XQuery; or according to the key performance metrics, such as maximum scale, highest running speed; or other classification methods. Regardless of the classification method used, some DBMS can cross categories. For example, they can support multiple query languages at the same time.
[0031] To better understand the solutions of the embodiments of the present invention, the following first introduces the relevant terms and concepts that may be involved in the embodiments of the present invention.
[0032] OLAP: Online Analytical Processing, which is specifically designed to support complex analysis operations. It can quickly and flexibly perform complex query and analysis processing on extremely large amounts of data according to the requirements of analysts.
[0033] OLTP: On-Line Transaction Processing. OLTP mainly performs operations such as adding, deleting, updating, and querying data, and has a relatively low processing delay.
[0034] Hive: A data warehouse platform based on Hadoop that can translate the Structured Query Language (SQL) written by users into corresponding MapReduce programs and execute them based on Hadoop.
[0035] Spark: An open-source general-purpose parallel computing framework.
[0036] ClickHouse: An online Multidimensional OLAP (MOLAP) software for multidimensional data storage, open-sourced by Yandex, which can query and analyze massive data through SQL.
[0037] PV: Page view, referring to the number of clicks and views of pictures and texts.
[0038] VV: Video view, referring to the number of video plays.
[0039] MySQL: An open-source database management system.
[0040] Please refer to Figure 1a , which is a schematic diagram of the architecture of a data processing system provided by an embodiment of the present invention and can be applied to a data management platform. The system includes: a client 100 and a server 101.
[0041] The client 100 may include multiple terminals as shown in Figure 1a , and corresponding users operate on the terminals. The terminals can be smartphones, tablets, laptops, desktop computers, smart speakers, smart watches, etc., but are not limited thereto. The client 100 is mainly used for interaction with the server 101, setting data screening conditions for users and presenting data query and analysis results to users, sending query and analysis requests and data import requests to the server 101 according to the screening conditions set by users, and receiving the processing results of data query and analysis sent by the server 101. Here, the processing results include any one or both of single-item detailed data and aggregated data.
[0042] The server 101 is used to query target data according to the query and analysis requests sent by the client 100, import content identifiers according to the data import requests sent by the client 100, and send the processing results of data query and analysis to the client 100.
[0043] Server 101 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present invention does not limit this.
[0044] In a possible embodiment, as Figure 1b shown, the services specifically running on server 101 may include: a data import service, a data management service, and a data processing service, where:
[0045] The data import service specifically includes a content identifier import interface 1011 and a metric data import interface 1012. These two types of interfaces are predefined data import interfaces and serve as a medium for synchronizing various types of data to the database management system 1013. The main purpose is to ensure the unity, integrity, and correctness of data import. Among them, the content identifier import interface 1011 is used to import content identifiers from different sources into the database management system 1013, such as data stored in various databases, local offline files, comma-separated value files from different sources, etc. The metric data import interface 1012 is used to synchronize metric data, that is, the task scheduling system is used to schedule the data synchronization service to synchronize metric data in different databases. The data synchronization service is preferably a Spark program. Using the Spark program to synchronize metric data in various databases can be paired with ClickHouse to achieve efficient data query tasks.
[0046] The data management service specifically includes a database management system 1013, which belongs to an OLAP system and is used to receive and store the data imported through the content identifier import interface 1011 and the metric data import interface 1012, calculate the target data according to the query statement sent by the data access layer service 1014, and return it to the data access layer service 1014. In addition, the database management system 1013 can automatically clean the data before importing the metric data to ensure the unity of data import.
[0047] The data processing service includes a data access layer service 1014, a query analysis engine 1015, a database 1016, and a query analysis service 1017.
[0048] The data access layer service 1014 is a low-level service, such as a DAO service, which is used to receive the query language forwarded by the query analysis service 1017 and forward the query statement to the database management system 1013 of the data management service, and then receive the target data queried in the database management system 1013 and return the target data to the query analysis service 1017.
[0049] The query analysis engine 1015 is used to receive the data query analysis request sent by the query analysis service 1017, find the target query interface in the database 1016, then parse the target query interface to generate a query statement, and send the query statement to the query analysis service 1017.
[0050] The database 1016 is used to store a variety of predefined query interfaces. The query analysis engine 1015 accesses the database 1016 according to the data query analysis request and obtains the target query interface from it.
[0051] The query analysis service 1017 is used to receive and forward relevant requests sent by the client 100, forward the query language parsed by the query analysis engine 1015, and send the processing result of the data query analysis to the client 100. Specifically, there is a graphic conversion adapter in the query analysis service 1017 that can convert the queried target data into chart-formatted data and send it to the client 100 as the processing result. After receiving the relevant request sent by the client 100, the query analysis service 1017 will forward it to the query analysis engine 1015.
[0052] In a specific implementation, before the query analysis operation, the metric data is automatically imported into the database management system 1013 at regular intervals through the metric data import interface 1012 for real-time data query analysis. The data export platform of the client 100 sends a data import request to the query analysis service 1017. After obtaining the data to be imported, the query analysis service 1017 automatically writes the content identifier of the data to be imported into the database, and then calls the data synchronization service to synchronize the content identifier to the database management system 1013 through the content identifier import interface 1011. Then, the client 100 sends a data query analysis request to the query analysis service 1017, and the query analysis service 1017 forwards the data query analysis request to the query analysis engine 1015. According to the data query analysis request, the query analysis engine 1015 searches for the target query interface in the database 1016 and parses the target query interface to generate a query statement. Then, the query analysis engine 1015 returns the query statement to the query analysis service 1017, and the query analysis service 1017 forwards the query statement to the data access layer service 1014. According to the query statement, the data access layer service 1014 forwards the query statement to the database management system 1013 to perform a data query operation. The database management system 1013 calculates according to the received query statement to obtain the target data, and then returns the target data to the data access layer service 1014. The data access layer service 1014 forwards it to the query analysis service 1017 again, and the graphical conversion adapter of the query analysis service 1017 can obtain the chart data and return it to the client 100 for visual display.
[0053] Using various predefined interfaces, including query interfaces, content identifier import interfaces, and metric data import interfaces, can improve the standardization of data query and import. Using an open-source online analytical processing system as the database management system can quickly and flexibly perform complex query analysis processing on extremely large amounts of data according to the requirements of analysts. Especially for temporarily imported content, it supports fast query analysis.
[0054] Please refer to Figure 2 , which is a schematic diagram of the step flow of a data processing method provided by an embodiment of the present invention and can be applied to a data management platform. The method includes:
[0055] S201, obtain a data query analysis request submitted by a client, where the data query analysis request includes a data processing rule.
[0056] In a possible embodiment, since the data volume level of a batch of temporarily demarcated content for query analysis is usually in the millions, in the face of such a large amount of data, query analysis requires a faster response speed and better query analysis efficiency. The client can directly experience the real-time nature and response speed of query analysis by providing users with diverse visual page displays and more intuitive data expressions. Specifically, the client can be a web front-end for data query analysis or an application software on a terminal device. The visual presentation on the client interface can directly allow users to experience all functions and set corresponding operations according to their needs, and then generate relevant data requests to be responded to and executed by the server. Optionally, corresponding data processing rules can be set through the query analysis interface of the client. Thus, the data query analysis request will include the data processing rules. Of course, the data query analysis request can also include other specific contents such as the address of the data query analysis, which is not limited here. The data processing rules are mainly selection conditions set for the data to be queried and analyzed. According to these selection conditions, users can customize the data processing rules, and in this way, more diverse data query analysis requests can also be obtained. Among them, the selection conditions can be generated by extracting fields from the full-volume index data or can be custom-set, which is not limited here. It should be noted that the selection conditions are different for different full-volume index data. For example, for index data of the video category, the selection conditions are the number of video plays, the number of likes, the number of collections, and the number of forwards. For index data of the shopping category, the selection conditions are the number of graphic and text clicks, the sales volume, the number of baby collections, and other index data. Further, multiple different clients can simultaneously submit different data query analysis requests for the same index data to the query analysis server for relevant devices to obtain and perform corresponding processing. The specific submission method and acquisition method are not limited here.
[0057] S202. Use the data processing rules to obtain target data from the database management system, and analyze and process the target data to generate a processing result that matches the data processing rules. The target data includes a content identifier and corresponding index data.
[0058] In a possible embodiment, the database management system preferably uses ClickHouse. As a distributed columnar database, it can achieve the distribution of data on different computer devices according to requirements, and supports structured query language and a variety of functions, including columnar compression technology, primary key index, MergeTree, single instruction multiple data stream instruction SIMD for vector calculation optimization, and support for the association between two large tables and other technical features, which can give full play to the computing power of computer devices and improve the speed of query analysis of massive data.
[0059] In a possible embodiment, the data included in the database management system is imported by the data synchronization service through the data import interface. Target data can be selected from these data according to the data processing rules. The finally obtained target data includes content identifiers and corresponding metric data. Among them, the data processing rules are further screening of the custom-defined circumscribed data, and the corresponding target data is the more detailed query data that the user wants to analyze. For example, the funny video data is circumscribed, but the data processing rules are set to analyze the like count or favorite count of this type of video data, involving real-time data query and analysis; the content identifier is used to associate the metric data and distinguish different target data, and it can be the index ID of the data. These content identifiers are extracted from the data specified by the user. After the content identifier and the metric data are associated, the database management system includes the metric data of the custom-defined circumscribed data. The metric data can be divided into static metrics and dynamic metrics. Among them, the static metrics do not change with time, such as the release date of a certain video or the label with "entertainment", etc.; the dynamic metrics include the access volume. For example, within a certain period of time, the access volume will change with the user's browsing or access, such as the video play count, the graphic and text click count, the data sharing volume, etc. The target data can be obtained by processing the metric data corresponding to the content identifier according to the data processing rules. For example, query the video sharing volume with the "entertainment" label imported yesterday as the target data. The original expression of the obtained target data is in tabular form, but this form alone is not sufficient to support the data analysis requirements. Therefore, specific analysis and processing need to be performed on the target data to generate the corresponding processing results. For example, the data sharing volume is statistically analyzed according to the sub-labels under the category label to generate the corresponding statistical chart.
[0060] S203. Send the processing result to the client, where the processing result includes one or both of the detailed data query result and the aggregated data analysis result.
[0061] In a possible embodiment, the query analysis server sends the processing result to the client. The processing result includes one or both of the detailed data query result and the aggregated data analysis result. Among them, the aggregated data is data of at least one dimension, such as the number of video plays or the number of graphic and text clicks under a certain classification on a certain date. Processing these data can obtain the aggregated data analysis result, and it can be displayed on the query analysis interface of the client in different forms such as line charts, pie charts, or bar charts to obtain the aggregated data analysis result more intuitively. Optionally, the user can choose to export the aggregated data analysis result as a local file for saving. The above method can better track user behavior and help decision-makers adjust product strategies according to the existing data analysis results. The detailed data is the query of specific index data for a single piece of content. For example, all or part of the index data of a specified ID is required, and finally it is displayed in an original table. Optionally, if there is a further need for the detailed data, statistical processing can also be performed to convert it into a more intuitive graphical expression. During the entire data query process, from the client submitting a query analysis request to finally receiving the processing result, it is all carried out on a visual interface. The client directly reflects the response speed of the entire query analysis system, and the cluster distribution of the database analysis system and some built-in tools can make the data query stable and quickly return the query result and complete the data analysis. Whether it is a detailed data query or an aggregated data query, the data volume of the query using this embodiment can achieve sub-second response and achieve true instant query.
[0062] In summary, the present embodiment has at least the following advantages:
[0063] In the database management system, the real-time query analysis target data is queried according to the query analysis request. By using the excellent characteristics of the database management system, especially ClickHouse: columnar storage and the implementation of the union between two large tables, sub-second response can be achieved for the query analysis of billions of data, load balancing can be achieved, the speed of real-time querying data is improved, and the response time of data query is reduced.
[0064] Please refer to Figure 3 , which is a schematic diagram of the step flow of a data processing method provided by an embodiment of the present invention and is applied to a data management platform. The method includes:
[0065] S301, obtain the data import request submitted by the client. The data import request includes a data filtering rule, and the data filtering rule includes one or more of a time limit condition and a category label;
[0066] In a possible embodiment, before querying and analyzing data, a batch of data content needs to be demarcated as the basis for screening target data. Therefore, data filtering rules need to be set in the data export platform of the client. There are multiple selection conditions in the filtering rules, and these selection conditions are automatically generated according to the retrieval of the full-content data by the data retrieval platform, covering all selectable fields of the full-content data. What is included in the content database are the static indicators corresponding to the fields, such as time limit conditions and category labels. The data filtering rules can set one or more of all the selection conditions to screen data. For example, screen out the content that was stored on November 11th and contains the "funny" label. Optionally, multiple data import requests submitted by the client can be obtained. The specific way to submit the data import request can be by clicking the corresponding operation button on the client or automatically submitting after the data filtering rules are set. There are no restrictions on the submission method and the method of obtaining the data import request here.
[0067] S302. Determine the data to be imported that meets the data filtering rules from the full data, where the full data includes data from multiple data sources.
[0068] In a possible embodiment, after the query analysis server responds to the data import request, it will determine the data that meets the data filtering rules in the full data as the data to be imported. Among them, the full data is the data in the full-content database mentioned above, and its magnitude can reach the hundreds of millions level, excluding some real-time statistical data. The full data can be stored in different databases to form the full-content database. The data in different databases corresponds to data from different data sources, such as comma-separated value files, local offline files, MySQL databases, Oracle databases, etc. Therefore, the data to be imported may also correspond to different data formats. Optionally, after determining the data to be imported, the data export platform will automatically write this batch of data to be imported into the database, so that the data to be imported enters the database management server as preparatory data.
[0069] S303. Call the content identifier import interface through the data synchronization service to import the content identifiers of the data to be imported into the database management system.
[0070] In a possible embodiment, after the query analysis server responds to the data import request and filters out the data to be imported, it will call the data synchronization service. The content identifier import interface is a pre-defined data import interface that can standardize the synchronization of data from various sources to the database management system. The definition of the interface well ensures the reliability of the data. The database management system can be ClickHouse, but ClickHouse itself does not have a mechanism for deduplicating primary keys. Therefore, when importing the data to be imported, it is necessary to ensure the idempotency and re-entrancy of the data import in the interface. Among them, the idempotency of the interface means that the result of a single request or multiple requests initiated by the user for the same operation is consistent, and there will be no side effects due to multiple clicks. In the four operations of addition, deletion, modification, and query, special attention should be paid to addition or modification, and the idempotency of the interface needs to be ensured; re-entrancy means that more than one task can concurrently use a re-entrant function without worrying about data errors, that is, the re-entrant function can be interrupted at any time and then continue to run later without losing data. Because the volume of the data to be imported is usually large, either all imports are successful or all imports fail. The specific method to achieve the above effects is to monitor the data to be imported in real time and detect whether there are failed data during the entire import process. If so, extract them and re-import them to ensure the integrity and correctness of the synchronized data. It should be noted that the content identifiers included in the data to be imported are imported into the database management system through the content identifier import interface in batches. The performance of ClickHouse for batch import can reach 10,000 records per second. After the synchronization is completed, the user can perform aggregation analysis on the index data corresponding to the data to be imported in the query analysis interface.
[0071] S304. Obtain a data query and analysis request submitted by the client, where the data query and analysis request includes a data processing rule.
[0072] S305. Use the data processing rule to obtain target data from the database management system, and perform analysis and processing on the target data to generate a processing result that matches the data processing rule. The target data includes content identifiers and corresponding index data.
[0073] S306. Send the processing result to the client, where the processing result includes one or both of a detailed data query result and an aggregated data analysis result.
[0074] In this embodiment, the specific implementation manners of steps S304 - S306 can refer to Figure 2 S201 - S203 in the embodiment shown, which will not be elaborated here.
[0075] In summary, the embodiments of the present invention have at least the following advantages:
[0076] Import the content identifiers included in the data to be imported from different sources into the database management system through the content data import interface, and then use the content identifiers for the association between tables. Filter the target data through the association operation and relevant rules to obtain the results of real-time query and analysis, realizing customized data query and analysis; the support for importing from multiple data sources also makes the query and analysis system highly scalable and supports more scenarios; by using a database that supports the association between large tables, it provides good support for application scenarios such as real-time query and analysis of PV / VV data of millions of batches of content.
[0077] Please refer to Figure 4 , which is a schematic diagram of the step process of another data processing method provided by an embodiment of the present invention, applied to a data management platform. The method includes:
[0078] S401, obtain the metric data of the full volume of data, and call the metric data import interface through the data synchronization service to import the metric data of the full volume of data into the database management system. The metric data includes the access volume.
[0079] Since the metric data is huge in volume, usually in the order of hundreds of millions of data, in order to ensure the integrity and reliability of the data synchronized to the database management system, it is necessary to call the metric data import interface through the data synchronization service to import the metric data of the full volume of data into the database management system, including the access volume. Here, the access volume refers to the amount of data of user browsing content, such as real-time statistical data of video play counts, graphic and text view counts, and favorite counts.
[0080] As a possible embodiment, please refer to Figure 5, the data synchronization service 502 can be scheduled regularly by the task scheduling system 501. For example, the data synchronization can be performed by using a Spark program through the metric data import interface. The specific process includes: First, when the scheduled time of the task scheduling system 501 arrives, the data synchronization service 502 is called to read the metadata information of the full-volume data's metric data from the data warehouse. Here, the data synchronization service 502 can be a Spark program, which will start to execute the synchronization task after being scheduled, that is, start to read the metadata from the data warehouse 503. Among them, the data warehouse 503 can be a Hive data warehouse, which stores the metric data. Then, a distributed table is created in the database management system according to the read metadata information. Specifically, after reading the metadata, it is first checked whether there is a corresponding table in the database management system 504. If not, the table creation statement will be executed to create a distributed table. Since the database management system 504 is usually deployed in a cluster, creating a distributed table requires executing the local table creation statement on each node of the cluster distribution. Finally, a table of the distributed engine is created, and subsequent data updates are all carried out through the distributed table; if it exists, the fields in the table, that is, whether the metadata information is consistent with that in the data warehouse 503, will be checked. If not, it will be modified to the metadata format in the data warehouse 503 to be compatible with the situation where the data table in the data warehouse 503 has been updated. After the data synchronization service 502 checks the metadata information, it will really start to synchronize the data: obtain the metric data of the full-volume data from the data warehouse 503, and through calling the metric data import interface, insert the metric data of the full-volume data into the database management system 504 in batches by using the distributed table. Specifically, before inserting the data, first, the corresponding date partition in the database management system 504 will be deleted to prevent the existence of dirty data, because when the task is rerun, there is still dirty data inserted during the previous task execution in the database management system 504; then, the metric data in the data warehouse 503 will be loaded. Since the metric data itself is very large in volume, reaching the order of hundreds of millions, it needs to be inserted in batches step by step to ensure that the server load of the database management system 504 cluster is not affected. The above steps of metric data import are applicable to synchronizing the full-volume content PV / VV data from the Hive data warehouse to the database management system ClickHouse.
[0081] As can be seen from the above process, the step of checking whether the metadata information is consistent before the data synchronization service 502 really synchronizes the data can ensure that the tables in the database management system 504 and the tables in the data warehouse 503 are automatically kept consistent, without the need for manual modification of the table information. In this data synchronization service 502, only the necessary information such as the IP addresses of the cluster nodes of the database management system 504 and the table names in the data warehouse 503 to be synchronized needs to be simply configured, and the data can be automatically synchronized without caring about the table creation statements on the database management system 504 side, realizing the automation and configurability of the synchronization program.
[0082] The import of indicator data is necessary before data query and analysis. However, the timed import of indicator data is for the real-time update of indicator data in the database management system to ensure the accuracy of query results. Therefore, the timed import of indicator data can continue during the data query and analysis process, independent of the data query and analysis, and will not affect the operation of data query and analysis.
[0083] S402. Obtain the data query and analysis request submitted by the client, where the data query and analysis request includes data processing rules.
[0084] In this embodiment, the specific implementation of this step can refer to Figure 2 S201 in the illustrated embodiment, which will not be elaborated here.
[0085] S403. The data processing rules include data screening conditions, and determine the query statement corresponding to the data screening conditions.
[0086] In a possible embodiment, the data processing rules include data screening conditions, and the data screening conditions here are the query conditions selected on the client. Since the communication between various devices or services in the data query and analysis system uses a unified language for data query, it can make the operation more efficient and convenient. Therefore, the primary task of obtaining target data from the database management system using the data processing rules is to determine the query statement corresponding to the data screening conditions, and then use the query statement to query the target data. Among them, the query statement is preferably Structured Query Language (SQL), and it can also be other languages that can achieve efficient query, which is not limited here. A series of query interfaces will be predefined in the database of the query analysis system. Each interface implements a specific query mode, and the mode definition contains all the information of a data query, including data screening conditions. By calling the query analysis engine according to the data screening conditions, the target query interface can be determined from multiple predefined query interfaces, and then the query analysis engine is used to parse the target query interface, and the corresponding query statement can be generated. Optionally, the front-end page can automatically render the front-end menu component according to the meta-information of the query mode to achieve the automation of the data query and analysis system. Users can assemble different query interfaces into a query view, and each time this view is opened, all the query interfaces included in the view will be automatically submitted to the query analysis engine to generate specific query statements. Defining query interfaces in the data query and analysis operation ensures the scalability and operability of data analysis.
[0087] S404. Call the data access layer service, and use the query statement to query the target data from the database management system.
[0088] The data access layer service can be the underlying DAO service. Invoking the data access layer service can identify the query statement and perform the most basic operations of adding, deleting, modifying, and querying. Therefore, the target data can be queried from the database management system using this query statement. Specifically, after receiving the query statement, the query analysis server passes the query statement to the data access layer service. The data access layer service queries the target data from the index data corresponding to the content identifier of the data to be imported included in the database management system according to the query statement. The database management system can be ClickHouse that supports the association (join) between large tables. The two large tables of the data to be imported and the index data are joined to obtain the index data corresponding to the content index, and then the target data can be queried from the index data corresponding to the content index according to the query statement. After that, the query analysis server will obtain the target data returned by the data access layer service to perform subsequent operations.
[0089] In this process, since the data access layer service and the data analysis engine are separated, the update of the capabilities of each module will not affect the operation of other modules. Therefore, when the definition of the query interface is upgraded to a more complex mode, only the parsing ability of the query analysis engine needs to be upgraded, but the query analysis engine always returns the specific query statement, and the data access layer service is unaware of this capability upgrade.
[0090] S405. Analyze and process the target data to generate a processing result that matches the data processing rule.
[0091] In a possible embodiment, the data processing rule includes the chart type. That is to say, when querying data, in addition to setting data filtering conditions, the way the processing result wants to be expressed can also be selected according to needs to achieve the analysis and processing of the target data. For example, the target data can be displayed on the analysis interface in the form of a line chart, a pie chart, a bar chart, etc. Specifically, the target data can be converted into the previously set chart type through the graphic conversion adapter in the query analysis server. Each chart type corresponds to the chart data that has been statistically processed, and these chart data can be used as the final processing result to help data analysts conduct relevant analysis.
[0092] S406. Send the processing result to the client, where the processing result includes one or both of the detailed data query result and the aggregated data analysis result.
[0093] In this embodiment, the specific implementation manner of this step can refer to Figure 2 S203 in the shown embodiment, which will not be elaborated here.
[0094] In summary, the embodiments of the present invention have at least the following advantages:
[0095] By defining the index data import interface and the data query interface, the scalability and operability of data import and data analysis are ensured. Among them, by the index data import interface, the full volume of index data is synchronously updated at regular intervals, making the synchronization of data more convenient; the data analysis engine and the data access layer service are separated. The data analysis engine parses the defined query interface, and then hands it over to the data access layer service to query the data. Then, the graphic conversion adapter in the query analysis service converts the target data returned by the data access layer service into the corresponding graphic format data. Such a design can make the responsibilities of each module clearer, reduce the coupling between modules, and also make the scalability of massive data query analysis operations larger, and the processing process more efficient and stable.
[0096] Please refer to Figure 6 , which is a schematic structural diagram of a data processing device provided in this embodiment. The device includes:
[0097] An acquisition module 601, configured to acquire a data query and analysis request submitted by a client, where the data query and analysis request includes a data processing rule;
[0098] A processing module 602, configured to use the data processing rule to acquire target data from a database management system, and perform analysis and processing on the target data to generate a processing result that matches the data processing rule, where the target data includes a content identifier and corresponding index data;
[0099] A sending module 603, configured to send the processing result to the client, where the processing result includes one or both of a detailed data query result and an aggregated data analysis result.
[0100] In a possible embodiment, the device further includes: a determination module 604 and an import module 605, where:
[0101] The acquisition module 601 is configured to acquire a data import request submitted by the client, where the data import request includes a data filtering rule, and the data filtering rule includes one or more of a time limit condition and a category label;
[0102] The determination module 604 is configured to determine the data to be imported that meets the data filtering rule from the full volume of data, where the full volume of data includes data from multiple data sources;
[0103] The import module 605 is configured to call the content identifier import interface through a data synchronization service to import the content identifier of the data to be imported into the database management system.
[0104] In a possible embodiment, the processing module 602 is further configured to:
[0105] Obtain the metric data of the full - volume data, and call the metric data import interface through the data synchronization service to import the metric data of the full - volume data into the database management system, where the metric data includes the number of visits.
[0106] In a possible embodiment, the processing module 602 is further configured to:
[0107] Determine a query statement corresponding to the data screening condition;
[0108] Call the data access layer service and use the query statement to query target data from the database management system.
[0109] In a possible embodiment, the processing module 602 is further configured to:
[0110] Determine a target query interface from a variety of predefined query interfaces according to the data screening condition;
[0111] Use a query analysis engine to parse the target query interface to generate a corresponding query statement.
[0112] In a possible embodiment, the processing module 602 is further configured to:
[0113] Use the chart type and the graphic conversion adapter to analyze and process the target data to generate corresponding chart data;
[0114] Take the chart data as the processing result.
[0115] In a possible embodiment, the processing module 602 is further configured to:
[0116] When the scheduled time of the task scheduling system arrives, call the data synchronization service to read the metadata information of the metric data of the full - volume data from the data warehouse;
[0117] Create a distributed table in the database management system according to the metadata information;
[0118] Obtain the metric data of the full - volume data from the data warehouse, and through calling the metric data import interface, use the distributed table to insert the metric data of the full - volume data into the database management system in batches.
[0119] In a possible embodiment, the processing module 602 is further configured to:
[0120] Pass the query statement to the data access layer service so that the data access layer service queries target data from the metric data corresponding to the content identifier of the data to be imported included in the database management system according to the query statement;
[0121] Obtain the target data returned by the data access layer service.
[0122] For the device embodiment, since it is basically similar to the method embodiment, the relevant parts can refer to the partial description of the method embodiment.
[0123] Please refer to Figure 7 , which is a schematic structural diagram of a computer device provided by an embodiment of the present invention. As Figure 7 shown, the computer device may include a processor 701, a memory 702, a network interface 703, and at least one communication bus 704. Among them, the processor 701 is used to schedule computer programs, and may include a central processing unit, a controller, and a microprocessor; the memory 702 is used to store computer programs, and may include a high-speed random access memory and a non-volatile memory, such as a disk storage device and a flash memory device; the network interface 703 provides data communication functions; the communication bus 704 is responsible for connecting each communication component.
[0124] Among them, the processor 701 may be used to call the computer program in the memory to perform the following operations:
[0125] Obtain a data query and analysis request submitted by the client, where the data query and analysis request includes a data processing rule;
[0126] Use the data processing rule to obtain target data from the database management system, and perform analysis and processing on the target data to generate a processing result that matches the data processing rule, where the target data includes a content identifier and corresponding metric data;
[0127] Send the processing result to the client, where the processing result includes one or both of a detailed data query result and an aggregated data analysis result.
[0128] Optionally, the processor 701 is specifically used for:
[0129] Obtain a data import request submitted by the client, where the data import request includes a data filtering rule, and the data filtering rule includes one or more of a time limit condition and a category label;
[0130] Determine the data to be imported that meets the data filtering rule from the full amount of data, where the full amount of data includes data from multiple data sources;
[0131] Call the content identifier import interface through the data synchronization service to import the content identifier of the data to be imported into the database management system.
[0132] Optionally, the processor 701 is specifically used for:
[0133] Obtain the metric data of the full amount of data, and call the metric data import interface through the data synchronization service to import the metric data of the full amount of data into the database management system, where the metric data includes the number of visits.
[0134] Optionally, the processor 701 is specifically configured to:
[0135] Determine a query statement corresponding to the data filtering condition;
[0136] Call the data access layer service and use the query statement to query target data from the database management system.
[0137] Optionally, the processor 701 is specifically configured to:
[0138] Determine a target query interface from multiple predefined query interfaces according to the data filtering condition;
[0139] Use a query analysis engine to parse the target query interface to generate a corresponding query statement.
[0140] Optionally, the processor 701 is specifically configured to:
[0141] Use the chart type and the graphic conversion adapter to analyze and process the target data to generate corresponding chart data;
[0142] Use the chart data as the processing result.
[0143] Optionally, the processor 701 is specifically configured to:
[0144] When the scheduled time of the task scheduling system arrives, call the data synchronization service to read the metadata information of the metric data of the full amount of data from the data warehouse;
[0145] Create a distributed table in the database management system according to the metadata information;
[0146] Obtain the metric data of the full amount of data from the data warehouse, and through calling the metric data import interface, use the distributed table to insert the metric data of the full amount of data into the database management system in batches.
[0147] Optionally, the processor 701 is specifically configured to:
[0148] Pass the query statement to the data access layer service, so that the data access layer service queries target data from the metric data corresponding to the content identifier of the data to be imported included in the database management system according to the query statement;
[0149] Obtain the target data returned by the data access layer service.
[0150] The computer device in the embodiment of the present invention can be used to execute the technical solutions in the above method embodiments. The implementation principles and technical effects are similar and will not be elaborated here.
[0151] The embodiment of the present invention further provides a storage medium, in which a computer program for the foregoing network access method is stored. The computer program includes program instructions. When one or more processors load and execute the program instructions, the description of the data processing method in the embodiment can be implemented and will not be elaborated here. The description of the beneficial effects of using the same method will also not be elaborated here. It can be understood that the program instructions can be deployed to be executed on one or multiple terminal devices that can communicate with each other.
[0152] The embodiment of the present invention further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps executed in the above method embodiments.
[0153] The foregoing disclosure is only for the preferred embodiments of the present invention, and of course, the scope of the rights of the present invention cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. A data processing method, characterized in that, Applied to a data management platform, the method includes: Obtain a data import request submitted by a client, where the data import request includes a data filtering rule, and the data filtering rule includes one or more of a time limit condition and a category label; Determine the data to be imported that meets the data filtering rule from the full-volume data, and call the content identifier import interface through the data synchronization service to import the content identifier of the data to be imported into the database management system; the full-volume data includes data from multiple data sources; Obtain the metric data of the full-volume data, and call the metric data import interface through the data synchronization service to import the metric data of the full-volume data into the database management system, where the metric data includes the access volume; the access volume includes at least one of the following: video playback volume, graphic view volume, data sharing volume, and favorite volume; Obtain a data query and analysis request submitted by the client, where the data query and analysis request includes a data processing rule; Use the data processing rule to obtain target data from the database management system, and perform analysis and processing on the target data to generate a processing result that matches the data processing rule, where the target data includes a content identifier and the corresponding metric data; Send the processing result to the client, where the processing result includes one or both of a detailed data query result and an aggregated data analysis result.
2. The method according to claim 1, wherein The data processing rule includes a data screening condition, and the step of using the data processing rule to obtain target data from the database management system includes: Determine a query statement corresponding to the data screening condition; Call the data access layer service and use the query statement to query the target data from the database management system.
3. The method according to claim 2, characterized in that, The step of determining a query statement corresponding to the data screening condition includes: Determine a target query interface from multiple predefined query interfaces according to the data screening condition; Use a query analysis engine to parse the target query interface to generate a corresponding query statement.
4. The method according to claim 1, wherein The data processing rule further includes a chart type, and the step of performing analysis and processing on the target data to generate a processing result that matches the data processing rule includes: Use the chart type and a graphic conversion adapter to perform analysis and processing on the target data to generate corresponding chart data; Use the chart data as the processing result.
5. The method according to claim 1, wherein The step of obtaining the metric data of the full-volume data and calling the metric data import interface through the data synchronization service to import the metric data of the full-volume data into the database management system includes: When the scheduled time of the task scheduling system arrives, call the data synchronization service to read the metadata information of the metric data of the full-volume data from the data warehouse; Create a distributed table in the database management system according to the metadata information; Obtain the metric data of the full-volume data from the data warehouse, and insert the metric data of the full-volume data into the database management system in batches by using the distributed table through calling the metric data import interface.
6. The method according to claim 2, wherein Invoking the data access layer service to query target data from the database management system using the query statement includes: Passing the query statement to the data access layer service so that the data access layer service queries target data from the metric data corresponding to the content identifier of the data to be imported included in the database management system according to the query statement; Obtaining the target data returned by the data access layer service.
7. A data processing device, characterized in that, Includes: An acquisition module for acquiring a data import request submitted by a client, where the data import request includes a data filtering rule, and the data filtering rule includes one or more of a time limit condition and a category label; A determination module for determining the data to be imported that conforms to the data filtering rule from the full-volume data; the full-volume data includes data from multiple data sources; An import module for invoking a content identifier import interface through a data synchronization service to import the content identifier of the data to be imported into the database management system; A processing module for obtaining the metric data of the full-volume data, and invoking a metric data import interface through the data synchronization service to import the metric data of the full-volume data into the database management system, where the metric data includes the access volume; the access volume includes at least one of the following: video playback volume, graphic and text view volume, data sharing volume, and collection volume; The acquisition module is further configured to acquire a data query and analysis request submitted by the client, where the data query and analysis request includes a data processing rule; The processing module is further configured to obtain target data from the database management system using the data processing rule, and perform analysis and processing on the target data to generate a processing result that matches the data processing rule, where the target data includes a content identifier and corresponding metric data; A sending module for sending the processing result to the client, where the processing result includes one or both of a detailed data query result and an aggregated data analysis result.
8. A storage medium, characterized in that, The storage medium stores a computer program, and the computer program includes program instructions, which are loaded and executed by one or more processors to execute the method according to any one of claims 1-6.
9. A computer program product, characterized in that, The computer program product includes computer instructions, and when the computer instructions are read and executed by a processor from a computer-readable storage medium, the method according to any one of claims 1-6 is executed.
Citation Information
Patent Citations
Data inquiry control method and device, computer device and storage medium
CN109446253A
Data query method and device based on rule configuration, storage medium and terminal
CN111259037A