Real-time data processing method, system and equipment of text travel big data platform and medium
Through real-time data processing methods, combined with Kafka cluster and ClickHouse engine, data cleaning and query are used for split-layer data warehouses, the problem of difficulty in digging deep laws and potential value behind cultural and tourism data is solved, realizing the immediacy and efficient processing of data, and improving the management and service level of scenic spots.
Patent Information
- Application Number
- CN202510057969.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-06-13
AI Technical Summary
The data analysis solutions of existing cultural and tourism scenic spot platforms are difficult to explore the deep laws and potential value behind the data based on actual needs, which limits the improvement of scenic spot management and service levels.
The user behavior log is captured through the page burial of the cultural and tourism information operation platform, combined with the scenic spot ticketing system to obtain ticket purchase business data, use the Kafka cluster real-time data processing, the ClickHouse engine quickly imports data, and cleans, converts and querys through the split-layer data warehouse, and calls the associated operation interface for further processing.
It realizes the immediacy and efficient processing of data, improves the data quality and query accuracy, explores the deep laws and potential value behind the data, and improves the management and service level of scenic spots.
Smart Images

Figure CN120144676A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of cultural tourism big data platform, and in particular to a real-time data processing method, system, device and medium for a cultural tourism big data platform. Background Art
[0002] In recent years, with the rapid development of science and technology and the continuous expansion of the Internet, the cultural and tourism industry is undergoing a profound information-based transformation. This transformation is not only reflected in tourists' demand for high-quality, personalized travel experiences, but also in the comprehensive upgrade of tourism companies' management and service models. The widespread application of information technology in food, accommodation, transportation, travel, shopping, entertainment and other aspects has greatly improved tourists' travel quality and satisfaction, and has also brought new development opportunities and challenges to cultural and tourism companies.
[0003] With the continuous deepening of informatization construction, the amount of data obtained by scenic spot platforms has shown explosive growth. In order to achieve the effective optimization of human and material resources for different scenic spots in the scenic area. At present, most of the mainstream scenic spot platform data analysis solutions on the market are based on historical data for trend prediction or simple real-time monitoring. Although they can help managers understand the flow of tourists to a certain extent, in the face of complex and changeable actual operation scenarios, these solutions are difficult to meet actual needs and difficult to explore the deep laws and potential value behind the data, thus limiting the improvement of scenic spot management and service levels. Summary of the invention
[0004] In view of the above-mentioned deficiencies in the prior art, the present application provides a real-time data processing method, system, device and medium for a cultural and tourism big data platform to solve the problem that existing solutions are difficult to explore the deep laws and potential values behind the data according to actual needs.
[0005] In the first aspect, the present application provides a real-time data processing method for a cultural and tourism big data platform, the method comprising: Capture user behavior logs through page embedding of the cultural and tourism information operation platform; obtain tourists' ticket purchase business data through the scenic spot ticketing system; use the data source of user behavior logs and ticket purchase business data as the data source of the Kafka cluster; obtain user behavior logs and ticket purchase business data from the data source of the Kafka cluster, and use the ClickHouse engine to import them into the preset split-layer data warehouse; use the ODS layer of the split-layer data warehouse to store the imported original user behavior logs and ticket purchase business data; use the preset cleaning algorithm in the DWD layer of the split-layer data warehouse to remove meaningless data from the original user behavior logs and ticket purchase business data; use the preset demand acquisition interface in the DIM layer of the split-layer data warehouse to obtain query conditions and associated operation interface names, and extract data that meets the preset query conditions from the user behavior logs and ticket purchase business data in the DWD layer; use the DM layer of the split-layer data warehouse to call the operation program corresponding to the associated operation interface, input the data that meets the preset query conditions into the operation program, and obtain the result data.
[0006] The real-time data processing method provided in the embodiment of the present application obtains user behavior logs and ticket purchase business data in real time through the Kafka cluster, ensuring the immediacy of the data. The ClickHouse engine is used for fast data import to improve the data processing speed. The split-layer data warehouse (including ODS, DWD, DIM, and DM layers) is used to store and process data in a hierarchical manner, making data management more orderly and efficient. This hierarchical structure facilitates data cleaning, conversion, and query, and improves the flexibility and scalability of data processing. The preset cleaning algorithm in the DWD layer removes meaningless data from the original data, improving the quality of the data. The DIM layer can obtain specific query conditions and associated operation interfaces, making data queries more accurate and targeted. The DM layer calls the operating procedures corresponding to the associated operation interfaces to further process the data that meets the preset query conditions, and digs out the deep laws and potential value behind the data.
[0007] In one implementation of the present application, after capturing user behavior logs through page embedding of the cultural tourism information operation platform and obtaining tourists' ticket purchase business data through the scenic spot ticketing system, the method further includes: Use the Nginx cluster to buffer user behavior logs and ticket purchase business data, and write structured business data into the MySQL database; Detect whether each data in the MySQL database has preset abnormal data and missing data; remove the data rows with preset abnormal data and missing data, and obtain the processed user behavior log and ticket purchase business data; Among them, the user behavior logs include user browsing, clicking, staying, commenting, liking, and favoriting behaviors; the ticket purchase business data includes structured data stored such as the current number of ticket purchases, time, age, number of people, and payment.
[0008] In one implementation manner of the present application, using a preset cleaning algorithm in the DWD layer of the split-layer data warehouse, meaningless data in the original user behavior logs and ticket purchase business data is removed, specifically including: Based on the user's unique IP in the original user behavior logs as an identifier, in the order of the log generation time, read user behaviors from consecutive original user behavior logs; According to the preset chain relationship between user behaviors, determine whether the current user behavior generates associated user behaviors corresponding to the preset chain relationship within a subsequent preset time period. When the current user behavior does not generate associated user behaviors corresponding to the preset chain relationship within the subsequent preset time period, determine the user behavior log corresponding to the current user behavior as meaningless data; According to the unique ticket purchase identifier in the ticket purchase business data, in the order of the time when the ticket purchase business data is generated, read ticket purchase business behaviors from consecutive ticket purchase business data; among them, the ticket purchase business behaviors include: page browsing, page querying, information input, and ticket purchase behaviors; Determine whether any non-ticket purchase behavior generates a ticket purchase behavior within a subsequent preset waiting time period; when no ticket purchase behavior is generated within the subsequent preset waiting time period, determine the ticket purchase business data corresponding to the current non-ticket purchase behavior as meaningless data; Remove meaningless data from the original user behavior logs and ticket purchase business data.
[0009] In one implementation manner of the present application, in the DM layer of the split-layer data warehouse, after calling the operation program corresponding to the association operation interface, inputting the data that meets the preset query conditions into the operation program, and obtaining the result data, the method further includes: Write the result data into the ClickHouse engine; Read the result data through several reading terminals associated with the ClickHouse engine.
[0010] In a second aspect, the present application provides a real-time data processing system for a cultural and tourism big data platform, and the system includes: An acquisition module, configured to capture user behavior logs through page embedding of the cultural and tourism information operation platform; and obtain ticket purchase business data of tourists through the scenic area ticketing system; A Kafka component, configured to use the data sources of user behavior logs and ticket purchase business data as the data sources of the Kafka cluster; The ClickHouse engine call module is used to obtain user behavior logs and ticket purchase business data from the data source of the Kafka cluster, and use the ClickHouse engine to import them into a preset split-layer data warehouse; The split-layer data warehouse module is used to store the imported original user behavior logs and ticket purchase business data by using the ODS layer of the split-layer data warehouse; use the preset cleaning algorithm in the DWD layer of the split-layer data warehouse to remove meaningless data from the original user behavior logs and ticket purchase business data; use the preset requirement acquisition interface in the DIM layer of the split-layer data warehouse to obtain query conditions and associated operation interface names, and extract data that meets the preset query conditions from the user behavior logs and ticket purchase business data in the DWD layer; use the DM layer of the split-layer data warehouse to call the operation program corresponding to the associated operation interface, input the data that meets the preset query conditions into the operation program, and obtain the result data.
[0011] In an implementation manner of the present application, the acquisition module includes a preprocessing unit, which is used to buffer the user behavior logs and ticket purchase business data using the Nginx cluster and write the structured business data into the MySQL database; detect whether each piece of data in the MySQL database has preset abnormal data and missing data; remove the data rows with preset abnormal data and missing data to obtain the processed user behavior logs and ticket purchase business data; wherein, the user behavior logs include user browsing, clicking, staying, commenting, liking, and favoriting behaviors; the ticket purchase business data includes structured data stored for the current number of ticket purchases, time, age, number of people, and payment.
[0012] In an implementation manner of the present application, the split-layer data warehouse module includes a meaningless identification unit, which is used to identify based on the user's unique IP in the original user behavior logs, and read user behaviors from consecutive original user behavior logs in the order of the log generation time; According to the preset chain relationship between user behaviors, determine whether the current user behavior generates an associated user behavior corresponding to the preset chain relationship within a subsequent preset time period. When the current user behavior does not generate an associated user behavior corresponding to the preset chain relationship within the subsequent preset time period, determine that the user behavior log corresponding to the current user behavior is meaningless data; According to the unique ticket purchase identifier in the ticket purchase business data, read the ticket purchase business behaviors from consecutive ticket purchase business data in the order of the ticket purchase business data generation time; wherein, the ticket purchase business behaviors include: page browsing, page query, information input, and ticket purchase behaviors; Determine whether any non-ticket purchasing behavior generates a ticket purchasing behavior within a subsequent preset waiting time period; when no ticket purchasing behavior is generated within the subsequent preset waiting time period, determine that the ticket purchasing business data corresponding to the current non-ticket purchasing behavior is meaningless data; Remove meaningless data from original user behavior logs and ticket purchase business data.
[0013] In one implementation of the present application, the system also includes a reading module for writing the result data into the ClickHouse engine; and reading the result data through a number of reading terminals associated with the ClickHouse engine.
[0014] In the third aspect, the present application provides a real-time data processing device for a cultural and tourism big data platform, the device comprising: a processor; and a memory on which executable code is stored, and when the executable code is executed, the processor executes a real-time data processing method for a cultural and tourism big data platform such as any one of the above items.
[0015] In a fourth aspect, the present application provides a non-volatile computer storage medium on which computer instructions are stored. When the computer instructions are executed, they implement a real-time data processing method for a cultural and tourism big data platform such as any of the above items.
[0016] Those skilled in the art can understand that the present application has at least the following beneficial effects: The present application discloses a real-time data processing method, system, device and medium for a cultural tourism big data platform, which obtains user behavior logs and ticket purchase business data in real time through a Kafka cluster to ensure the immediacy of data. The ClickHouse engine is used for fast data import to improve the data processing speed. The split-layer data warehouse (including ODS, DWD, DIM, and DM layers) is used to store and process data in layers, making data management more orderly and efficient. This hierarchical structure facilitates data cleaning, conversion, and query, and improves the flexibility and scalability of data processing. Through the preset cleaning algorithm in the DWD layer, meaningless data in the original data is removed, improving the quality of the data. The DIM layer can obtain specific query conditions and associated operation interfaces (actual needs), so that data queries can be associated with actual needs. The DM layer calls the operation program corresponding to the associated operation interface to further process the data that meets the preset query conditions, and digs out the deep laws and potential value (result data) behind the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Some embodiments of the present disclosure are described below with reference to the accompanying drawings, in which: Figure 1 It is a flow chart of a real-time data processing method for a cultural and tourism big data platform provided in an embodiment of the present application.
[0018] Figure 2 It is a schematic diagram of the internal structure of a real-time data processing system for a cultural and tourism big data platform provided by an embodiment of the present application.
[0019] Figure 3 It is a schematic diagram of the internal structure of a real-time data processing device for a cultural and tourism big data platform provided by an embodiment of the present application. Detailed implementation manners
[0020] Those skilled in the art should understand that the embodiments described below are only the preferred embodiments of the present disclosure, and do not mean that the present disclosure can only be implemented through these preferred embodiments. These preferred embodiments are only used to explain the technical principles of the present disclosure and are not used to limit the protection scope of the present disclosure. Based on the preferred embodiments provided by the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts should still fall within the protection scope of the present disclosure.
[0021] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the element.
[0022] The technical solutions proposed by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0023] The embodiment provides a real-time data processing method for a cultural and tourism big data platform. As Figure 1 shown, the method provided by the embodiment of the present application mainly includes the following steps: Step 110: Capture user behavior logs through page tagging of the cultural and tourism informatization operation platform; obtain the ticket purchase business data of tourists through the scenic spot ticketing system.
[0024] It should be noted that in this step, through the page tagging technology of the cultural and tourism informatization operation platform, various user behavior logs on the platform are captured, such as browsing, clicking, staying, commenting, liking, collecting and other user behavior logs. Through the scenic spot ticketing system, the ticket purchase business data of tourists is obtained, including structured data such as the current number of ticket purchasers, time, age, payment information, etc.
[0025] In some embodiments, after capturing user behavior logs through page tagging of the cultural and tourism informatization operation platform and obtaining the ticket purchase business data of tourists through the scenic spot ticketing system, the method further includes: Buffer the user behavior logs and ticket purchase business data using an Nginx cluster, and write the structured business data into a MySQL database; Detect whether each piece of data in the MySQL database has preset abnormal data and missing data; remove the data rows with preset abnormal data and missing data to obtain the processed user behavior logs and ticket purchase business data; Among them, the user behavior logs include user browsing, clicking, staying, commenting, liking, and favoriting behaviors; the ticket purchase business data includes structured data such as the current number of ticket purchases, time, age, number of people, and payment stored.
[0026] Those skilled in the art can understand that through data buffering using an Nginx cluster in this application, high-concurrency data requests can be effectively handled, improving the efficiency and stability of data processing. Using a MySQL database to store structured data facilitates data query, management, and analysis. Through the data detection and cleaning steps, anomalies and missing situations in the data can be discovered and processed in a timely manner, ensuring the accuracy and integrity of the data. High-quality data is the basis for subsequent data analysis and mining, helping to improve the reliability and accuracy of analysis results. By capturing user behavior logs and ticket purchase business data, the needs and behavior habits of users can be deeply understood, providing strong support for optimizing the cultural and tourism informatization operation platform.
[0027] Step 120: Use the data sources of the user behavior logs and ticket purchase business data as the data sources of the Kafka cluster.
[0028] Those skilled in the art can understand that as a high-throughput distributed message system, Kafka can receive and process data streams from different data sources in real time. Using the user behavior logs and ticket purchase business data as the data sources of the Kafka cluster can ensure that these data are captured and processed in a timely manner, thus meeting the requirements of real-time data analysis.
[0029] Step 130: Obtain the user behavior logs and ticket purchase business data from the data sources of the Kafka cluster, and use the ClickHouse engine to import them into a preset split-tier data warehouse.
[0030] Those skilled in the art can understand that when using the ClickHouse engine for data import, its efficient data writing ability can be fully utilized to quickly import a large amount of data from the Kafka cluster into the split-level data warehouse, thereby shortening the data processing cycle. The split-level data warehouse usually adopts a hierarchical storage structure, such as the ODS (Operational DataStore) layer, DWD (Data Warehouse Detail) layer, DIM (Dimension) layer, and DM (Data Mining) layer, etc., which can support different data processing and analysis requirements.
[0031] Step 140: Use the ODS layer of the split-level data warehouse to store the imported original user behavior logs and ticket purchase business data; use the preset cleaning algorithm in the DWD layer of the split-level data warehouse to remove the meaningless data in the original user behavior logs and ticket purchase business data.
[0032] In some embodiments, using the preset cleaning algorithm in the DWD layer of the split-level data warehouse to remove the meaningless data in the original user behavior logs and ticket purchase business data can specifically be: Based on the unique user IP in the original user behavior logs as the identifier, and in the order of the log generation time, read the user behavior from the consecutive original user behavior logs; According to the preset chain relationship between user behaviors, determine whether the current user behavior generates the associated user behavior corresponding to the preset chain relationship within the subsequent preset time period. When the current user behavior does not generate the associated user behavior corresponding to the preset chain relationship within the subsequent preset time period, determine the user behavior log corresponding to the current user behavior as meaningless data; According to the unique ticket purchase identifier in the ticket purchase business data, and in the order of the time when the ticket purchase business data is generated, read the ticket purchase business behaviors from the consecutive ticket purchase business data; among them, the ticket purchase business behaviors include: page browsing, page query, information input, and ticket purchase behavior; Determine whether any non-ticket purchase behavior generates a ticket purchase behavior within the subsequent preset waiting time period; when no ticket purchase behavior is generated within the subsequent preset waiting time period, determine the ticket purchase business data corresponding to the current non-ticket purchase behavior as meaningless data; Remove the meaningless data in the original user behavior logs and ticket purchase business data.
[0033] It should be added that the preset chain relationship can be to monitor whether there is a click behavior in a series of subsequent user behaviors within the preset time period after a browsing behavior occurs.
[0034] Step 150: Use the preset demand acquisition interface in the DIM layer of the split-layer data warehouse to obtain query conditions and the names of associated operation interfaces, and extract data that meets the preset query conditions from the user behavior logs and ticket purchase business data in the DWD layer; use the DM layer of the split-layer data warehouse to call the operation program corresponding to the associated operation interface, input the data that meets the preset query conditions into the operation program, and obtain the result data.
[0035] It should be noted that the operation program can be a preset analysis algorithm, and those skilled in the art can adjust the specific content corresponding to the operation program according to actual needs.
[0036] In addition, this application can publish the result data to each reading terminal. The specific process can be: after using the DM layer of the split-layer data warehouse to call the operation program corresponding to the associated operation interface, input the data that meets the preset query conditions into the operation program, and obtain the result data, write the result data into the ClickHouse engine; read the result data through several reading terminals associated with the ClickHouse engine.
[0037] In addition to this, this application Figure 2 is a real-time data processing system for a cultural and tourism big data platform provided by an embodiment of this application. As Figure 2 shown, the system provided by the embodiment of this application mainly includes: An acquisition module 210, configured to capture user behavior logs through page burying points of the cultural and tourism information operation platform; obtain the ticket purchase business data of tourists through the scenic area ticketing system.
[0038] The acquisition module 210 includes a preprocessing unit, configured to buffer the user behavior logs and ticket purchase business data using an Nginx cluster, and write the structured business data into a MySQL database; detect whether there is preset abnormal data and missing data in each piece of data in the MySQL database; remove the data rows with preset abnormal data and missing data, and obtain the processed user behavior logs and ticket purchase business data; Among them, the user behavior logs include user browsing, clicking, staying, commenting, liking, and favorite behaviors; the ticket purchase business data includes structured data such as the current number of ticket purchases, time, age, number of people, and payment.
[0039] A Kafka component 220, configured to use the data sources of the user behavior logs and ticket purchase business data as the data sources of the Kafka cluster; A ClickHouse engine calling module 230, configured to obtain the user behavior logs and ticket purchase business data from the data sources of the Kafka cluster, and use the ClickHouse engine to import them into a preset split-layer data warehouse; The split-layer data warehouse module 240 is used to store the imported original user behavior logs and ticket purchase business data by using the ODS layer of the split-layer data warehouse; use the preset cleaning algorithm in the DWD layer of the split-layer data warehouse to remove the meaningless data in the original user behavior logs and ticket purchase business data; use the preset requirement acquisition interface in the DIM layer of the split-layer data warehouse to obtain the query conditions and the associated operation interface names, and extract the data that meets the preset query conditions from the user behavior logs and ticket purchase business data in the DWD layer; use the DM layer of the split-layer data warehouse to call the operation program corresponding to the associated operation interface, input the data that meets the preset query conditions into the operation program, and obtain the result data.
[0040] The split-layer data warehouse module 240 includes a meaningless identification unit, which is used to identify the user's unique IP in the original user behavior log as the identifier, and read the user behavior from the continuous original user behavior logs in the order of the log generation time; According to the preset chain relationship between user behaviors, determine whether the current user behavior generates the associated user behavior corresponding to the preset chain relationship within the subsequent preset time period. When the current user behavior does not generate the associated user behavior corresponding to the preset chain relationship within the subsequent preset time period, determine that the user behavior log corresponding to the current user behavior is meaningless data; According to the unique ticket purchase identifier in the ticket purchase business data, read the ticket purchase business behavior from the continuous ticket purchase business data in the order of the generation time of the ticket purchase business data; among them, the ticket purchase business behavior includes: page browsing, page querying, information input, and ticket purchase behavior; Determine whether any non-ticket purchase behavior generates a ticket purchase behavior within the subsequent preset waiting time period; when no ticket purchase behavior is generated within the subsequent preset waiting time period, determine that the ticket purchase business data corresponding to the current non-ticket purchase behavior is meaningless data; Remove the meaningless data in the original user behavior logs and ticket purchase business data.
[0041] The system further includes a reading module, which is used to write the result data into the ClickHouse engine; read the result data through several reading terminals associated with the ClickHouse engine.
[0042] The above is the method embodiment in this application. Based on the same inventive concept, the embodiment of this application also provides a real-time data processing device for a cultural and tourism big data platform. As Figure 3 shown, the device includes: a processor; and a memory, on which executable code is stored. When the executable code is executed, the processor executes a real-time data processing method for a cultural and tourism big data platform as described in the above embodiment.
[0043] Specifically, the server side captures user behavior logs through page tagging of the cultural and tourism informatization operation platform, obtains the ticket purchase business data of tourists through the scenic area ticketing system, uses the data sources of the user behavior logs and the ticket purchase business data as the data sources of the Kafka cluster, obtains the user behavior logs and the ticket purchase business data from the data sources of the Kafka cluster, and uses the ClickHouse engine to import them into the preset split-layer data warehouse. The ODS layer of the split-layer data warehouse is used to store the imported original user behavior logs and ticket purchase business data. The preset cleaning algorithm in the DWD layer of the split-layer data warehouse is used to remove the meaningless data in the original user behavior logs and ticket purchase business data. The preset requirement acquisition interface in the DIM layer of the split-layer data warehouse is used to obtain the query conditions and the associated operation interface names, and extract the data that meets the preset query conditions from the user behavior logs and the ticket purchase business data in the DWD layer. The DM layer of the split-layer data warehouse is used to call the operation program corresponding to the associated operation interface, input the data that meets the preset query conditions into the operation program, and obtain the result data.
[0044] In addition, the embodiment of the present application also provides a non-volatile computer storage medium, on which executable instructions are stored. When the executable instructions are executed, the real-time data processing method of a cultural and tourism big data platform as described above is implemented.
[0045] So far, the technical solutions of the present disclosure have been described in combination with multiple embodiments above. However, it is easy for those skilled in the art to understand that the protection scope of the present disclosure is not limited to these specific embodiments. Without departing from the technical principle of the present disclosure, those skilled in the art can split and combine the technical solutions in the above embodiments, and can also make equivalent changes or replacements to the relevant technical features. Any changes, equivalent replacements, improvements, etc. made within the technical concept and / or technical principle of the present disclosure will fall within the protection scope of the present disclosure.
Claims
1. A real-time data processing method for a cultural and tourism big data platform, characterized in that: The method comprises: Capture user behavior logs through page embedding on the cultural tourism information operation platform; obtain tourists’ ticket purchase business data through the scenic spot ticketing system; Use the data source of user behavior logs and ticket purchase business data as the data source of the Kafka cluster; Obtain user behavior logs and ticket purchase business data from the data source of the Kafka cluster, and use the ClickHouse engine to import them into the preset split-layer data warehouse; The ODS layer of the split-layer data warehouse is used to store the imported original user behavior logs and ticket purchase business data; the preset cleaning algorithm in the DWD layer of the split-layer data warehouse is used to remove meaningless data in the original user behavior logs and ticket purchase business data; Utilize the preset demand acquisition interface in the DIM layer of the split-layer data warehouse to obtain the query conditions and the name of the associated operation interface, and extract the data that meets the preset query conditions from the user behavior log and ticket purchase business data in the DWD layer; utilize the DM layer of the split-layer data warehouse to call the operation program corresponding to the associated operation interface, input the data that meets the preset query conditions into the operation program, and obtain the result data.
2. The real-time data processing method of the cultural and tourism big data platform according to claim 1 is characterized in that: After capturing user behavior logs through page embedding of the cultural tourism information operation platform and obtaining tourists' ticket purchase business data through the scenic spot ticketing system, the method further includes: Use the Nginx cluster to buffer user behavior logs and ticket purchase business data, and write structured business data into the MySQL database; Detect whether each data in the MySQL database has preset abnormal data and missing data; remove the data rows with preset abnormal data and missing data, and obtain the processed user behavior log and ticket purchase business data; Among them, user behavior logs include user browsing, clicking, staying, commenting, liking, and collecting behaviors; ticket purchase business data includes structured data stored based on the current number of ticket purchases, time, age, number of people, and payment.
3. The real-time data processing method of the cultural and tourism big data platform according to claim 1 is characterized in that: Use the preset cleaning algorithm in the DWD layer of the split-layer data warehouse to remove meaningless data from the original user behavior logs and ticket purchase business data, including: Based on the unique IP address of the user in the original user behavior log, the user behavior is read from the continuous original user behavior log in the order of the log generation time; According to the preset chain relationship between user behaviors, determine whether the current user behavior generates associated user behaviors corresponding to the preset chain relationship within a subsequent preset time period; when the current user behavior does not generate associated user behaviors corresponding to the preset chain relationship within the subsequent preset time period, determine that the user behavior log corresponding to the current user behavior is meaningless data; According to the unique ticket purchase identifier in the ticket purchase business data, the ticket purchase business behavior is read from the continuous ticket purchase business data in the order of the time when the ticket purchase business data is generated; wherein the ticket purchase business behavior includes: page browsing, page query, information input, and ticket purchase behavior; Determine whether any non-ticket purchasing behavior generates a ticket purchasing behavior within a subsequent preset waiting time period; when no ticket purchasing behavior is generated within the subsequent preset waiting time period, determine that the ticket purchasing business data corresponding to the current non-ticket purchasing behavior is meaningless data; Remove meaningless data from original user behavior logs and ticket purchase business data.
4. The real-time data processing method of the cultural and tourism big data platform according to claim 1 is characterized in that: In the DM layer of the split-layer data warehouse, the operation program corresponding to the associated operation interface is called, data that meets the preset query conditions is input into the operation program, and after the result data is obtained, the method further includes: Write the resulting data into the ClickHouse engine; The result data is read through several reading terminals associated with the ClickHouse engine.
5. A real-time data processing system for a cultural and tourism big data platform, characterized in that: The system comprises: The acquisition module is used to capture user behavior logs through the page embedding of the cultural and tourism information operation platform; and obtain tourists' ticket purchase business data through the scenic spot ticketing system; Kafka component, used to use the data source of user behavior logs and ticket purchase business data as the data source of the Kafka cluster; The ClickHouse engine call module is used to obtain user behavior logs and ticket purchase business data from the data source of the Kafka cluster, and use the ClickHouse engine to import them into the preset split-layer data warehouse; The split-layer data warehouse module is used to utilize the ODS layer of the split-layer data warehouse to store the imported original user behavior logs and ticket purchase business data; utilize the preset cleaning algorithm in the DWD layer of the split-layer data warehouse to remove meaningless data in the original user behavior logs and ticket purchase business data; utilize the preset demand acquisition interface in the DIM layer of the split-layer data warehouse to obtain the query conditions and the names of the associated operation interfaces, and extract the data that meets the preset query conditions from the user behavior logs and ticket purchase business data in the DWD layer; utilize the DM layer of the split-layer data warehouse to call the operation program corresponding to the associated operation interface, input the data that meets the preset query conditions into the operation program, and obtain the result data.
6. The real-time data processing system of the cultural and tourism big data platform according to claim 5 is characterized in that: The acquisition module includes a preprocessing unit, Used to buffer user behavior logs and ticket purchase business data using the Nginx cluster, and write structured business data into the MySQL database; Check whether each piece of data in the MySQL database contains preset abnormal data and missing data; Remove data rows with preset abnormal data and missing data to obtain processed user behavior logs and ticket purchase business data; Among them, user behavior logs include user browsing, clicking, staying, commenting, liking, and collecting behaviors; ticket purchase business data includes structured data stored based on the current number of ticket purchases, time, age, number of people, and payment.
7. The real-time data processing system of the cultural and tourism big data platform according to claim 5 is characterized in that: The split-layer data warehouse module includes a meaningless identification unit, Used to read user behavior from continuous original user behavior logs based on the user's unique IP in the original user behavior log and in the order of log generation time; According to the preset chain relationship between user behaviors, determine whether the current user behavior generates associated user behaviors corresponding to the preset chain relationship within a subsequent preset time period; when the current user behavior does not generate associated user behaviors corresponding to the preset chain relationship within the subsequent preset time period, determine that the user behavior log corresponding to the current user behavior is meaningless data; According to the unique ticket purchase identifier in the ticket purchase business data, the ticket purchase business behavior is read from the continuous ticket purchase business data in the order of the time when the ticket purchase business data is generated; wherein the ticket purchase business behavior includes: page browsing, page query, information input, and ticket purchase behavior; Determine whether any non-ticket purchasing behavior generates a ticket purchasing behavior within a subsequent preset waiting time period; when no ticket purchasing behavior is generated within the subsequent preset waiting time period, determine that the ticket purchasing business data corresponding to the current non-ticket purchasing behavior is meaningless data; Remove meaningless data from original user behavior logs and ticket purchase business data.
8. The real-time data processing system of the cultural and tourism big data platform according to claim 5 is characterized in that: The system also includes a reading module, Used to write result data into the ClickHouse engine; read result data through several read terminals associated with the ClickHouse engine.
9. A real-time data processing device for a cultural and tourism big data platform, characterized in that: The device comprises: processor; and a memory having executable codes stored thereon, which, when the executable codes are executed, enable the processor to execute a real-time data processing method for a cultural and tourism big data platform as described in any one of claims 1-4.
10. A non-volatile computer storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed, they implement a real-time data processing method for a cultural and tourism big data platform as described in any one of claims 1-4.
Citation Information
Cited By
Dual-mode big data platform for text, travel and body industry
CN122240719A