A Real-time Data Update and Management Method Based on Spark Streaming

Through real-time data update and management methods based on Spark Streaming, the problem of complex and time-consuming deployment of existing real-time data warehouses when adding Kafka topics is solved, real-time and automated processing of Kafka topic updates are realized.

CN113590667BActive Publication Date: 2025-05-27SHENZHEN SEN5 TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110600651.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-31
Publication Date
2025-05-27
Estimated Expiration
2041-05-31

AI Technical Summary

Technical Problem

The existing real-time data warehouse needs to be redeployed and launched when adding Kafka topics. There are problems such as environmental uncertainty and diversity of component sources, which leads to complex and time-consuming deployment, and it is impossible to ensure the real-time nature of Kafka topic updates.

Method used

Real-time data update and management methods based on Spark Streaming are adopted to realize automated processing and data synchronization by configuring resource parameters, establishing metadata information database, analyzing Kafka data source parameters, reading and updating metadata information, modifying Hive data information, and parsing Kafka data into database tables.

Benefits of technology

The Kafka topic update process is simplified, the deployment risk and time is reduced, the real-time update of Kafka topic is achieved, and the efficiency and reliability of real-time data processing is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113590667B_ABST
    Figure CN113590667B_ABST
Patent Text Reader

Abstract

The present invention proposes a real-time data update and management method based on Spark Streaming, including: performing parameter configuration of configuration resources; establishing a metadata information library table; parsing the source parameters of Kafka data to obtain real-time data updates; reading the metadata information in the metadata information library, including reading the description information of Kafka data in the metadata information library, and reading after updating the corresponding metadata in the metadata information library; modifying the hive data information; reading Kafka data, partitioning the read batches of Kafka data, and parsing and mapping the Kafka data into database tables according to the partitions. This real-time data storage and management method only needs to modify the metadata information for new tasks, create a new hive table, and Spark Streaming synchronizes the metadata information to obtain new and changed data, parses the Kafka data one by one to correspond to the data types of hive, writes the data into hive, and updates the offset information of the corresponding data at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet technology, and in particular to a real-time data updating and management method based on Spark Streaming. Background Art

[0002] "Data Intelligence" has an essential and basic link, which is the construction of a data warehouse. At the same time, a data warehouse is also a basic service that a company will inevitably provide after its data grows to a certain scale. From the perspective of smart business, the results of the data represent the feedback from users or devices, and the timeliness of obtaining the results is particularly important. Rapidly obtaining data feedback can help companies make decisions faster and iterate products better. Real-time data warehouses play an irreplaceable role in this process.

[0003] The real-time data warehouse established in the existing technology mainly performs real-time ETL (Extraction-Transformation-Loading) on ​​traffic data, does not calculate real-time indicators, and has not established a real-time data warehouse system. The real-time scenario is relatively simple, and the processing of real-time data streams is to improve the service capabilities of the data platform. The processing of real-time data depends on the collection of data upward and is related to the query and visualization of data downward. The first part is data collection, which is collected by the three-end SDK and sent to Kafka through the Log Collector Server. The second part is data ETL, which is used to complete the cleaning and processing of the original data and import it into Druid in real time and offline. The third part is data visualization, which is responsible for calculating indicators by Druid and completing data visualization through the Web Server in conjunction with the front end.

[0004] The above real-time data warehouse has the problem that adding new Kafka topics requires redeploying new codes. However, due to environmental uncertainty and the diversity of component sources, software deployment is very risky. During the deployment process, attention should be paid to the dependencies and coordination between builds, the variability of the deployment process, and the integration and security of the Internet. The deployment process is complex, difficult, and time-consuming. In addition, the real-time data warehouse has many new requirements for Kafka topics during the application process. It can be seen that the shortcomings of the existing real-time data warehouse will lead to the inability of the real-time data warehouse to guarantee the real-time update of Kafka topics in today's complex real-time scenarios and fast data stream updates. Summary of the invention

[0005] In view of this, in order to overcome the above-mentioned defects of the prior art, the present invention proposes a real-time data updating and management method based on Spark Streaming.

[0006] Specifically, the real-time data update and management method based on Spark Streaming includes:

[0007] Perform parameter configuration of configuration resources;

[0008] Establish metadata information base table;

[0009] Parse the source parameters of Kafka data;

[0010] Reading metadata information in a metadata information repository, including reading description information of the Kafka data in the metadata information repository, and updating corresponding metadata in the metadata information repository before reading;

[0011] Modify hive data information;

[0012] The Kafka data is read, the read batches of the Kafka data are partitioned, and the Kafka data are parsed and mapped into database tables according to the partitions.

[0013] Furthermore, the method further includes the creation of a metadata information base: the data forwarding center receives and classifies the data generated by the device terminal, and then creates the metadata information base according to the pre-established metadata information base table;

[0014] The “creating a metadata information base according to the pre-established metadata information base table” includes:

[0015] Read hive table data, obtain the schema information of the data, and write the schema information into the hive table field description information table in the metadata information library;

[0016] Read all data in the hive table field description information table whose database name is equal to the database name and table name when the parameter configuration of the configuration resource is executed;

[0017] Loop through each piece of data, match it according to the corresponding fieldType type, build a single StructField data according to the matching rules, and then assemble the Struct Schema information corresponding to the entire table.

[0018] The “updating the corresponding metadata in the metadata information library and then reading” includes:

[0019] The metadata information base is updated, and then the corresponding metadata information is read and broadcasted as a broadcast variable;

[0020] The updating of the broadcast variables includes:

[0021] Get the current system execution time; get the last hive schema update time;

[0022] If the remainder of the current system time divided by 10 is greater than 1 and the current system time minus the last hive schema update time is greater than 10 minutes, perform hive schema information update;

[0023] Obtain the SQL statement related to the task in the modification hive data structure description information table in the latest metadata information library and execute the SQL statement;

[0024] Update the field information of the corresponding table in the Kafka data description information table and update the field type information of the corresponding table in the hive table field description information table;

[0025] Obtain updated data in the Kafka data description information table and the hive table field description information table;

[0026] Clear the Spark broadcast data and then reassign the Spark broadcast data.

[0027] The "modify hive data information" includes: modifying the structural description information of the hive data; and determining whether the hive data meets the update conditions, and executing the previous data update operation when the hive data meets the update conditions; when the hive data information does not meet the update conditions, executing the subsequent reading of the Kafka data parsing and mapping operation. Among them, the update conditions of the hive data are: the current system execution time is ten minutes apart from the previous system execution time, and the hive data structure description information table is not empty.

[0028] The “mapping the Kafka data into database tables according to the partitions” includes:

[0029] Get all the data of a certain partition;

[0030] Construct an array with the same length as the hive field information;

[0031] Parse the partitioned data in Json format to obtain hash type data, and put it into the corresponding position in the array according to the position of the field in hive;

[0032] Serialize array data;

[0033] The partitioned data generates RDD[Row];

[0034] The RDD[Row] and struct schema information generated by the partitioned data are constructed into Spark's DataFrame data, and then written into the corresponding table in the hive database.

[0035] The “constructing an array array with the same length as the hive field information; performing Json parsing on the partition data to obtain hash type data, and placing the data into the corresponding position in the array according to the position of the field in hive; and serializing the array data” further includes:

[0036] Parse the Json data into a hashmap structure, obtain the corresponding data according to the key and record it as a hashmap;

[0037] Get the data structure of all fields of the corresponding table from the Hive table field description information table, and combine them into a hashmap data record as hashmapSchema;

[0038] The array constructed by sequential field information of the hive table creation statement recorded in the Kafka data description information table is recorded as array;

[0039] Create an array with the same data length as array and record it as arrayValue;

[0040] Loop through each value in the array, get the value corresponding to the key in the hashmap, get the value corresponding to the key in the hashmapSchema, convert the value obtained in the hashmap into data of the data type obtained in the hashmapSchema, and then assign it to the data at the same position in the arrayValue at this time;

[0041] Convert arrayValue to serialized Row type.

[0042] Preferably, it also includes: after each partition data is processed, determining whether the Kafka data of the batch is processed;

[0043] If the judgment result is yes, the next batch of data is read for processing; if the judgment result is no, the remaining partition data is read for processing.

[0044] The present invention also provides a real-time data storage and management method system, which is used to perform data processing of the real-time data storage and management method, including:

[0045] Parameter configuration module: used to configure the parameters of resources in the visualization panel;

[0046] Metadatabase building module: used to build metadatabase information;

[0047] Parameter parsing module: used to parse Kafka data source parameters;

[0048] Metadata acquisition module: used to read metadata information in the metadata information library;

[0049] Hive data information modification module: used to modify hive data structure information and related data;

[0050] Kafka data batch processing module: used to read the Kafka data and related information, partition the read Kafka data according to the batch of data acquisition, and parse and map the data according to the partitions.

[0051] The Kafka data batch processing module also includes:

[0052] Kafka data parsing and mapping module: used to parse and map each partition of a batch of Kafka data into a database table.

[0053] Preferably, it also includes:

[0054] Judgment module: used to judge whether the batch data has been processed after the data of one partition has been processed: if the judgment result is yes, read the next batch of data for processing; if the judgment result is no, obtain the data of the remaining partitions of the batch data for processing.

[0055] In summary, the real-time data storage and management method of the present invention synchronizes data to the metadata information library after adding fields, adding or deleting kafkatopic, and Spark Streaming synchronizes metadata information to obtain data additions and changes, and then automatically adds the newly added data during parsing to achieve automatic perception and automated processing. The real-time data storage and management method only needs to modify the metadata information and create a new hive table when adding a new task. The code will automatically update the hive metadata information in the metadata information library, parse the Kafka data one by one to the corresponding hive data type, write the data to hive, and update the offset information of the corresponding data at the same time. By simply configuring the Kafka data and hive data mapping information, the testing, development, and deployment of online tasks can be quickly realized. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0057] Figure 1 A schematic diagram of a metadata update architecture of a real-time data update and management method based on Spark Streaming of the present invention;

[0058] Figure 2 A schematic diagram for describing the data processing flow of the real-time data update and management method based on Spark Streaming of the present invention;

[0059] Figure 3 A structural schematic diagram of a real-time data storage and management method system using a data processing flow of the real-time data update and management method based on Spark Streaming of the present invention;

[0060] Figure 4 The present invention is a schematic diagram of applying the real-time data updating and management method based on Spark Streaming to a computer device.

[0061] Reference numerals:

[0062] 1-computer device; 11-processing unit; 12-system memory; 13-bus; 2-external devices. DETAILED DESCRIPTION

[0063] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0064] The present invention provides a real-time data update and management method based on Spark Streaming, and the system includes two main modules: a data acquisition module and a data processing module. The data acquisition module reads multiple topic data of Kafka data and writes them into a metadata information library, such as hive, ES, Druid, MySQL, etc. The data processing module can parse and map the complex data types of most Kafka data into database tables, read multiple Kafka data at a time, and save a lot of machine processing resources. When adding new tasks online, it only needs to simply initialize metadata resources, and the real-time task scheduling information can be quickly deployed online without any modification to the code.

[0065] The real-time data storage and management method of the present invention includes: the data forwarding center (Skyway IOT) is used to receive and classify the data generated by the device terminal: the data forwarding center verifies the authority of the device terminal to ensure that the device terminal has the authority to report data, and the data sent by the device terminal that passes the verification is received by the data forwarding center, and the received data is classified and the same data is written into the same topic. First, a metadata information library is created according to the pre-established metadata information library table, and then the full amount of data of the previously created metadata information library is read and broadcast as a broadcast variable. At this time, the program starts to read Kafka data, and the read data will have separate topic information, and different topic data will be processed differently.

[0066] See the instruction manual Figure 1 , which is a schematic diagram of metadata update. The data forwarding center interacts with the database by first creating a metadata repository based on the table corresponding to the pre-created hive database, and then reading all the data in the previously created metadata repository and broadcasting it as a broadcast variable.

[0067] The metadata information includes four tables, among which sdx_hive_table_schema_dec records the hive table field description information. Its data structure is shown in Table 1:

[0068] name type length Decimal Point Not null virtual key Notes db varchar 255 0 √ □1 Database Name tb varchar 255 0 √ □2 Table name fieldname varchar 255 0 √ □3 Field Name fieldtype varchar 1000 0 √ Field Type

[0069] Table 1 sdx_hive_table_schema_dec data structure

[0070] sdx_Kafka_hive_eventname records the description information of the Kafka topic consumed by Spark. Its data structure is shown in Table 2:

[0071]

[0072] Table 2 sdx_Kafka_hive_eventname data structure

[0073] sdx_partition_offset records the offset of the topic consumed by Spark from Kafka. Its data structure is shown in Table 3:

[0074] name type length Decimal Point Not null virtual key Notes topic varchar 128 0 √ partitionid int 16 0 √ offset varchar 255 0 √ createtime datetime 0 0 √ consumergroup varchar 255 0 √ MySQLstatus int 16 0 √ hivestatus int 16 0 eventname varchar 255 0

[0075] Table 3 sdx_partition_offset

[0076] The sdx_hive_table_add_columns table is used to modify the description information of the hive data structure.

[0077] Example 1

[0078] This embodiment provides a data processing flow of a real-time data update and management method based on Spark Streaming. Before the solution is executed, it is necessary to first configure the resource parameters in the visualization panel. The specific configuration information includes:

[0079] a. Task name: that is, the event name. This information corresponds to the eventname information in the data, that is, the values ​​are the same;

[0080] b. Kafka data source: Kafka topic name is the name of the Kafka topic;

[0081] c. Number of Kafka partitions: the number of partitions corresponding to the Kafka topic;

[0082] d. Target database: db target database name;

[0083] e. Target table: tb target table name

[0084] f.isevent: whether it is a Kafka data stream, 1 for yes, 0 for no.

[0085] When the configuration information is saved, a piece of information will be inserted into the metadata information table sdx_Kafka_hive_eventname synchronously. At this time, the fieldnames field information is empty. For example: INSERT INTO `sage_task_metadata`.`sdx_kafka_hive_eventname`(`eventname`,`db`,`tb`, `topicname`,`partitions`,`isevent`)VALUES('app_installation','ods_sdx_safe','ods_sdx_app_installation','sdb_sdx_app_installation',1,0).

[0086] Create hive table. Example: create table ods_sdx_safe.ods_sdx_app_installation(idint,name string)PARTITIONED BY(subregion string).

[0087] See the instruction manual Figure 2 , is a schematic diagram describing the data processing flow of the real-time data storage and management method of this embodiment. Specifically, the data processing flow of the real-time data storage and management method includes the following steps:

[0088] S1: Start the Spark Streaming program and perform initial configuration: initialize SparkSession and initialize the configuration parameters in the configuration resource conf. In this embodiment, the dynamic resource configuration is as follows:

[0089] --conf Spark.Streaming.dynamicAllocation.enabled=true\

[0090] --conf Spark.Streaming.dynamicAllocation.minExecutors=1\

[0091] --conf Spark.Streaming.dynamicAllocation.maxExecutors=6\

[0092] S2: Parse Kafka data source parameters. In this embodiment, the parameters required for task execution include:

[0093] args(0)SparkStreaming seconds

[0094] args(1)consumerGroup

[0095] args(2)MySQLEnvironment

[0096] args(3)hiveEnvironment

[0097] Among them, SparkStreaming second: defines how often to process the data stream to create a StreamingContext object; consumerGroup: the same system can use different consumer groups to consume the same data to distinguish between online testing, development, and production; MySQLEnvironment: different platforms deploy and rely on inconsistent external metadata environments. To specify the environment through parameters, you only need to enter the environment identification value; hiveEnvironment: determine which database the data enters based on the specified environment parameters.

[0098] S3: Read metadata information in the metadata repository, including: reading description information of Kafka data in the metadata repository, and obtaining it after updating the corresponding metadata in the metadata repository.

[0099] In this embodiment, it specifically includes:

[0100] S31: Update the metadata information base, then read the corresponding metadata information and broadcast it as a broadcast variable. The specific execution process is: automatically obtain the data in sdx_Kafka_hive_eventname where isevent is 1 and eventname is app_installation, and do the following:

[0101] S311: Delete all data under the same db and tb in sdx_hive_table_schema_dec;

[0102] S312: Re-append the data to sdx_hive_table_schema_dec according to the structure information of the corresponding table in hive. At this time, three data are inserted into the table; namely:

[0103] INSERT INTO`sage_task_metadata`.`sdx_hive_table_schema_dec`(`db`,`tb`, `fieldname`,`fieldtype`)

[0104] VALUES('ods_sdx_safe','ods_sdx_app_installation','id','IntegerType');

[0105] INSERT INTO`sage_task_metadata`.`sdx_hive_table_schema_dec`(`db`,`tb`, `fieldname`,`fieldtype`)

[0106] VALUES('ods_sdx_safe','ods_sdx_app_installation','name','StringType');

[0107] INSERT INTO`sage_task_metadata`.`sdx_hive_table_schema_dec`(`db`,`tb`, `fieldname`,`fieldtype`)

[0108] VALUES('ods_sdx_safe','ods_sdx_app_installation','subregion','StringType');

[0109] S313: Update the fieldnames value in sdx_kafka_hive_eventname. At this time, the value becomes: id, name, subregion.

[0110] S32: Read the updated Kafka data description information, that is, read the latest sdx_Kafka_hive_eventname information;

[0111] S33: Read the offset information of Kafka data, that is, read the historical Kafka offset information in sdx_partition_offset; if the partition offset data does not exist, consumption will start from the earliest data of the partition, that is, read the initial data.

[0112] S4: Modify hive data information:

[0113] S41: Modify the structural description information of hive data: The specific execution process is to modify the property panel including db (data source library), tb (data source table), fieldname (new field name), fieldtype (new field type). At this time, choose to add a new field to the ods_sdx_app_installation table.

[0114] This will trigger the insertion of a message into the hive data structure description information table (sdx_hive_table_add_columns).

[0115] That is: INSERT INTO`sage_task_metadata`.`sdx_hive_table_add_columns`(`db`,`tb`,`sql`)VALUES('ods_sdx_safe','ods_sdx_app_installation','alter table ods_sdx_app_installation add columns(city string)');

[0116] S42: Determine whether the hive metadata meets the update conditions:

[0117] When the hive metadata meets the update conditions (once every ten minutes and the sdx_hive_table_add_columns table is not empty), the previous data update operation is performed, that is, updating the configuration information in the visualization panel, creating the hive table, and reading the metadata information in the metadata information library. At this time, the ods_sdx_app_installation in the database will trigger the execution of the SQL statement and add the city field.

[0118] If the hive metadata information does not meet the update conditions, perform subsequent operations.

[0119] S5: Read Kafka data and process the batch of data. Reading Kafka data includes: Spark reads a single Kafka topic data or multiple Kafka data. Processing the batch of data includes: obtaining data from a specified offset position of Kafka, obtaining information about the batch of data, and partitioning the batch of data according to topic.

[0120] S501: Obtain the offest information of the batch data and the metadata information of the hive table, including: all offset information, the fields string information composed of the hive table fields (the information is consistent with the order of the fields in the hive table creation statement), and the Struct Schema information constructed by the hive table field information in the table creation order.

[0121] S502: The batch of data is partitioned according to the topic name, and after the partitioning, the data in the same partition is the data of the same topic. The Kafka data is parsed and mapped into a database table according to the partition.

[0122] S6: Build a thread pool to process each partition of the batch data and parse and map the Kafka data into a database table. The thread pool can be used to process data quickly.

[0123] S601: Acquire all data of a partition of the batch data, and acquire corresponding hive struct schema and hive fields information according to the partition field of the partition.

[0124] S602: Process a single task in the thread pool. Specifically including:

[0125] Construct an array with the same length as the hive field information;

[0126] Parse the partitioned data in Json format to obtain hash type data, and put it into the corresponding position in the Array according to the position of the field in Hive.

[0127] Serialize array data;

[0128] The partition data generates RDD[Row];

[0129] The RDD[Row] and struct schema information generated by the partitioned data are constructed into Spark's DataFrame data, and then written into the corresponding table in the Hive database;

[0130] Write the partition's offset information into MySQL as historical information to ensure data consistency.

[0131] S7: After processing the data of one partition, determine whether the batch of data has been processed. If the judgment result is yes, read the next batch of data for processing; if the judgment result is no, obtain the data of the remaining partitions of the batch of data for processing, forming a batch loop processing in the thread pool.

[0132] Example 2

[0133] This embodiment provides a solution for automatically constructing Scala data types in Hive Schema, which is used to implement "creating a metadata information library based on a pre-established metadata information library table". The solution execution process of this embodiment includes:

[0134] (1) Read hive table data through Spark, then obtain the schema information of the data, and write the schema information into the hive table field description information table in the metadata information library, that is, into the sdx_hive_table_schema_dec table;

[0135] (2) Read all data in the hive table field description information table (sdx_hive_table_schema_dec) where the database name (db) is equal to the database name and table name (tb) when the parameter configuration of the configuration resource is executed;

[0136] (3) Loop through each piece of data, match it according to the corresponding fieldType type, build a single StructField data according to the matching rules, and then assemble the Struct Schema information corresponding to the entire table.

[0137] For example, if there are db, tb, name, and StringType data in sdx_hive_table_schema_dec, StructField(name, StringType, true) will be generated after mapping.

[0138] Example 3

[0139] This embodiment provides a solution for constructing DateFrame from Kafka Json, which is used to implement the parsing and mapping of Kafka data into a database table. Specifically, Kafka Json can be used to construct DataFrame.

[0140] The solution execution process of this embodiment includes:

[0141] (1) The Kafka Json data format is {"key1":"value1","key2":"value2","key13":"value3"}. At this time, the Json data is parsed into a hashmap structure, and the corresponding data can be obtained according to the key. This data is recorded as a hashmap;

[0142] (2) The data structure of all fields of the corresponding table can be obtained from the hive table field description information table (sdx_hive_table_schema_dec), and the data is combined into a hashmap data recorded as hashmapSchema;

[0143] (3) The Kafka data description information table (sdx_Kafka_hive_eventname) records the sequential field information of the hive table creation statement, which is cut into an array by commas and recorded as array;

[0144] (4) Create a new array with the same data length as array, recorded as arrayValue;

[0145] (5) Loop through each value in the array, get the value corresponding to the key in the hashmap (the key value here is the same as the value looped in the array), get the value corresponding to the key in the hashmapSchema (the key value here is the same as the value looped in the array), convert the value obtained in the hashmap into data of the data type obtained in the hashmapSchema, and then assign it to the data at that position in the arrayValue at this time;

[0146] (6) Convert arrayValue to serialized Row type;

[0147] (7) Process the RDD data to obtain an RDD[Row];

[0148] (8) You can construct a Spark DataFrame by combining RDD[Row] and the StructType corresponding to the table.

[0149] Example 4

[0150] This embodiment provides a Hive Schema information dynamic update solution for updating broadcast variables. The solution execution process of this embodiment includes:

[0151] (1) Get the current system execution time; get the last hive schema update time;

[0152] (2) If the remainder of the current system time divided by 10 is greater than 1 and the system current time minus the last update time is greater than 10 minutes, execute the hive schema information update;

[0153] (3) Obtain the SQL statement related to the task in the modification hive data structure description information table (sdx_hive_table_add_columns) in the latest metadata information library and execute the statement;

[0154] (4) Update the field information of the corresponding table in the Kafka data description information table (sdx_Kafka_hive_eventname) and update the field type information of the corresponding table in the hive table field description information table (sdx_hive_table_schema_dec);

[0155] (5) Get the updated data in sdx_Kafka_hive_eventname and sdx_hive_table_schema_dec;

[0156] (6) Clear the Spark broadcast data and reassign the Spark broadcast data;

[0157] (7) At this time, the hive schema is updated to the latest version.

[0158] Example 5

[0159] This embodiment provides a real-time data storage and management system that uses the data processing flow of the real-time data update and management method based on Spark Streaming provided in Embodiment 1.

[0160] See the instruction manual Figure 3 , real-time data storage and management system includes:

[0161] Parameter configuration module: used to configure resource parameters in the visualization panel. Specific configuration information includes: task name, Kafka data source, number of Kafka partitions, target database, target table, and isevent.

[0162] Metadatabase building module: used to establish metadatabase information, including: establishing a metadata information base table, and creating a metadata information base according to the metadata information base table;

[0163] Parameter parsing module: used to parse Kafka data source parameters.

[0164] Metadata acquisition module: used to read metadata information in the metadata information library. Including:

[0165] Read the description information of Kafka data in the metadata repository, update the corresponding metadata in the metadata repository, and then obtain the corresponding metadata and broadcast it as a broadcast variable.

[0166] Hive data information modification module: used to modify hive data structure information and related data;

[0167] It also includes a selection module: when the hive metadata information meets the update conditions, the data update operation is performed, and the information is inserted into the hive data structure description information table (sdx_hive_table_add_columns); when the hive metadata information does not meet the update conditions, the Kafka data is read for processing.

[0168] Kafka data batch processing module: used to read Kafka data and related information, partition the read Kafka data according to the batch of data acquisition, and parse and map the data according to the partition. The Kafka data batch processing module includes the Kafka data parsing and mapping module: used to build a thread pool to parse and map each partition of a batch of Kafka data into a database table.

[0169] Judgment module: used to judge whether the batch of data has been processed after the data of one partition has been processed: if the judgment result is yes, read the next batch of data for processing; if the judgment result is no, obtain the data of the remaining partitions of the batch of data for processing.

[0170] Example 6

[0171] Figure 4 A schematic diagram of the structure of a computer device provided in this embodiment. Figure 4 The computer device 1 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0172] The computer device 1 is embodied in the form of a general-purpose computing device. The components of the computing device include: one or more processors or processing units 11, system memory 12, and a bus 13 connecting different system components (including system memory 12 and processing unit 11). The computing device 1 typically includes a variety of computer system readable media, which can be any available media that can be accessed by the device computer, including volatile and non-volatile media, removable and non-removable media. The system memory 12 may include computer system readable media in the form of volatile memory 12.

[0173] The computer device 1 may also communicate with one or more external devices 2 (e.g., keyboards, pointing devices, displays, etc.), with one or more devices that enable a user to interact with the computer, and / or with any device that enables the computer device 1 to communicate with one or more other computing devices.

[0174] The processing unit executes various functional applications and data processing by running the images stored in the system memory, for example, implementing a data processing flow of a real-time data update and management method based on Spark Streaming provided in Example 1, including:

[0175] Updates the configuration information in the Visualization Panel.

[0176] Create metadata repository tables.

[0177] Initializes the configuration parameters of the configuration resource.

[0178] Parse the Kafka data source parameters.

[0179] Read metadata information in the metadata repository. This includes: reading description information of Kafka data in the metadata repository; and updating corresponding metadata in the metadata repository, and then obtaining corresponding metadata and broadcasting it as a broadcast variable.

[0180] Modify hive data information. Including: modify attribute information and insert information into the hive data structure description information table (sdx_hive_table_add_columns), and determine whether hive data information meets the update conditions: execute data update operation when hive metadata information meets the update conditions, and read Kafka data for processing when hive metadata information does not meet the update conditions.

[0181] Read Kafka data and process the batch of data read. This includes: reading Kafka data and related information, partitioning the read Kafka data according to the batch of data obtained, and passing the data to the thread pool for parsing and mapping according to the partition.

[0182] Build a thread pool to parse and map each partition of a batch of Kafka data into a database table.

[0183] After processing the data of one partition, determine whether the batch of data has been processed: if the judgment result is yes, read the next batch of data for processing; if the judgment result is no, obtain the data of the remaining partitions of the batch of data for processing.

[0184] This embodiment applies a data processing flow of a real-time data update and management method based on Spark Streaming to a specific computer device, stores the method in a memory, and when an executor executes the memory, runs the processing flow to process information of the real-time data storage and management method. The method is quick and convenient to use and has a wide range of applications.

[0185] Of course, those skilled in the art will appreciate that the processor may also implement the data processing solution of the real-time data storage and management method provided by any embodiment of the present invention.

[0186] Example 7

[0187] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the data processing flow of the real-time data update and management method based on Spark Streaming provided in any embodiment of the present invention is implemented, including:

[0188] Updates the configuration information in the Visualization Panel.

[0189] Create metadata repository tables.

[0190] Initializes the configuration parameters of the configuration resource.

[0191] Parse the Kafka data source parameters.

[0192] Read metadata information in the metadata repository. This includes: reading description information of Kafka data in the metadata repository; and updating corresponding metadata in the metadata repository, and then obtaining corresponding metadata and broadcasting it as a broadcast variable.

[0193] Modify hive data information. Including: modify attribute information and insert information into the hive data structure description information table (sdx_hive_table_add_columns), and determine whether hive data information meets the update conditions: execute data update operation when hive metadata information meets the update conditions, and read Kafka data for processing when hive metadata information does not meet the update conditions.

[0194] Read Kafka data and process the batch of data read. This includes: reading Kafka data and related information, partitioning the read Kafka data according to the batch of data obtained, and passing the data to the thread pool for parsing and mapping according to the partition.

[0195] Build a thread pool to parse and map each partition of a batch of Kafka data into a database table.

[0196] After processing the data of one partition, determine whether the batch of data has been processed: if the judgment result is yes, read the next batch of data for processing; if the judgment result is no, obtain the data of the remaining partitions of the batch of data for processing.

[0197] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.

[0198] This embodiment applies the data processing flow of the real-time data update and management method based on Spark Streaming to a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the eyelash segmentation method with an adaptive threshold provided by the present invention is implemented, which is simple, fast, easy to store, and not easy to lose.

[0199] In summary, the real-time data update and management method based on Spark Streaming provided by the present invention synchronizes data to the metadata information library after adding fields, adding or deleting kafka topics through the front-end interface. Spark Streaming synchronizes metadata information to obtain data additions and changes, and then automatically adds the newly added data during parsing to achieve automatic perception and automated processing. The real-time data storage and management method only needs to modify the metadata information and create a new hive table when adding a new task. The code will automatically update the hive metadata information in the metadata information library, parse the Kafka data one by one to the corresponding hive data type, write the data into hive, and update the offset information of the corresponding data at the same time. By simply configuring the Kafka data and hive data mapping information, the testing, development, and deployment of online tasks can be quickly implemented.

[0200] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. In addition to the above embodiments, there may also be different variations. The technical features of the above embodiments may be combined with each other. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A real-time data update and management method based on Spark Streaming, characterized in that, the real-time data storage and management method includes: Performing parameter configuration of configuration resources; Establishing a metadata information library table; Parsing the source parameters of Kafka data to obtain real-time data updates; Reading the metadata information in the metadata information library, including reading the description information of the Kafka data in the metadata information library, and reading after updating the corresponding metadata in the metadata information library; Modifying hive data information; Reading the Kafka data, partitioning the read batches of the Kafka data, and parsing and mapping the Kafka data into a database table according to the partition; The reading after updating the corresponding metadata in the metadata information library includes: Updating the metadata information library, then reading the corresponding metadata information, and broadcasting it as a broadcast variable; Among them, the update of the broadcast variable includes: Obtaining the current system execution time; obtaining the previous hive schema update time; If the remainder of dividing the minute number of the current system time by 10 is greater than 1 and the current system time minus the previous hive schema update time is greater than 10 minutes, then perform hive schema information update; Obtaining the sql statements related to the tasks in the modified hive data structure description information table in the current latest metadata information library and executing the sql statements; Updating the field information of the corresponding table in the Kafka data description information table and updating the field type information of the corresponding table in the hive table field description information table; Obtaining the updated data in the Kafka data description information table and the hive table field description information table; Clearing the data of Spark broadcast, and then re-assigning the data of Spark broadcast; The parsing and mapping the Kafka data into a database table according to the partition includes: Obtaining all the data of a certain partition; Constructing an array array as long as the hive field information; Performing Json parsing on the data of the partition to obtain hash type data, and putting it into the corresponding position in the array according to the position of the field in hive; Serializing the array data; Generating RDD[Row] from the data of the partition; The RDD[Row] generated from the data of the partition and the struct schema information are constructed into a Spark DataFrame data, and then written into the corresponding table in the hive database.

2. The real-time data update and management method according to claim 1, characterized in that, It also includes the creation of a metadata information library: the data forwarding center accepts the data generated by the device terminal and classifies it, and then creates the metadata information library according to the pre-established metadata information library table; Among them, the creating the metadata information library according to the pre-established metadata information library table includes: Read the Hive data, then obtain the schema information of the data, and write the schema information into the Hive table field description information table in the metadata information library; Read all the data in the Hive table field description information table where the database name is equal to the database name and the table name is equal to the table name when executing the parameter configuration of the configuration resource; Loop through each piece of data, make a match according to the corresponding fieldType type, form a single StructField data according to the matching rule, and then form the Struct Schema information corresponding to the entire table.

3. According to the real-time data update and management method described in claim 1, characterized in that, the modification of the Hive data information includes: including: modifying the structure description information of the Hive data; and, judging whether the Hive data meets the update condition. When the Hive data meets the update condition, perform the previous data update operation; when the Hive data information does not meet the update condition, perform the subsequent operation of reading and parsing and mapping the Kafka data.

4. According to the real-time data update and management method described in claim 3, characterized in that, the update condition of the Hive data is: the interval between the current system execution time and the previous system execution time is ten minutes, and the Hive data structure description information table is not empty.

5. According to the real-time data update and management method described in claim 4, characterized in that, construct an array with the same length as the Hive field information; perform Json parsing on the data of this partition to obtain hash type data, and put it into the corresponding position in the Array according to the position of this field in Hive; The serialization of the array data further includes: parse the Json data into a hashmap structure, and obtain the corresponding data according to the key and record it as hashmap; obtain the data structures of all fields of the corresponding table from the Hive table field description information table, and combine them into a hashmap data and record it as hashmapSchema; construct an array of the sequential field information of the Hive table creation statement recorded in the Kafka data description information table and record it as array; create an array with the same data length as array and record it as arrayValue; loop through each value in array, obtain the value corresponding to the key in hashmap, obtain the value corresponding to the key in hashmapSchema, convert the value obtained in hashmap into the data of the data type obtained in hashmapSchema, and then assign it to the data in the same position in arrayValue at this time; convert arrayValue into a serialized Row type.

6. According to the real-time data update and management method described in claim 1, characterized in that, further includes: after processing the data of each partition, judge whether the Kafka data of this batch has been processed; if the judgment result is yes, read the next batch of data for processing; If the judgment result is negative, read the remaining partition data for processing.

7. A real-time data update and management method system for performing data processing of the real-time data update and management method described in any one of claims 1-5. Characterized in that Comprising: Parameter configuration module: used to configure the parameters of resources on the visualization panel; Meta-database construction module: used to establish meta-database information; Parameter parsing module: used to parse Kafka data source parameters; Meta-data acquisition module: used to read meta-data information from the meta-data information library; Hive data information modification module: used to modify Hive data structure information and related data; Kafka data batch processing module: used to read the Kafka data and related information, partition the read Kafka data according to the batches of data acquisition, and perform parsing and mapping processing on the data according to the partitions.

8. The real-time data update and management method system according to claim 7. Characterized in that Further comprising: Judgment module: used to judge whether the batch data is processed after processing the data of one of the partitions: if the judgment result is positive, read the next batch of data for processing; If the judgment result is negative, obtain the data of the remaining partitions of this batch of data for processing.

Citation Information

Patent Citations

  • Big data real-time storage, processing and querying system

    CN106815338A

  • Data storage method and device for persisting kafka data to hdfs, and computer equipment

    CN111177271A