Method and system for loading data into data lake, electronic equipment and storage medium

By designing an automated method and system for data loading into a data lake, the problems of low data loading efficiency and relying on manual operations in the existing technology are solved, and an efficient and automated data entry process is realized, and the task is automatically rerun when it fails.

CN120066694APending Publication Date: 2025-05-30E-SURFING DIGITAL LIFE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311628575.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, data loading into the data lake is relatively low, and most of them rely on manual operations, resulting in manual modification of the configuration when data entering the lake fails to achieve task reruns, which affects efficiency.

Method used

By designing a method and system for loading data into a data lake, it includes obtaining the lake entry task, traversing the lake entry task table, generating task queues, and executing the lake entry script through multi-threading to generate task receipts and processing task failures.

Benefits of technology

It realizes the automation and efficiency of data loading into the data lake, reduces manual operations, improves the efficiency of data entering the lake, and automatically reruns the task when the task fails.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066694A_ABST
    Figure CN120066694A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for loading data into a data lake, electronic equipment and a storage medium, and the method comprises the steps: obtaining a first lake entering task, setting the task state of the first lake entering task as an initial task state, and adding the initial task state into a preset lake entering task table; traversing the lake entering task table, obtaining a first lake entering task in an initial task state, adding the first lake entering task in the initial task state into a first task queue, and updating the corresponding task state into the first task state; performing multi-thread execution on the first task queue through a preset lake entering script, returning a task execution result to the lake entering task table, and updating a corresponding task state into a second task state; generating a lake entering task receipt according to the task execution result and returning the lake entering task receipt to the user; and determining a first lake entering task which fails to be executed in the lake entering task table, and setting a task state as an initial task state. The efficiency of loading the data into the data lake is improved, and the method can be widely applied to the technical field of data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a method and system for loading data into a data lake, an electronic device, and a storage medium. Background Art

[0002] When the enterprise's collection of various types of data reaches a relatively large scale, it will inevitably involve big data analysis and processing technologies. At this time, it is necessary to gather these data to facilitate subsequent statistical analysis, so the concept of a data lake needs to be introduced. A data lake can converge structured, semi-structured, unstructured, and binary data. In recent years, more and more big data technologies have been integrated with Spring Boot, and Spring Boot is also a service framework that is convenient for rapid deployment and easy to integrate with various technologies. Spring Boot uses a large number of reflection technologies at the bottom layer to build IOC and AOP, decoupling the code to the greatest extent, and is suitable for the extension and maintenance development of each function. In the prior art, the related functions for data convergence into the lake are relatively lacking, and mostly rely on manual operations, resulting in low efficiency of data into the lake. Especially when the data into the lake fails, it is necessary to manually modify the configuration to achieve task rerun, which further affects the efficiency of data into the lake.

[0003] Term Explanation:

[0004] Spring Boot: Spring Boot is a brand-new framework provided by the Pivotal team. Its design purpose is to simplify the initial setup and development process of new Spring applications. This framework uses a specific method for configuration, so that developers no longer need to define boilerplate configurations.

[0005] IOC: IOC (Inversion of control) is the inversion of control. It is an idea rather than a technical implementation. IOC aims to facilitate project maintenance and testing. It provides a method for unified configuration and management of Java objects through Java's reflection mechanism. The Spring framework uses a container to manage the life cycle of objects. The container can configure objects by scanning XML files or specific Java annotations on classes. Developers can obtain objects through dependency lookup or dependency injection.

[0006] AOP: AOP (Aspect Oriented Programming) is a technology for aspect-oriented programming, which realizes the unified maintenance of program functions through pre-compilation and dynamic proxy during runtime.

[0007] Spark: Spark is a multi-language (Python / Java / Scala / R) engine for performing data engineering, data experimentation, and machine learning on a single node or cluster. It is commonly used for large-scale data analysis.

[0008] Hive: Hive is a distributed, fault-tolerant data warehouse system that supports large-scale data analysis. Hive Metastore (HMS) provides a central repository for metadata, which can be easily analyzed to make informed, data-driven decisions. Therefore, it is a key component of many data lake architectures. Hive is built on top of Apache Hadoop and supports storage such as S3, ADLS, and GS through HDFS. Hive allows users to read, write, and manage petabytes of data using SQL. Summary of the Invention

[0009] An object of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.

[0010] To this end, an object of an embodiment of the present invention is to provide a method for loading data into a data lake, which improves the efficiency of loading data into the data lake.

[0011] Another object of an embodiment of the present invention is to provide a system for loading data into a data lake.

[0012] In order to achieve the above technical object, the technical solutions adopted in the embodiments of the present invention include:

[0013] On the one hand, an embodiment of the present invention provides a method for loading data into a data lake, including the following steps:

[0014] Obtain a first data lake loading task, set the task status of the first data lake loading task to the initial task status, and add the first data lake loading task to a preset data lake loading task table;

[0015] Traverse the data lake loading task table, obtain the first data lake loading task in the initial task status, add the first data lake loading task in the initial task status to a first task queue, and update the corresponding task status to a first task status;

[0016] Execute the first task queue in multiple threads through a preset data lake loading script, then return the task execution result to the data lake loading task table, and update the corresponding task status to a second task status;

[0017] Generate a data lake loading task receipt according to the task execution result and return it to the user;

[0018] Determine the first lake-in task that fails to execute in the lake-in task list, and set the task status of the first lake-in task that fails to execute to the initial task status.

[0019] Further, in an embodiment of the present invention, the step of obtaining the first lake-in task, setting the task status of the first lake-in task to the initial task status, and adding the first lake-in task to a preset lake-in task list specifically includes:

[0020] Generate a first timing task according to a preset cron expression;

[0021] When the first timing task is triggered, obtain lake-in product information according to a preset product information table and basic information table, and generate a plurality of the first lake-in tasks according to the lake-in product information;

[0022] Determine the task ID and task batch of each of the first lake-in tasks, set the task status of the first lake-in tasks to the initial task status, and then add the first lake-in tasks to the lake-in task list according to the task ID and the task batch.

[0023] Further, in an embodiment of the present invention, the method for loading data into the data lake further includes the step of pre-configuring the product information table and the basic information table, which specifically includes:

[0024] Configure product information, the cron expression, and receipt recipient information through the product information table;

[0025] Configure the table name, table type, data output SQL, and data output path through the basic information table.

[0026] Further, in an embodiment of the present invention, the step of traversing the lake-in task list, obtaining the first lake-in tasks in the initial task status, adding the first lake-in tasks in the initial task status to the first task queue, and updating the corresponding task status to the first task status specifically includes:

[0027] Traverse the lake-in task list according to a preset time period, and obtain the first lake-in tasks in the initial task status;

[0028] Add the first lake-in tasks in the initial task status to the first task queue according to the task batch;

[0029] Update the task status of the first lake-in tasks added to the first task queue to the first task status, update the corresponding task status in the lake-in task list to the first task status, and record the corresponding task start time.

[0030] Further, in an embodiment of the present invention, the step of performing multi-threaded execution on the first task queue through a preset lake-in script, and then returning the task execution result to the lake-in task table and updating the corresponding task status to a second task status specifically includes:

[0031] Call the lake-in script to generate structured data files corresponding to each of the first lake-in tasks in the first task queue, where the structured data files include lake-in data files and verification files;

[0032] Based on the data output SQL and the data output path, load the lake-in data files to the corresponding positions in the data lake and perform verification according to the verification files;

[0033] Generate the task execution result according to the verification result, return the task execution result to the lake-in task table, and then update the corresponding task status in the lake-in task table to the second task status and record the corresponding task end time and the task execution result.

[0034] Further, in an embodiment of the present invention, the step of generating a lake-in task receipt according to the task execution result and returning it to the user specifically includes:

[0035] Perform task monitoring on the lake-in information table based on a preset monitoring queue, and obtain the first lake-in tasks that have been executed in the lake-in information table according to the task execution result;

[0036] Determine the task batch and task execution information of the first lake-in tasks that have been executed, where the task execution information includes the task start time, task end time, and task execution result, and then generate a lake-in task receipt according to the task batch and the task execution information;

[0037] Send the lake-in task receipt to the corresponding user based on the receipt recipient information.

[0038] Further, in an embodiment of the present invention, the step of determining the first lake-in tasks that have failed to execute in the lake-in task table and setting the task status of the first lake-in tasks that have failed to execute to the initial task status specifically includes:

[0039] Determine that the first lake-in tasks that have failed to execute in the lake-in task table are tasks to be re-run according to the task execution result;

[0040] Set the task status of the tasks to be re-run to the initial task status, so that the tasks to be re-run are re-executed when the lake-in task table is traversed next time.

[0041] On the other hand, an embodiment of the present invention provides a system for loading data into a data lake, including:

[0042] A task acquisition module, configured to acquire a first lake-in task, set the task status of the first lake-in task to an initial task status, and add the first lake-in task to a preset lake-in task table;

[0043] A task queue generation module, configured to traverse the lake-in task table, acquire the first lake-in task in the initial task status, add the first lake-in task in the initial task status to a first task queue, and update the corresponding task status to a first task status;

[0044] A task execution module, configured to perform multi-threaded execution on the first task queue through a preset lake-in script, then return the task execution result to the lake-in task table, and update the corresponding task status to a second task status;

[0045] A receipt sending module, configured to generate a lake-in task receipt according to the task execution result and return it to the user;

[0046] A task re-run module, configured to determine the first lake-in task that fails to execute in the lake-in task table, and set the task status of the first lake-in task that fails to execute to the initial task status.

[0047] On the other hand, an embodiment of the present invention provides an electronic device, where the electronic device includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for implementing connection communication between the processor and the memory. When the program is executed by the processor, it implements the method for loading data into a data lake as described above.

[0048] On the other hand, an embodiment of the present invention further provides a storage medium, where the storage medium is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method for loading data into a data lake as described above.

[0049] The advantages and beneficial effects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention:

[0050] In an embodiment of the present invention, a first lake-inlet task is obtained, the task status of the first lake-inlet task is set to the initial task status, and the first lake-inlet task is added to a preset lake-inlet task table. Then, the lake-inlet task table is traversed to obtain the first lake-inlet task in the initial task status, the first lake-inlet task in the initial task status is added to the first task queue, and the corresponding task status is updated to the first task status. Next, the first task queue is executed in a multi-threaded manner through a preset lake-inlet script. Furthermore, the task execution result is returned to the lake-inlet task table, and the corresponding task status is updated to the second task status. Then, a lake-inlet task receipt is generated based on the task execution result and returned to the user. Furthermore, the first lake-inlet task that fails to execute in the lake-inlet task table is determined, and the task status of the first lake-inlet task that fails to execute is set to the initial task status. Through generating a lake-inlet task, invoking a lake-inlet script to execute the task, and updating the task status, the embodiment of the present invention realizes the automatic control of data loading into the data lake, without manual operation by the user, improving the efficiency of data loading into the data lake; realizing the multi-threaded asynchronous execution of multiple lake-inlet tasks based on the task queue, further improving the efficiency of data loading into the data lake; when the lake-inlet task fails to execute, setting its task status to the initial task status, and automatically re-executing the task in the next traversal without manual modification of the configuration, further improving the efficiency of data loading into the data lake. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following introduces the drawings required to be used in the embodiments of the present invention. It should be understood that the drawings introduced below are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0052] Figure 1 It is a flowchart of the steps of the method for loading data into the data lake provided by the embodiment of the present invention;

[0053] Figure 2 It is a flowchart of step S101 provided by the embodiment of the present invention;

[0054] Figure 3 It is a flowchart of pre-configuring a product information table and a basic information table provided by the embodiment of the present invention;

[0055] Figure 4 It is a flowchart of step S102 provided by the embodiment of the present invention;

[0056] Figure 5 It is a flowchart of step S103 provided by the embodiment of the present invention;

[0057] Figure 6A flowchart of step S104 provided by an embodiment of the present invention;

[0058] Figure 7 A flowchart of step S105 provided by an embodiment of the present invention;

[0059] Figure 8 A schematic diagram of the logical architecture of the method for loading data into a data lake provided by an embodiment of the present invention;

[0060] Figure 9 A schematic diagram of the structure of the system for loading data into a data lake provided by an embodiment of the present invention;

[0061] Figure 10 A schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0062] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present application, and cannot be understood as a limitation to the present application. It should be noted that although the functional modules are divided in the system schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division from that in the system schematic diagram or a different order from that in the flowchart. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0063] In the description of the present invention, "a plurality" means two or more. If the first and second are described, it is only for the purpose of distinguishing technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0064] The method for loading data into a data lake provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a set-top box, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application for implementing the method for loading data into a data lake, etc., but is not limited to the above forms.

[0065] The present application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0066] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with the relevant laws, regulations, and standards of the relevant countries and regions. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.

[0067] The embodiments of the present invention can be used in the technology of supporting multi-table configuration into the lake for smart community products, specifically involving implementing various functions of configuring, executing, landing, and pushing the tasks into the lake by maintaining the product information table, basic information table, and the task table for entering the lake through the Spring Boot service. For the docking between the smart community and the group data lake, since it is necessary to output the data in the Hive table before the agreed time and split, compress, and generate verification files for the output data according to the specification document, the embodiments of the present invention split the functions of entering the lake into three pieces of data: product information, basic information, and task information for maintenance and monitoring. To solve the problems of uneven functions of entering the lake, lack of monitoring, and low maintainability, the embodiments of the present invention overall control the process of data entering the lake and achieve the monitoring of each batch of tasks entering the lake and the feedback on the task execution status, and timely and efficiently load the data into the data lake.

[0068] As Figure 1 shown in the following is a flowchart of steps of a method for loading data into the data lake provided by an embodiment of the present invention. Referring to Figure 1 , the embodiments of the present invention provide a method for loading data into the data lake, which specifically includes the following steps:

[0069] S101. Obtain the first task for entering the lake, set the task status of the first task for entering the lake to the initial task status, and add the first task for entering the lake to the preset task table for entering the lake.

[0070] Specifically, the deployment of the service for entering the lake requires the server to have some preconditions, including java, spark, and hive. After the preconditions are met, only the start.sh / stop.sh scripts are needed to complete the start / stop operations. After starting, a running log directory logs will be generated, and the administrator can view the logs to observe the running status of the tasks. In addition, the service for entering the lake also provides some startup configurations. One is to configure the output logs through logback-spring.xml, one is to provide the service startup port, the maximum number of parallel tasks, etc. through application.properties, and the last one is to configure the database properties through jdbc.properties.

[0071] As Figure 2 shown in the following is a flowchart of step S101 provided by an embodiment of the present invention. Referring to Figure 2 , further as an optional implementation manner, the step of obtaining the first task for entering the lake, setting the task status of the first task for entering the lake to the initial task status, and adding the first task for entering the lake to the preset task table for entering the lake specifically includes:

[0072] S1011. Generate the first scheduled task according to the preset cron expression;

[0073] S1012. When the first scheduled task is triggered, obtain the product information for lake input according to the preset product information table and basic information table, and generate multiple first lake input tasks based on the product information for lake input.

[0074] S1013. Determine the task IDs and task batches of the first lake input tasks, set the task status of the first lake input tasks to the initial task status, and then add the first lake input tasks to the lake input task table according to the task IDs and task batches.

[0075] Specifically, for the initialization of the task, it is necessary to obtain the effective (is_effective = 1) product information from the product information table. A scheduled task is generated according to the cron expression configured by the user (the cron expression fetches and updates from the table every 10 seconds). Until the scheduled task is triggered, the effective (is_effective = 1) product information for lake input in the product information table and the basic information table will be obtained. Based on the product information for lake input, batch tasks are generated. At the same time, a unique task ID is generated for each task, and the task status is initialized to 0 (i.e., the initial task status), and the task batch is 00. They are generated and filled into the lake input task table according to the accounting period. It should be noted that the cron expression is a syntax rule for setting scheduled tasks, which consists of 6 fields, representing seconds, minutes, hours, date, month, and week respectively.

[0076] As Figure 3 shown is a flowchart of pre-configuring the product information table and the basic information table provided by an embodiment of the present invention. Referring to Figure 3 , further as an optional implementation manner, the method for loading data into the data lake further includes the step of pre-configuring the product information table and the basic information table, which specifically includes:

[0077] S201. Configure product information, cron expression, and receipt recipient information through the product information table;

[0078] S202. Configure the table name, table type, data output SQL, and data output path through the basic information table.

[0079] Specifically, the lake input service of the embodiment of the present invention needs to maintain the product information table, the basic information table, and the lake input task table. The following will separately explain these three types of tables.

[0080] The product information table is mainly used to configure some basic information of the product, scheduled task configuration, email receipt carbon copy recipients, whether it is effective, data source, etc. Among them, the scheduled task is a cron expression, and dynamic configuration is also realized (reconfiguration does not require restarting the program).

[0081] The basic information table corresponds to the configuration of each table's information under the product, including table name, table type, data landing path, data output SQL, whether it is effective, etc. Among them, the data landing path points to the final landing position of the data, and the data output SQL supports custom time formats, such as #(yyyyMMdd), #(yyyy - MM - dd), #(yyyy year MM month dd day), etc. The default generated time is the current date minus 1 day, and #() is the only identifier for the system to determine the time format.

[0082] Most information in the lake - incoming task table does not support user configuration. It is only used to record the execution status of tasks and provide some basic information for task receipts. When the task is configured with a schedule, the task will first be initialized into this table. At this time, the lake - incoming batch is 00, the task status is 0, and the email status is 0. When the task is submitted for execution, the task status changes to 1, and the task start time is recorded. When the task is completed, the task status may be 2 / 3 / 4. At this time, the task is marked as ended, and the task end time and total duration are recorded. The task is transferred to the queue waiting for emails, and email receipts will be sent after batch completion. In addition, some information is available for users to change, including the lake - incoming batch and task status.

[0083] S102. Traverse the lake - incoming task table, obtain the first lake - incoming task in the initial task state, add the first lake - incoming task in the initial task state to the first task queue, and update the corresponding task status to the first task status.

[0084] Specifically, for the execution of the lake - incoming task, first, the Spring Boot - provided scheduled scheduling scans the lake - incoming tasks with a task table status of 0, registers the scanned tasks into the DAY_TOTAL_TASKS queue (i.e., the first task queue), and at the same time submits the spark tasks of each lake - incoming task to calculate the corresponding accounting period data volume update (insert or update semantics) to the lake - incoming task table.

[0085] As Figure 4 shown is a flowchart of step S102 provided by an embodiment of the present invention. Referring to Figure 4 , further as an optional implementation, the step of traversing the lake - incoming task table, obtaining the first lake - incoming task in the initial task state, adding the first lake - incoming task in the initial task state to the first task queue, and updating the corresponding task status to the first task status specifically includes:

[0086] S1021. Traverse the lake - incoming task table according to a preset time period and obtain the first lake - incoming task in the initial task state;

[0087] S1022. Add the first lake - incoming task in the initial task state to the first task queue according to the task batch;

[0088] S1023. Update the task status of the first lake-in task added to the first task queue to the first task status, update the corresponding task status in the lake-in task table to the first task status, and record the corresponding task start time.

[0089] Specifically, the time period for traversing the lake-in task table can be pre-configured, such as 6h or 12h; according to the task batches, the first lake-in tasks of the same batch in the initial task status are added to the same first task queue for subsequent multi-threaded asynchronous execution; update the corresponding task status to 1 (i.e., the first task status), and at the same time record the corresponding task start time in the lake-in task table.

[0090] S103. Perform multi-threaded execution on the first task queue through a preset lake-in script, then return the task execution result to the lake-in task table, and update the corresponding task status to the second task status.

[0091] As Figure 5 shown is a flowchart of step S103 provided by an embodiment of the present invention. Referring to Figure 5 , further as an optional implementation manner, the step of performing multi-threaded execution on the first task queue through a preset lake-in script, then returning the task execution result to the lake-in task table, and updating the corresponding task status to the second task status specifically includes:

[0092] S1031. Call the lake-in script to generate structured data files corresponding to each first lake-in task in the first task queue. The structured data files include lake-in data files and verification files.

[0093] S1032. Load the lake-in data files to the corresponding positions in the data lake based on the data output SQL and the data output path, and perform verification according to the verification files.

[0094] S1033. Generate a task execution result according to the verification result, return the task execution result to the lake-in task table, then update the corresponding task status in the lake-in task table to the second task status, and record the corresponding task end time and task execution result.

[0095] Specifically, perform multi-threaded asynchronous execution on the lake-in tasks. By calling the general lake-in script, generate structured data files (including lake-in data files and verification files) for each task, output the result files to the specified positions, and feedback the execution status and information of the tasks. Finally, update all the information (task status, task end time, task execution result, etc.) after the tasks are completed to the lake-in task table.

[0096] S104. Generate a lake-in task receipt according to the task execution result and return it to the user.

[0097] Specifically, task monitoring is only provided to service maintainers for observation in the form of logs, including the scheduled status of product tasks, the number of tasks being executed, the number of tasks generated on the same day, the execution status of tasks on the same day, and so on. By obtaining the status of each message queue to provide feedback for logging to implement the task monitoring function, it can help administrators quickly obtain the execution status of tasks and achieve the effect of quickly locating task anomalies.

[0098] As Figure 6 shown is a flowchart of step S104 provided by an embodiment of the present invention. Referring to Figure 6 , further as an optional implementation manner, the step of generating a lake-in task receipt based on the task execution result and returning it to the user specifically includes:

[0099] S1041. Perform task monitoring on the lake-in information table based on a preset monitoring queue, and obtain the first lake-in tasks that have been executed in the lake-in information table according to the task execution result;

[0100] S1042. Determine the task batch and task execution information of the first lake-in tasks that have been executed. The task execution information includes the task start time, task end time, and task execution result, and then generate a lake-in task receipt based on the task batch and task execution information;

[0101] S1043. Send the lake-in task receipt to the corresponding user based on the receipt recipient information.

[0102] Specifically, after all lake-in tasks of the product are completed and the task execution results are updated to the lake-in task table, the lake-in service will use each record in the product information table as the basic unit and the tasks under the product association as batch information to send the overall lake-in status to the user. The user can locate whether each task is successfully completed by viewing the email information.

[0103] S105. Determine the first lake-in tasks that have failed to execute in the lake-in task table, and set the task status of the first lake-in tasks that have failed to execute to the initial task status.

[0104] As Figure 7 shown is a flowchart of step S105 provided by an embodiment of the present invention. Referring to Figure 7 , further as an optional implementation manner, the step of determining the first lake-in tasks that have failed to execute in the lake-in task table and setting the task status of the first lake-in tasks that have failed to execute to the initial task status specifically includes:

[0105] S1051. Determine that the first lake-in tasks that have failed to execute in the lake-in task table are tasks to be re-run according to the task execution result;

[0106] S1052. Set the task status of the task to be rerun to the initial task status, so that the task to be rerun will be executed again when the lake-inbound task table is traversed next time.

[0107] Specifically, in the embodiments of the present invention, rerunning a task only requires modifying the task status in the lake-inbound task table. When the task status is 0, it means the task has not been executed. The lake-inbound service will automatically scan the task and update the status to 1 to indicate that it is being executed. For subsequent steps, no further operations are required by the user. The lake-inbound service will automatically complete all steps of lake-inbound according to the snapshot when the task was generated, and finally inform the user that the lake-inbound task rerun has been completed.

[0108] As Figure 8 shown is a schematic diagram of the logical architecture of the method for loading data into a data lake provided by the embodiments of the present invention. It can be recognized that the embodiments of the present invention implement the overall process of loading data into the data lake by generating tasks, invoking general scripts to execute tasks, completing tasks, generating receipt emails, and rerunning tasks. Similarly, each step of data lake-inbound is monitored in real time by monitoring queue information.

[0109] The above has described the method steps of the embodiments of the present invention. It can be understood that the embodiments of the present invention achieve automatic and controllable data loading into the data lake by generating lake-inbound tasks, invoking lake-inbound scripts to execute tasks, and updating task status, without manual operation by the user, improving the efficiency of data loading into the data lake; realizing multi-threaded asynchronous execution of multiple lake-inbound tasks based on the task queue, further improving the efficiency of data loading into the data lake; when a lake-inbound task fails, setting its task status to the initial task status, and automatically executing task rerun during the next traversal without manual configuration modification, further improving the efficiency of data loading into the data lake.

[0110] As Figure 9 shown is a schematic diagram of the structure of the system for loading data into a data lake provided by the embodiments of the present invention. Referring to Figure 9 , the embodiments of the present invention provide a system for loading data into a data lake, including:

[0111] A task acquisition module, configured to acquire a first lake-inbound task, set the task status of the first lake-inbound task to the initial task status, and add the first lake-inbound task to a preset lake-inbound task table;

[0112] A task queue generation module, configured to traverse the lake-inbound task table, acquire the first lake-inbound task in the initial task status, add the first lake-inbound task in the initial task status to a first task queue, and update the corresponding task status to a first task status;

[0113] A task execution module, which is used to execute the first task queue in a multi-threaded manner through a preset lake-inbound script, and then return the task execution result to the lake-inbound task table and update the corresponding task status to the second task status;

[0114] A receipt sending module, which is used to generate a lake-inbound task receipt according to the task execution result and return it to the user;

[0115] A task re-running module, which is used to determine the first lake-inbound task that fails to execute in the lake-inbound task table and set the task status of the first lake-inbound task that fails to execute to the initial task status.

[0116] The content in the above method embodiments is applicable to the present system embodiment. The functions specifically implemented by the present system embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0117] The embodiment of the present invention also provides an electronic device, which includes: a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing the connection and communication between the processor and the memory. When the program is executed by the processor, it realizes the method for loading the above data into the data lake. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0118] As Figure 10 shown is a schematic hardware structure diagram of the electronic device provided by the embodiment of the present invention. Referring to Figure 10 , the embodiment of the present invention provides an electronic device, including:

[0119] A processor 1001, which can be implemented by using a general-purpose CPU (Central Processing Unit, central processor), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solution provided by the embodiment of the present invention;

[0120] A memory 1002, which can be implemented in the form of a read-only memory (ReadOnly Memory, ROM), a static storage device, a dynamic storage device, or a random access memory (Random Access Memory, RAM), etc. The memory 1002 can store an operating system and other application programs. When implementing the technical solution provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1002, and the processor 1001 is used to call and execute the method for loading the data of the embodiment of the present invention into the data lake;

[0121] An input / output interface 1003 for implementing information input and output;

[0122] A communication interface 1004 for implementing communication interaction between this device and other devices, which can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0123] A bus 1005 for transmitting information between various components of the device (such as the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004);

[0124] Among them, the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 achieve communication connections with each other inside the device through the bus 1005.

[0125] An embodiment of the present invention also provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above method of loading data into the data lake.

[0126] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0127] An embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the method shown.

[0128] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may in fact be executed substantially simultaneously or the blocks may sometimes be executed in the reverse order. Further, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for purposes of providing a more thorough understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are envisioned in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.

[0129] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the above-described functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Accordingly, those of ordinary skill in the art will be able to implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are illustrative only and are not intended to limit the scope of the present invention, the scope of which is determined by the full scope of the appended claims and their equivalents.

[0130] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods of the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a removable hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc.

[0131] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definitional sequence of executable instructions for implementing a logical function, which can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with the instruction execution system, apparatus, or device.

[0132] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the above-mentioned program can be printed, as the above-mentioned program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then storing it in a computer memory.

[0133] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0134] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0135] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

[0136] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A method for loading data into a data lake, characterized in that, it includes the following steps: Obtain a first data lake loading task, set the task status of the first data lake loading task to the initial task status, and add the first data lake loading task to a preset data lake loading task table; Traverse the data lake loading task table, obtain the first data lake loading task in the initial task status, add the first data lake loading task in the initial task status to a first task queue and update the corresponding task status to the first task status; Execute the first task queue in a multi-threaded manner through a preset data lake loading script, then return the task execution result to the data lake loading task table, and update the corresponding task status to the second task status; Generate a data lake loading task receipt according to the task execution result and return it to the user; Determine the first data lake loading task that fails to execute in the data lake loading task table, and set the task status of the first data lake loading task that fails to execute to the initial task status.

2. The method for loading data into a data lake according to claim 1, characterized in that, The step of obtaining the first data lake loading task, setting the task status of the first data lake loading task to the initial task status, and adding the first data lake loading task to a preset data lake loading task table specifically includes: Generate a first scheduled task according to a preset cron expression; When the first scheduled task is triggered, obtain data lake loading product information according to a preset product information table and basic information table, and generate a plurality of the first data lake loading tasks according to the data lake loading product information; Determine the task ID and task batch of each of the first data lake loading tasks, set the task status of the first data lake loading task to the initial task status, and then add the first data lake loading task to the data lake loading task table according to the task ID and the task batch.

3. The method for loading data into a data lake according to claim 2, characterized in that, The method for loading data into a data lake further includes a step of pre-configuring the product information table and the basic information table, which specifically includes: Configure product information, the cron expression, and receipt recipient information through the product information table; Configure the table name, table type, data output SQL, and data output path through the basic information table.

4. The method for loading data into a data lake according to claim 2, characterized in that, The step of traversing the data lake loading task table, obtaining the first data lake loading task in the initial task status, adding the first data lake loading task in the initial task status to a first task queue and updating the corresponding task status to the first task status specifically includes: Traverse the data lake loading task table according to a preset time period, and obtain the first data lake loading task in the initial task status; Add the first data lake loading task in the initial task status to the first task queue according to the task batch; Update the task status of the first data lake loading task added to the first task queue to the first task status, update the corresponding task status in the data lake loading task table to the first task status, and record the corresponding task start time.

5. A method for loading data into a data lake according to claim 3, characterized in that, the step of multi-threadedly executing the first task queue through a preset lake-in script, and then returning the task execution result to the lake-in task table and updating the corresponding task status to a second task status specifically includes: invoking the lake-in script to generate structured data files corresponding to each of the first lake-in tasks in the first task queue, where the structured data files include lake-in data files and verification files; loading the lake-in data files to corresponding positions in the data lake based on the data output SQL and the data output path, and performing verification according to the verification files; generating the task execution result according to the verification result, returning the task execution result to the lake-in task table, and then updating the corresponding task status in the lake-in task table to the second task status and recording the corresponding task end time and the task execution result.

6. A method for loading data into a data lake according to claim 3, characterized in that, the step of generating a lake-in task receipt according to the task execution result and returning it to the user specifically includes: performing task monitoring on the lake-in information table based on a preset monitoring queue, and obtaining the first lake-in tasks that have been executed in the lake-in information table according to the task execution result; determining the task batch and task execution information of the executed first lake-in tasks, where the task execution information includes the task start time, task end time, and task execution result, and then generating a lake-in task receipt according to the task batch and the task execution information; sending the lake-in task receipt to the corresponding user based on the receipt recipient information.

7. A method for loading data into a data lake according to any one of claims 1 to 6, characterized in that, the step of determining the first lake-in tasks that have failed to execute in the lake-in task table and setting the task status of the first lake-in tasks that have failed to execute to the initial task status specifically includes: determining, according to the task execution result, that the first lake-in tasks that have failed to execute in the lake-in task table are tasks to be re-run; setting the task status of the tasks to be re-run to the initial task status, so that the tasks to be re-run are re-executed when the lake-in task table is traversed next time.

8. A system for loading data into a data lake, characterized in that, comprising: a task acquisition module, configured to acquire a first lake-in task, set the task status of the first lake-in task to the initial task status, and add the first lake-in task to a preset lake-in task table; a task queue generation module, configured to traverse the lake-in task table, acquire the first lake-in tasks in the initial task status, add the first lake-in tasks in the initial task status to a first task queue and update the corresponding task status to a first task status; a task execution module, configured to multi-threadedly execute the first task queue through a preset lake-in script, and then return the task execution result to the lake-in task table and update the corresponding task status to a second task status; A receipt sending module, configured to generate a lake-inbound task receipt according to the task execution result and return it to the user; A task re-running module, configured to determine the first lake-inbound task that fails to execute in the lake-inbound task table, and set the task status of the first lake-inbound task that fails to execute to the initial task status.

9. An electronic device, characterized in that, the electronic device includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory, and when the program is executed by the processor, it realizes the steps of the method for loading data into a data lake as described in any one of claims 1 to 7.

10. A storage medium, which is a computer-readable storage medium for computer-readable storage, characterized in that, the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to realize the steps of the method for loading data into a data lake as described in any one of claims 1 to 7.