Data import method and apparatus, electronic device, storage medium, and program product

By reusing the target data files in the recycling bin in the database, the problem of repeated resource consumption when retrying import after the import fails, improving the efficiency of data import.

WO2025123848A1PCT designated stage expired Publication Date: 2025-06-19CHINA TELECOM CORP LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/120740
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-14
Filing Date
2024-09-24
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

In the database, when external data sources import the database, when network resources are unstable or system resources are occupied, the probability of import failure increases, resulting in duplicate resource consumption during retry import, which is inefficient.

Method used

Avoid duplicate imports by obtaining the target data file of the failed source file imported in the recycle bin of the target database and updating it to the row collection structure of the target data import request.

Benefits of technology

Reduces resource overhead caused by retrying import tasks, improves the efficiency of data import, and reduces the consumption of CPU, memory and IO resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024120740_19062025_PF_FP_ABST
    Figure CN2024120740_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a data import method and apparatus, an electronic device, and a non-volatile storage medium. The method comprises: when a target database receives a target data import request, acquiring a source file set corresponding to the target data import request; on the basis of historical import information, determining a source file, which has been successfully imported into the target database, in the source file set as a first source file; and acquiring a target data file generated by the first source file in a recycle bin of the target database during historical import, and updating the target data file into a row set structure corresponding to the target data import request.
Need to check novelty before this filing date? Find Prior Art

Description

Data import method, device, electronic device, storage medium and program product

[0001] Related applications

[0002] This application claims priority to Chinese patent application number 2023117237210, filed on December 14, 2023, entitled “Data import method, device, electronic device and non-volatile storage medium,” the entire text of which is hereby incorporated by reference. Technical Field

[0003] The present application relates to the field of data processing and storage technology, and in particular to a data import method, device, electronic device and non-volatile storage medium. Background Art

[0004] During database use, as more businesses are connected, after a new table is successfully created, the database usually needs to import a large amount of data from an external data source. During the import process, there is a certain probability that the external data source will fail if the network resources are unstable or the system resource usage is high.

[0005] Summary of the Invention

[0006] Embodiments of the present application provide a data import method, device, electronic device, non-volatile storage medium, and computer program product.

[0007] An embodiment of the present application provides a data import method, including: when a target database receives a target data import request, obtaining a source file set corresponding to the target data import request, wherein the target data import request is used to request to import source files in a source file set outside the target database into the target database; based on historical import information, determining a source file in the source file set that has been successfully imported into the target database as a first source file, wherein the historical import information is used to record the import status of the historical source file corresponding to the historical data import request received by the target database; obtaining a target data file generated by the first source file in the recycle bin of the target database during the historical import, and updating the target data file to a row set structure corresponding to the target data import request, wherein the recycle bin stores historical source files that have been successfully imported from the historical source file set corresponding to the historical data import request on which the target database failed to execute.

[0008] Optionally, after updating the target data file to the row set structure corresponding to the target data import request, the method further includes: determining hierarchical relationship information corresponding to the target data file in the target database, wherein the hierarchical relationship information is used to characterize the physical storage unit and row set structure corresponding to the target data file, and the segment position in the row set structure, wherein one physical storage unit corresponds to multiple row set structures, and one row set structure corresponds to multiple segment positions; determining the file path and hash value of the source file corresponding to the target data file; determining the target timestamp corresponding to the target data file, wherein the target timestamp is used to characterize the moment when the target data file is updated to the row set structure corresponding to the target data import request; and storing the hierarchical relationship data, file path, hash value and target timestamp as historical import information.

[0009] Optionally, based on historical import information, determining a source file in the source file collection that has been successfully imported into the target database as the first source file includes: determining the file path and hash value of the source file in the source file collection, and judging whether there is a historical source file in the historical import information with a file path and hash value consistent with the source file in the source file collection; in the case that there is a historical source file in the historical import information with a file path and hash value consistent with the source file, obtaining a target timestamp corresponding to the target data file corresponding to the historical source file in the historical import information; based on the target timestamp, determining the retention time of the target data file in the recycle bin, and judging whether the retention time is greater than a preset retention time threshold; in the case that the retention time is not greater than the preset retention time threshold, determining the source file as the first source file.

[0010] Optionally, updating the target data file to the row set structure corresponding to the target data import request includes: creating a row set structure at a location corresponding to the hierarchical relationship information corresponding to the first source file in the target database; storing the target data file generated by the first source file in the recycle bin during historical import to the segment location corresponding to the hierarchical relationship information in the row set structure, and updating the file name of the target data file based on the identifier of the row set structure.

[0011] Optionally, the row set structure created at the location corresponding to the hierarchical relationship information in the target database has the same hierarchical relationship as the row set structure stored during the historical import of the first source file.

[0012] Optionally, the method also includes: when there is no historical source file in the historical import information whose file path and hash value are consistent with the source file, or when the retention time is greater than a preset retention time threshold, determining the source file as a second source file; generating a target data file corresponding to the second source file, and storing the target data file in the target database.

[0013] Optionally, the method further includes: storing hierarchical relationship information corresponding to the target data file corresponding to the second source file, a target timestamp, and a file path and a hash value of the second source file as historical import information.

[0014] Optionally, the method also includes: when there are source files in the source file collection that fail to be imported, determining that the target data import request has failed to execute; and cleaning the target data files corresponding to the successfully imported source files in the source file collection corresponding to the target data import request in the target database into the recycle bin.

[0015] Optionally, the information recorded in the historical import information includes at least one of the following: a unique label of the target data import request, the file path and hash value of the imported source file, the execution status of the imported source file, the target timestamp when the source file is successfully imported, and the hierarchical relationship information and file path of the target data file corresponding to the source file.

[0016] An embodiment of the present application provides a data import device, including: a file acquisition module, used to obtain a source file set corresponding to a target data import request when a target database receives a target data import request, wherein the target data import request is used to request to import source files in a source file set outside the target database into the target database; a file matching module, used to determine, based on historical import information, a source file in the source file set that has been successfully imported into the target database as a first source file, wherein the historical import information is used to record the import status of the historical source file corresponding to the historical data import request received by the target database; a fast import module, used to obtain a target data file generated by the first source file in a recycle bin of the target database during the historical import, and update the target data file to a row set structure corresponding to the target data import request, wherein the recycle bin stores historical source files that have been successfully imported from the historical source file set corresponding to historical data import requests that failed to be executed by the target database.

[0017] An embodiment of the present application provides an electronic device, comprising: a memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the above-mentioned data import method is executed when the program is run.

[0018] An embodiment of the present application provides a non-volatile storage medium, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the above-mentioned data import method by running the computer program.

[0019] An embodiment of the present application provides a computer program product, including a computer program, wherein the computer program implements the above-mentioned data import method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.

[0021] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0022] FIG1 is a hardware structure block diagram of a computer terminal (or electronic device) for implementing a method for data importing according to an embodiment of the present application;

[0023] FIG2 is a schematic diagram of a data import method flow according to an embodiment of the present application;

[0024] FIG3 is a schematic diagram of a method flow for accelerating import retry according to an embodiment of the present application;

[0025] FIG4 is a schematic diagram of a logical hierarchical relationship of data files in a database provided according to an embodiment of the present application;

[0026] FIG5 is a schematic diagram of interaction between logical components according to an embodiment of the present application;

[0027] FIG6 is a schematic structural diagram of a data import device provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0030] To facilitate those skilled in the art to better understand the embodiments of the present application, some technical terms or nouns involved in the embodiments of the present application are explained as follows:

[0031] Doris database: It is a high-performance, real-time analytical database that can return query results under massive data with only sub-second response time. It can support not only high-concurrency point query scenarios, but also high-throughput complex analysis scenarios.

[0032] Database systems (for example, the Doris database) often need to import large amounts of data from external data sources. During the import process, there is a certain probability of import failure if network resources are unstable or system resources are heavily occupied. Users often initiate the same import task again to re-import data because they cannot tell which files or data were imported successfully or failed.

[0033] However, when current database systems process retry requests for failed import tasks of the same data source, they treat the retried import request as a new import task and re-import all the external source files corresponding to the task. This will repeatedly consume resources such as the Central Processing Unit (CPU), memory, and disk input and output (IO), and there is still a risk of import failure, resulting in low overall import efficiency. Moreover, as the amount of data increases, the resource consumption occupied by this retried import increases proportionally, resulting in problems such as low database data import efficiency.

[0034] In order to solve the above problems, relevant solutions are provided in the embodiments of the present application, which are described in detail below.

[0035] According to an embodiment of the present application, an embodiment of a method for importing data is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0036] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 shows a hardware structure block diagram of a computer terminal (or electronic device) for implementing a data import method. As shown in Figure 1, the computer terminal 10 (or electronic device) may include one or more (102a, 102b, ..., 102n are used in the figure to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that the structure shown in Figure 1 is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may also include more or fewer components than those shown in Figure 1, or have a configuration different from that shown in Figure 1.

[0037] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or electronic device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0038] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the data import method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned data import method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0039] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0040] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or electronic device).

[0041] In the above operating environment, an embodiment of the present application provides a data import method. FIG2 is a schematic diagram of a data import method flow according to an embodiment of the present application. As shown in FIG2 , the method includes the following steps:

[0042] Step S202: When the target database receives a target data import request, a source file set corresponding to the target data import request is obtained, wherein the target data import request is used to request to import source files in the source file set outside the target database into the target database;

[0043] In this embodiment, the target database is described by taking the Doris database as an example. It should be understood that the database in this application is not limited to the Doris database, and may also be other types of databases.

[0044] Step S204: Determine, based on the historical import information, a source file in the source file set that has been successfully imported into the target database as a first source file, wherein the historical import information is used to record the import status of the historical source file corresponding to the historical data import request received by the target database;

[0045] Step S206, obtain the target data file generated by the first source file in the recycle bin of the target database during the historical import, and update the target data file to the row set structure corresponding to the target data import request, wherein the recycle bin stores the historical source files that were successfully imported in the historical source file set corresponding to the historical data import request that failed to be executed by the target database.

[0046] Through the above steps, by retrying the import request for the failed import task when the data source is consistent, the imported data that has entered the recycle bin due to task failure is reused according to the unique signature of the import task and the file identifiers involved and the unique storage format of the data file in the database storage system, thereby achieving the purpose of reducing the resource overhead caused by a large number of retries of the import task, improving the overall import efficiency, and accelerating the successful execution of the import task, thereby solving the technical problem of low data import efficiency caused by the need to re-import all source files corresponding to the same task that failed to import in the process of importing an external data source into the database.

[0047] The data import method in steps S202 to S206 of the embodiment of the present application is further introduced below.

[0048] Figure 3 is a schematic diagram of a method flow for accelerating import retry provided according to an embodiment of the present application. As shown in Figure 3, in the embodiment of the present application, the acceleration of the import of duplicate data is file-granular. Combined with Doris's unique data file storage format and the uniqueness of the import task signature, when receiving the client's target data import request, it is determined whether the source file has been successfully imported, and a decision is made on whether to re-import it later, thereby reducing the import of duplicate data as much as possible and reducing resource consumption.

[0049] The following further introduces the method flow of accelerating import retry shown in Figure 3.

[0050] In order to better understand the present application scheme, the hierarchical relationship of data file storage in the database is first introduced. As shown in Figure 4, database subsets (shards) are used to divide data into smaller parts in a distributed database system. Each shard represents a part of the table, allowing the system to distribute storage and processing data on multiple nodes. Each data import request corresponds to a row set structure, namely Rowset. Rowset is a data set of a data change in Tablet (shard). Data changes include: data import, deletion, update, etc. Rowset is recorded according to version information, and a version is generated for each change. Tablet is the actual physical storage unit of a data table in the database. A table is stored in the distributed storage layer of the database in units of Tablets after partitioning and bucketing. Each Tablet includes metadata and several consecutive Rowsets, that is, one physical storage unit corresponds to multiple row set structures.

[0051] Usually, a data import request will correspond to multiple source files that need to be imported, that is, each import task may generate multiple Segment files. Segment (segment location) is the smallest unit stored in the target database. The target data file corresponding to a source file corresponds to a segment location. The Segment files generated by the same import task logically belong to the same Rowset, that is, a row set structure corresponds to multiple segment locations.

[0052] For repeated imports of the same source file, under the same cluster configuration, except for the different generated rowset_id (identifier of the row set structure), the content and size of the target data file are the same. Therefore, the embodiment of the present application can complete the accelerated process of import retry shown in Figure 3 by constructing a Manager (management component), a Filter (filtering component) and an Executor (processing component). The interaction logic between the above three components is shown in Figure 5.

[0053] Among them, Manager is used to record the import status of each source file in the source file set corresponding to the target data import request, Filter is used to filter the first source file in the source file set that has been successfully imported into the target database, and Executor is used to quickly import the first source file that has been judged as a duplicate (successfully imported into the target database) by Filter, that is, to obtain the target data file generated by the first source file in the recycle bin of the target database during the historical import, and update the target data file to the row set structure corresponding to the target data import request. Through the collaborative work of the above three components, unnecessary repeated imports are avoided, resource consumption is saved and import efficiency is improved. The following is a further introduction to this process.

[0054] First, when receiving a target data import request (load task), the Filter queries the Manager based on the unique label of the import task of the data import request and the source file information to be pulled to see whether there is a source file in the source file collection corresponding to the target data import request that has been successfully imported into the target database. It also needs to determine whether the retention time of the corresponding target data file in the recycle bin has expired; the Manager determines the source file in the source file collection that has been successfully imported into the target database as the first source file based on the historical import information. The specific steps are as follows.

[0055] In some embodiments of the present application, based on historical import information, determining a source file in a source file collection that has been successfully imported into a target database as a first source file includes the following steps: determining the file path and hash value of the source file in the source file collection, and judging whether there is a historical source file in the historical import information whose file path and hash value are consistent with the source file in the source file collection; in the case that there is a historical source file in the historical import information whose file path and hash value are consistent with the source file, obtaining the target timestamp corresponding to the target data file corresponding to the historical source file in the historical import information; based on the target timestamp, determining the retention time of the target data file in the recycle bin, and judging whether the retention time is greater than a preset retention time threshold; in the case that the retention time is not greater than the preset retention time threshold, determining the source file as the first source file.

[0056] After determining that the duplicate source file has been successfully imported and the first source file in the recycle bin has not expired in the retention time of the corresponding target data file, the Executor can quickly import the first source file (Executor fast load), that is, create a new row set structure Rowset according to the import logic of the first source file, and update the target data file generated by the first source file in the recycle bin of the target database during the historical import to the row set structure. The specific steps are as follows.

[0057] In some embodiments of the present application, updating the target data file to the row set structure corresponding to the target data import request includes the following steps: creating a row set structure at a location corresponding to the hierarchical relationship information corresponding to the first source file in the target database; storing the target data file generated by the first source file in the recycle bin during historical import to the segment location corresponding to the hierarchical relationship information in the row set structure, and updating the file name of the target data file based on the identifier of the row set structure.

[0058] Specifically, the hierarchical relationship corresponding to the newly created row set structure Rowset is the same as the hierarchical relationship of the Rowset stored during the historical import of the first source file. The only difference is that the identifier (rowset id) of the current row set structure is different from the identifier of the row set structure generated during the historical import. After the row set structure is created, the Executor can move the target data file generated during the historical import in the recycle bin directly to the segment location directory corresponding to the row set structure generated by this import based on the information of the target data file corresponding to the first source file in the target database recorded in the Manager, and update the file name of the target data file based on the identifier of this row set structure, synchronously update the metadata, and update the latest import record to the Manager to complete the short path fast import process of the first source file judged to be duplicated.

[0059] The embodiment of the present application records the import status of the source file and the unique mapping between the source file and the target file. When the same import task is re-initiated, the source file that has been successfully imported can be quickly skipped to speed up its import speed. By reusing the data files in the recycle bin, the fast import of the same source file in the retried import task can be completed, which can greatly reduce the consumption of CPU, memory and IO resources and improve the import efficiency. At the same time, by creating a new Rowset to reuse the Rowset hierarchical relationship generated by the historical import, the metadata update can be quickly completed in combination with the existing logic.

[0060] The following is a further introduction to the process of updating the latest imported records to the Manager. The specific steps are as follows.

[0061] In some embodiments of the present application, after updating the target data file to the row set structure corresponding to the target data import request, the method further includes the following steps: determining the hierarchical relationship information corresponding to the target data file in the target database, wherein the hierarchical relationship information is used to characterize the physical storage unit and row set structure corresponding to the target data file, and the segment position in the row set structure, wherein one physical storage unit corresponds to multiple row set structures, and one row set structure corresponds to multiple segment positions; determining the file path and hash value of the source file corresponding to the target data file; determining the target timestamp corresponding to the target data file, wherein the target timestamp is used to characterize the moment when the target data file is updated to the row set structure corresponding to the target data import request; and storing the hierarchical relationship data, file path, hash value and target timestamp as historical import information.

[0062] In this embodiment, the Manager component can be used to record the import status of each source file, that is, the above-mentioned historical import information. Specifically, when the business is accessing data, a target data import request (import task) may specify multiple source files from an external data source for import, and a source file may generate multiple target data files when writing to the target database. The information recorded in the above-mentioned historical import information includes at least one of the following: a unique label of the target data import request, the file path and hash value of the imported source file, the execution status of the imported source file, the target timestamp when the source file is successfully imported, and the hierarchical relationship information of the target data file written to the target database corresponding to the source file (including the physical storage unit and row set structure corresponding to the target data file, and the segment position in the row set structure) and file path (possibly multiple).

[0063] In addition, for source files in the source file set corresponding to the target data import request that have not been successfully imported into the target database, as well as source files that have been successfully imported into the target database but whose corresponding target data files in the recycle bin have expired retention time, the normal import process (normal load) can be used for import. The specific steps are as follows.

[0064] In some embodiments of the present application, the method also includes the following steps: when there is no historical source file in the historical import information whose file path and hash value are consistent with the source file, or when the retention time is greater than a preset retention time threshold, the source file is determined to be a second source file; a target data file corresponding to the second source file is generated, and the target data file is stored in the target database.

[0065] After the second source file is successfully imported, the Manager component can be used to record the import status of the second source file. The specific steps are as follows.

[0066] In some embodiments of the present application, the method further includes the following steps: storing hierarchical relationship information corresponding to the target data file corresponding to the second source file, the target timestamp, and the file path and hash value of the second source file as historical import information.

[0067] If all source files (including the first source file and the second source file) in the source file set corresponding to the target data import request have been successfully imported into the target database, the target data import request is successfully executed; if there are source files in the source file set that fail to be imported, the target data import request is determined to have failed to be executed; the target data files corresponding to the successfully imported source files in the source file set corresponding to the target data import request in the target database are cleaned up to the recycle bin.

[0068] This application solution records the source files successfully imported in each import task and the target file information generated by the target database, combined with the effective length of time the generated target files are retained in the recycle bin. For the import of duplicate source files, the fast import is completed by reusing the data of the target data files in the recycle bin and synchronously updating the metadata, which can greatly reduce the consumption of CPU, memory and disk IO resources.

[0069] According to an embodiment of the present application, an embodiment of a data import device is also provided. FIG6 is a structural diagram of a data import device provided according to an embodiment of the present application. As shown in FIG6 , the device includes:

[0070] A file acquisition module 60 is configured to, when the target database receives a target data import request, acquire a source file set corresponding to the target data import request, wherein the target data import request is used to request to import source files in the source file set outside the target database into the target database;

[0071] A file matching module 62 is configured to determine, based on historical import information, a source file in the source file set that has been successfully imported into the target database as a first source file, wherein the historical import information is used to record the import status of the historical source file corresponding to the historical data import request received by the target database;

[0072] Optionally, based on historical import information, determining a source file in the source file collection that has been successfully imported into the target database as the first source file includes: determining the file path and hash value of the source file in the source file collection, and judging whether there is a historical source file in the historical import information with a file path and hash value consistent with the source file in the source file collection; in the case that there is a historical source file in the historical import information with a file path and hash value consistent with the source file, obtaining a target timestamp corresponding to the target data file corresponding to the historical source file in the historical import information; based on the target timestamp, determining the retention time of the target data file in the recycle bin, and judging whether the retention time is greater than a preset retention time threshold; in the case that the retention time is not greater than the preset retention time threshold, determining the source file as the first source file.

[0073] The quick import module 64 is used to obtain the target data file generated by the first source file in the recycle bin of the target database during the historical import, and update the target data file to the row set structure corresponding to the target data import request, wherein the recycle bin stores the historical source files that were successfully imported in the historical source file set corresponding to the historical data import request that failed to execute the target database.

[0074] Optionally, updating the target data file to the row set structure corresponding to the target data import request includes: creating a row set structure at a location corresponding to the hierarchical relationship information corresponding to the first source file in the target database; storing the target data file generated by the first source file in the recycle bin during historical import to the segment location corresponding to the hierarchical relationship information in the row set structure, and updating the file name of the target data file based on the identifier of the row set structure.

[0075] Optionally, the row set structure created at the location corresponding to the hierarchical relationship information in the target database has the same hierarchical relationship as the row set structure stored during the historical import of the first source file.

[0076] Optionally, after updating the target data file to the row set structure corresponding to the target data import request, the quick import module 64 is further used to: determine the hierarchical relationship information corresponding to the target data file in the target database, wherein the hierarchical relationship information is used to characterize the physical storage unit and row set structure corresponding to the target data file, and the segment position in the row set structure, wherein one physical storage unit corresponds to multiple row set structures, and one row set structure corresponds to multiple segment positions; determine the file path and hash value of the source file corresponding to the target data file; determine the target timestamp corresponding to the target data file, wherein the target timestamp is used to characterize the moment when the target data file is updated to the row set structure corresponding to the target data import request; store the hierarchical relationship data, file path, hash value and target timestamp as historical import information.

[0077] Optionally, the data import device is also used to: determine the source file as the second source file when there is no historical source file whose file path and hash value are consistent with the source file in the historical import information, or when the retention time is greater than a preset retention time threshold; generate a target data file corresponding to the second source file, and store the target data file in the target database.

[0078] Optionally, the data import device is further used to store hierarchical relationship information corresponding to the target data file corresponding to the second source file, a target timestamp, and a file path and hash value of the second source file as historical import information.

[0079] Optionally, the data import device is also used to: determine that the target data import request has failed to execute when there are source files in the source file collection that fail to be imported; and clean up the target data files corresponding to the successfully imported source files in the source file collection corresponding to the target data import request in the target database into the recycle bin.

[0080] Optionally, the information recorded in the historical import information includes at least one of the following: a unique label of the target data import request, the file path and hash value of the imported source file, the execution status of the imported source file, the target timestamp when the source file is successfully imported, and the hierarchical relationship information and file path of the target data file corresponding to the source file.

[0081] It should be noted that the various modules in the above-mentioned data import device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.

[0082] It should be noted that the data import device provided in this embodiment can be used to execute the above-mentioned data import method. Therefore, the relevant explanations and instructions on the above-mentioned data import method are also applicable to the embodiments of this application and will not be repeated here.

[0083] An embodiment of the present application also provides a non-volatile storage medium, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the data import method according to the above embodiment by running the computer program: when the target database receives a target data import request, a source file set corresponding to the target data import request is obtained, wherein the target data import request is used to request to import source files in a source file set outside the target database into the target database; based on historical import information, a source file in the source file set that has been successfully imported into the target database is determined as a first source file, wherein the historical import information is used to record the import status of the historical source file corresponding to the historical data import request received by the target database; the target data file generated by the first source file in the recycle bin of the target database during the historical import is obtained, and the target data file is updated to the row set structure corresponding to the target data import request, wherein the recycle bin stores the historical source files that have been successfully imported from the historical source file set corresponding to the historical data import request on which the target database failed to execute.

[0084] The features described in the aforementioned data import method embodiment are applicable to the data import method executed by the device where the non-volatile storage medium is located by running a computer program, and will not be described in detail here.

[0085] In one embodiment, a computer program product is provided, including a computer program, which is executed by a processor according to the data import method of the above embodiment: when a target database receives a target data import request, a source file set corresponding to the target data import request is obtained, wherein the target data import request is used to request to import source files in a source file set outside the target database into the target database; based on historical import information, a source file in the source file set that has been successfully imported into the target database is determined as a first source file, wherein the historical import information is used to record the import status of the historical source file corresponding to the historical data import request received by the target database; a target data file generated by the first source file in the recycle bin of the target database during the historical import, and the target data file is updated to the row set structure corresponding to the target data import request, wherein the recycle bin stores the historical source files that have been successfully imported from the historical source file set corresponding to the historical data import request that failed to be executed by the target database.

[0086] The features described in the aforementioned data import method embodiment are applicable to the data import method in which a computer program is executed by a processor and will not be described in detail here.

[0087] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0088] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0089] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0090] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0091] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0092] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0093] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

[0094] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A data import method, comprising: When the target database receives a target data import request, obtaining a source file set corresponding to the target data import request, wherein the target data import request is used to request to import source files in the source file set outside the target database into the target database; According to the historical import information, the source file in the source file set that has been successfully imported into the target database is determined as the first source file, wherein the historical import information is used to record the import status of the historical source file corresponding to the historical data import request received by the target database; Obtain the target data file generated by the first source file in the recycle bin of the target database during historical import, and update the target data file to the row set structure corresponding to the target data import request, wherein the recycle bin stores the historical source files that were successfully imported from the historical source file set corresponding to the historical data import request that failed to be executed by the target database.

2. The data import method according to claim 1, wherein: After updating the target data file to the row set structure corresponding to the target data import request, the method further includes: Determine hierarchical relationship information corresponding to the target data file in the target database, wherein the hierarchical relationship information is used to characterize the physical storage unit and the row set structure corresponding to the target data file, and the segment position in the row set structure, wherein one physical storage unit corresponds to multiple row set structures, and one row set structure corresponds to multiple segment positions; Determine the file path and hash value of the source file corresponding to the target data file; Determine a target timestamp corresponding to the target data file, wherein the target timestamp is used to represent the time when the target data file is updated to the row set structure corresponding to the target data import request; The hierarchical relationship data, the file path, the hash value, and the target timestamp are stored as the historical import information.

3. The data import method according to claim 2, wherein: According to the historical import information, determining the source file in the source file set that has been successfully imported into the target database as the first source file includes: Determine the file path and hash value of the source file in the source file set, and judge whether there is a historical source file in the historical import information whose file path and hash value are consistent with those of the source file in the source file set; If the historical source file whose file path and hash value are consistent with the source file exists in the historical import information, the target data file corresponding to the historical source file in the historical import information is obtained. The corresponding target timestamp; Determining the retention time of the target data file in the recycle bin according to the target timestamp, and judging whether the retention time is greater than a preset retention time threshold; When the retention time is not greater than the preset retention time threshold, the source file is determined as the first source file.

4. The data import method according to claim 3, wherein: Updating the target data file to the row set structure corresponding to the target data import request includes: According to the hierarchical relationship information corresponding to the first source file, the row set structure is created at a position corresponding to the hierarchical relationship information in the target database; The target data file generated by the first source file in the recycle bin during historical import is stored in the segment position corresponding to the hierarchical relationship information in the row set structure, and the file name of the target data file is updated according to the identifier of the row set structure.

5. The data import method according to claim 4, wherein: The row set structure created at a position in the target database corresponding to the hierarchical relationship information has the same hierarchical relationship as the row set structure stored when the first source file is historically imported.

6. The data import method according to claim 3, further comprising: If the historical source file whose file path and hash value are consistent with those of the source file does not exist in the historical import information, or if the retention time is greater than the preset retention time threshold, the source file is determined as the second source file; Generate the target data file corresponding to the second source file, and store the target data file in the target database.

7. The data import method according to claim 6, further comprising: The hierarchical relationship information corresponding to the target data file corresponding to the second source file, the target timestamp, and the file path and the hash value of the second source file are stored as the historical import information.

8. The data import method according to any one of claims 1 to 7, further comprising: If the source file that fails to be imported exists in the source file set, determining that the target data import request fails to be executed; In the source file set corresponding to the target data import request in the target database, the target data files corresponding to the successfully imported source files are cleaned up into the recycle bin.

9. The data import method according to any one of claims 2 to 6, wherein: The information recorded in the historical import information includes at least one of the following: the unique label of the target data import request, the file path and hash value of the imported source file, the execution status of the imported source file, the target timestamp when the source file is successfully imported, and the hierarchical relationship information and file path of the target data file corresponding to the source file.

10. A data import device, comprising: A file acquisition module, configured to acquire a source file set corresponding to a target data import request when a target database receives the target data import request, wherein the target data import request is used to request to import source files in the source file set outside the target database into the target database; A file matching module, configured to determine, according to historical import information, the source file in the source file set that has been successfully imported into the target database as the first source file, wherein the historical import information is used to record the import status of the historical source file corresponding to the historical data import request received by the target database; A quick import module is used to obtain the target data file generated by the first source file in the recycle bin of the target database during the historical import, and update the target data file to the row set structure corresponding to the target data import request, wherein the recycle bin stores the historical source files successfully imported from the historical source file set corresponding to the historical data import request that failed to execute the target database.

11. An electronic device, comprising: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes the data import method described in any one of claims 1 to 9 when running.

12. A non-volatile storage medium comprising a stored computer program, wherein: The device where the non-volatile storage medium is located executes the data import method described in any one of claims 1 to 9 by running the computer program.

13. A computer program product comprising a computer program, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Mass data processing method, database server and application server

    CN103942287A

  • Data transmission anomaly processing method and device, electronic equipment and storage medium

    CN107918564A

  • Data import method, device and equipment of Greenplum database

    CN112711627A

  • Data importing method and device, electronic equipment and nonvolatile storage medium

    CN117708210A

  • Method and apparatus for managing data exchange among systems in a network

    US20020073236A1