Batch data extraction method, device, equipment and storage medium
Patent Information
- Application Number
- CN202610990338.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]本申请的主要目的在于提供一种批量数据的提取方法、装置、设备以及存储介质,旨在解决如何提高批量数据的提取效率的技术问题
[0016]本申请提供了一种批量数据的提取方法,本申请获取数据提取请求,所述数据提取请求包含业务类型和日期范围;根据所述数据提取请求统计每个分组标识对应的分组数据量;获取续传信息,根据所述数据提取请求,所述分组数据量、所述续传信息以及所述分组标识对业务数据库进行分次查询,得到查询到的单次数据库实体集合;将所述单次数据库实体集合转换为多个目标格式文件;将所述多个目标格式文件打包成压缩包,并将所述压缩包上传至存储系统。本申请通过续传信息和分组标识对业务数据库进行分次查询,每次查询仅获取满足预设单次查询数量要求的部分数据,使得单次仅处理部分数据,避免了数据一次性全部加载所导致的内存资源紧张和查询超时,从而提高了批量数据的提取效率。
Smart Images

Figure CN122838440A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, device, and storage medium for extracting batch data. Background Technology
[0002] In existing large-scale data extraction solutions, the data extraction module typically queries all the data to be extracted directly from the business database and loads it into memory all at once for processing and transformation. With the continuous growth of business data volume, the amount of data returned in a single query can reach hundreds of millions of records, leading to frequent problems such as database query timeouts, memory overflows, and excessively long processing times. Therefore, improving the efficiency of batch data extraction remains a problem that needs to be solved.
[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of this application is to provide a method, apparatus, device, and storage medium for extracting batch data, aiming to solve the technical problem of how to improve the efficiency of batch data extraction.
[0005] To achieve the above objectives, this application proposes a method for extracting batch data, the method comprising: Obtain a data extraction request, which includes a business type and a date range; The amount of group data corresponding to each group identifier is calculated based on the data extraction request. Obtain resume information, and query the business database in several steps based on the data extraction request, the group data volume, the resume information, and the group identifier to obtain the set of database entities retrieved in a single step. Convert the single database entity set into multiple target format files; The multiple target format files are packaged into a compressed file, and the compressed file is uploaded to the storage system.
[0006] In one embodiment, the step of calculating the amount of group data corresponding to each group identifier based on the data extraction request includes: According to the service type, obtain the preset configuration item information, which includes at least the data table name, group identifier field name, date field name, sending status field name, and resume sequence number field name; The amount of data successfully sent by each group identifier within the date range is calculated based on the date range and the configuration item information, and this amount is taken as the data volume of the group corresponding to each group identifier.
[0007] In one embodiment, the step of obtaining resume information includes: Read the resume context information from the cache system, the resume context information including the maximum resume sequence number of the processed data; If the resume context information is successfully read, then the resume context information is used as the resume information. If it does not exist, the resume information will be set to empty.
[0008] In one embodiment, the step of performing multiple queries on the business database based on the data extraction request, the group data volume, the resume information, and the group identifier to obtain the queried single database entity set includes: The current pending start sequence number of the resume transmission is determined based on the resume transmission information; A single query is performed on the business database based on the group identifier, the starting resume sequence number, and the preset single query quantity requirement to obtain a single database entity and update the resume information. The query is completed based on the amount of grouped data. If the query is not completed, the query continues based on the updated continuation information to obtain the single database entity until the query is completed. The single database entities obtained from the query are determined as a single database entity set.
[0009] In one embodiment, the step of performing a single query on the business database based on the group identifier, the starting and continuing sequence number, and the preset single query quantity requirement to obtain a single database entity includes: Obtain the associated table configuration information of the basic information table in the business database. The associated table configuration information includes at least the name of the associated table and the associated fields between the tables. The initial basic data is retrieved from the basic information table based on the group identifier, the starting and continuing sequence number, and the preset single query quantity requirement. Based on the associated fields and the preset single query quantity requirement, query the corresponding associated data from the associated table; The basic data and the associated data are identified as a single database entity.
[0010] In one embodiment, the step of packaging the plurality of target format files into a compressed package includes: Obtain the preset number of files to be packaged, and divide the multiple target format files into multiple file groups according to the number of files to be packaged; Compress each file group separately to generate a compressed file group package.
[0011] In one embodiment, after the step of uploading the compressed package to the storage system, the method further includes: Receive download instructions; The target service type is determined based on the download instruction; The compressed package of the target business type is compressed a second time to obtain the final compressed package.
[0012] Furthermore, to achieve the above objectives, this application also proposes a batch data extraction device, which includes: The acquisition module is used to acquire data extraction requests, which include business type and date range; The statistics module is used to count the amount of group data corresponding to each group identifier based on the data extraction request; The query module is used to obtain resume information. Based on the data extraction request, the group data volume, the resume information, and the group identifier, the business database is queried in multiple steps to obtain a single set of database entities. The conversion module is used to convert the single database entity set into multiple target format files; The compression module is used to package the multiple target format files into a compressed package and upload the compressed package to the storage system.
[0013] In addition, to achieve the above objectives, this application also proposes a batch data extraction device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the batch data extraction method described above.
[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the batch data extraction method described above.
[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the batch data extraction method described above.
[0016] This application provides a method for extracting batch data. The method involves obtaining a data extraction request, which includes a business type and a date range; calculating the data volume of each group corresponding to a group identifier based on the data extraction request; obtaining resume information; and performing multiple queries on the business database based on the data extraction request, the group data volume, the resume information, and the group identifier to obtain a single set of database entities; converting the single set of database entities into multiple target format files; packaging the multiple target format files into a compressed package; and uploading the compressed package to a storage system. This application uses resume information and group identifiers to perform multiple queries on the business database, with each query only retrieving a portion of the data that meets a preset single query quantity requirement. This ensures that only a portion of the data is processed at a time, avoiding memory resource constraints and query timeouts caused by loading all data at once, thereby improving the efficiency of batch data extraction. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the method for extracting batch data in this application, as provided in Embodiment 1. Figure 2 This is a schematic diagram of the quantity statistics process provided in Embodiment 1 of the batch data extraction method of this application; Figure 3 This is a schematic diagram of the data query and file generation process provided in Embodiment 1 of the batch data extraction method of this application; Figure 4 This is a flowchart illustrating Embodiment 2 of the method for extracting batch data in this application. Figure 5 This is a schematic diagram of the specific data query process provided in Embodiment 2 of the method for extracting batch data in this application; Figure 6 A simplified flowchart illustrating the batch data extraction method provided in Embodiment 1 of this application; Figure 7 This is a schematic diagram of the module structure of the batch data extraction device according to an embodiment of this application; Figure 8 This is a schematic diagram of the hardware operating environment involved in the batch data extraction method in this application embodiment.
[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0023] This application obtains a data extraction request, which includes a business type and a date range; calculates the amount of group data corresponding to each group identifier based on the data extraction request; obtains resume information; and queries the business database in multiple sessions based on the data extraction request, the amount of group data, the resume information, and the group identifier to obtain a single set of database entities; converts the single set of database entities into multiple target format files; packages the multiple target format files into a compressed package, and uploads the compressed package to the storage system.
[0024] In existing large-scale data extraction solutions, the data extraction module typically queries all the data to be extracted directly from the business database and loads it into memory all at once for processing and transformation. With the continuous growth of business data volume, the amount of data returned in a single query can reach hundreds of millions of records, leading to frequent problems such as database query timeouts, memory overflows, and excessively long processing times. Therefore, improving the efficiency of batch data extraction remains a problem that needs to be solved.
[0025] This application performs segmented queries on the business database by using continuation information and group identifiers. Each query retrieves only a portion of the data that meets the preset single query quantity requirement, thus processing only a portion of the data at a time. This avoids memory resource constraints and query timeouts caused by loading all the data at once, thereby improving the efficiency of batch data extraction.
[0026] All user-related data involved in this application were obtained with the user's permission or consent; that is, when this application is applied to specific products or technologies, user permission is required to obtain and process the relevant data, and the processing of the relevant data must comply with the relevant laws, regulations and regulatory standards of the relevant countries and regions.
[0027] Based on this, embodiments of this application provide a method for extracting batch data, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the batch data extraction method of this application.
[0028] In this embodiment, the method for extracting batch data includes steps S10 to S50: Step S10: Obtain a data extraction request, wherein the data extraction request includes a business type and a date range; It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or batch data extraction device capable of performing the above functions. The following description uses a batch data extraction device as an example to illustrate this embodiment and the subsequent embodiments.
[0029] It's important to note that data extraction requests are initiated by users through the front-end interface. Users select the business type and enter the date range, then click the submit button, at which point the system receives the data extraction request. The business type refers to the business category of the data, which varies depending on the scenario. In the financial sector, the business type could be transaction details, customer information management data, etc. The date range specifies the time period during which the data to be extracted occurred. Alternatively, data extraction requests can be automatically triggered by a scheduled task. The system automatically generates a data extraction request containing the business type and date range according to a preset scheduling cycle.
[0030] Step S20: Calculate the amount of group data corresponding to each group identifier based on the data extraction request; It's important to note that in batch data extraction, the total data volume is often very large, so data is typically extracted in groups. Group identifiers are used to group the data; these identifiers can be organization numbers, line codes, or department numbers, etc. Group identifiers distinguish different groups of data and allow for the calculation of the data volume in each group, facilitating subsequent queries. Furthermore, the data volumes of each group can be added together to obtain the total data volume, which the system will store. (See reference...) Figure 2 , Figure 2 This is a flowchart illustrating the quantity statistics process. Figure 2 The process begins by receiving a query request: business type and date range; then, it queries the relevant configuration items based on the business type and switches to the corresponding data source; next, it calculates the data volume for each organization number within the corresponding business based on the table and field information of the configuration items; finally, it sums all the data volumes and records them.
[0031] In one feasible approach, the step of calculating the amount of group data corresponding to each group identifier based on the data extraction request includes: obtaining preset configuration item information according to the service type, wherein the configuration item information includes at least the data table name, group identifier field name, date field name, sending status field name, and continuation sequence number field name; and calculating the amount of data successfully sent by each group identifier within the date range based on the date range and the configuration item information, which is taken as the amount of group data corresponding to each group identifier.
[0032] It should be noted that the quantity statistics are based on specific query conditions, counting the amount of data that meets the conditions to provide basic support for subsequent data queries. In this embodiment, the statistics are grouped by institution number, counting the amount of data successfully sent by each institution number within a specified time period. Therefore, the quantity statistics include at least three variables: date, sending status, and institution number. Considering the need for data continuation, a continuation field can be added. After completing the quantity statistics, the data volume for each group is obtained.
[0033] Step S30: Obtain resume information. Based on the data extraction request, the group data volume, the resume information, and the group identifier, perform multiple queries on the business database to obtain the set of single database entities queried. It's important to note that data querying retrieves data that meets specific query criteria. This process requires several things to be done. First, data with different identifiers must be stored separately, with each identifier corresponding to different data in a separate file. The identifier is a higher-level classification than the group identifier. Group identifiers, such as organization numbers, may have several organization numbers corresponding to one identifier, or one organization number to one identifier. Storing data separately by identifier allows grouping similar data together to meet query needs. Alternatively, data can be stored directly based on the group identifier. The file format can be determined according to user requirements or a preset format, such as XML. Second, the stored data is typically data from successfully sent data; therefore, query results can include a "successfully sent" condition. Finally, the number of data entries in each file needs to be limited. Due to the large volume of data, the number of data entries in each file must be controlled within a certain range.
[0034] It should be noted that the resume information is used to record the position when the last query was interrupted, including the maximum resume sequence number processed (such as an auto-incrementing primary key ID or a timestamp), to ensure that the query can continue from the breakpoint after the task is interrupted, avoiding data duplication or omission. The system performs multiple queries on the business database according to the group identifier and the starting resume sequence number in the resume information, for example, if the last query processed up to the 50,000th record, according to the preset number of queries per session. Each query yields a batch of database entities, i.e., a single-session database entity set, until all data under that group identifier has been queried, resulting in a single-session database entity set.
[0035] In one feasible approach, the step of obtaining the resume information includes: reading resume context information from the caching system, the resume context information containing the maximum resume sequence number of the processed data; if the resume context information is successfully read, then the resume context information is used as the resume information; if it does not exist, then the resume information is set to empty.
[0036] It's important to note that the resume context information can be read from a caching system (such as Redis). This resume context information is stored as key-value pairs. The key name can include a fixed prefix and a business type identifier, such as "EXPORT:CONTINUE:Transaction Details Report". The value stores the maximum resume sequence number of the processed data, such as 50000. This sequence number corresponds to a unique, incrementing sequence number of a record in the database table (such as the value of the auto-incrementing primary key seq_id). The system uses this information to determine which data was processed when the last query was interrupted, thus enabling query continuation. If the resume context information is successfully read from the caching system (i.e., the corresponding key-value pair exists in the cache), it indicates that there was an incomplete query task. The system uses this resume context information as the resume information for the current query. In this case, the system will continue the query from the position after this resume sequence number, instead of starting from the beginning. This avoids data duplication and saves time on repeated queries. If the system fails to read the resume context information from the caching system (i.e., the corresponding key-value pair does not exist in the cache), it means there is no incomplete query task, and it is a completely new data retrieval task. At this point, the system sets the resume information to empty, indicating that the query needs to start from the first data record.
[0037] Step S40: Convert the single database entity set into multiple target format files; It should be noted that in this embodiment, the target format file is an XML file. This step converts the data generated by the data query into an XML file. This is the most critical and complex part of the entire file generation process. Since the fields for each data item are different, conversion methods from database entity classes to XML entity classes need to be implemented separately for over 100 data types. When generating the files, the maximum data size limit for a single XML file must also be considered to prevent the files from becoming too large, and the generated files are named sequentially using serial numbers. See reference [link / reference]. Figure 3 , Figure 3 This is a schematic diagram of the data query and file generation process. Figure 3 The process first retrieves the data from a single query, obtaining the database entity. Then, it converts the database entity into a corresponding XML entity. Next, it generates an XML file sequence number and then generates the XML file. Finally, it checks if a next page exists. If it does, it returns to the step of retrieving the data from the single query and obtaining the database entity; otherwise, it terminates.
[0038] Step S50: Package the multiple target format files into a compressed package and upload the compressed package to the storage system.
[0039] It's important to note that after generating the files, due to the large number of files, they need to be packaged and uploaded. File packaging and uploading refers to packaging the generated files into a ZIP archive according to a certain quantity, uploading this archive to object storage, and simultaneously obtaining the corresponding download links. This step is applicable to all types of data and has universality. The final output is all successfully sent files, which need to be clearly divided according to business modules. The file download step involves downloading the files uploaded in the previous step in batches according to business modules and then reassembling them into packages to complete the delivery. Specifically, the download links for all data in the corresponding module are queried based on the business type, and the files are downloaded in batches to the same folder. After the files for the same business module have been downloaded, this folder is compressed to generate a compressed archive containing all the sent files for that business module.
[0040] In one feasible approach, the step of packaging the plurality of target format files into a compressed package includes: obtaining a preset packaging quantity, dividing the plurality of target format files into multiple file groups according to the packaging quantity; compressing each file group separately to generate a compressed file group package.
[0041] It's important to note that the system retrieves a preset number of files to package, such as 100 files per package. This preset number can be set by the user through a configuration file or dynamically adjusted as needed. The system first counts the total number of target format files generated, then groups this total according to the preset number: files 1-100 form the first group, files 101-200 form the second group, and files 201-250 form the third group. Each group contains no more than the preset number of target format files. In another implementation, the system can also group files based on size (e.g., each compressed file cannot exceed a specific size) to meet the limitations of different storage systems. Then, each file group is compressed separately to generate a compressed file group package.
[0042] In one feasible approach, after the step of uploading the compressed package to the storage system, the method further includes: receiving a download instruction; determining the target service type based on the download instruction; and performing secondary compression on the compressed package of the target service type to obtain a final compressed package.
[0043] It's important to note that after completing all data extraction, file generation, packaging, and uploading, the system awaits download instructions from users or downstream systems. Specifically, the front-end interface displays a list of downloadable business modules. Users generate download instructions by selecting the desired modules and clicking the download button. The system parses the target business type information from the received download instructions, determining the business type to be processed in this download operation. This allows the system to locate all compressed packages corresponding to the target business type in the storage system. Finally, the system queries the object storage system for all compressed packages and their download links corresponding to the target business type. These compressed packages are then downloaded in batches to a local temporary folder. After downloading, the system recompresses the entire temporary folder (a second compression) to generate a single, complete final compressed package for users to download all at once.
[0044] This embodiment obtains a data extraction request, which includes a business type and a date range; it calculates the data volume corresponding to each group identifier based on the data extraction request; it obtains resume information; and based on the data extraction request, the data volume of each group, the resume information, and the group identifier, it performs multiple queries on the business database to obtain a single set of database entities; it converts the single set of database entities into multiple target format files; it packages the multiple target format files into a compressed package and uploads the compressed package to the storage system. This embodiment performs multiple queries on the business database using resume information and group identifiers, with each query only retrieving a portion of the data that meets the preset single query quantity requirement. This ensures that only a portion of the data is processed at a time, avoiding memory resource constraints and query timeouts caused by loading all data at once, thereby improving the efficiency of batch data extraction.
[0045] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 4 Step S30 also includes steps S301 to S304: Step S301: Determine the current pending start sequence number of the resume transmission based on the resume transmission information; It should be noted that the system first reads the resume information. If the resume information is not empty, the system extracts the resume sequence number of the processed data from it and uses this sequence number as the starting resume sequence number for the current pending data. For example, if the maximum resume sequence number recorded in the resume information is 50000, the query will start from the 50001st record. If the resume information is empty, the starting resume sequence number is set to a preset initial value (such as 0), indicating that the query will start from the first record.
[0046] Step S302: Perform a single query on the business database according to the group identifier, the starting resume sequence number and the preset single query quantity requirement to obtain the single database entity, and update the resume information; It should be noted that when executing a single query, the system first filters the data in the current group using the group identifier field in the Structured Query Language (SQL) statement, filters the data after the starting position using the continuation sequence number field, uses a preset single query limit as the query count limit, and sorts the data in ascending order according to the continuation sequence number field. After the query, a batch of database entities is obtained, i.e., the single database entities. Subsequently, the system extracts the largest continuation sequence number from this batch of database entities, such as 55000, updates the continuation information to this largest continuation sequence number, and saves it to the cache system so that it can be resumed from this position if the task is interrupted.
[0047] In one feasible approach, the step of performing a single query on the business database based on the group identifier, the start / continue sequence number, and a preset single query quantity requirement to obtain a single database entity includes: obtaining the association table configuration information of the basic information table in the business database, wherein the association table configuration information includes at least the name of the association table and the association fields between the tables; querying the initial basic data from the basic information table based on the group identifier, the start / continue sequence number, and the preset single query quantity requirement; querying the corresponding associated data from the association table based on the association fields and the preset single query quantity requirement; and determining the basic data and the associated data as the single database entity.
[0048] It should be noted that when business data for a particular business type involves multiple table joins, such as data from a business module distributed across three tables: a basic information table, an additional information table, and a management information table, the system will first retrieve the configuration information for the join tables. This configuration information can be pre-defined in the configuration items, including but not limited to: the names of the join tables and the join fields between the tables. Then, the system queries data from the basic information table. Specifically, the system filters data for the current group based on the group identifier, filters data after the starting position based on the start / resume sequence number, limits the number of records queried based on a preset single query quantity (e.g., 5000 records), and sorts the data in ascending order by the resume sequence number field. After executing the query, a batch of initial basic data is obtained. This basic data contains the join key values for each record, which are used to supplement the corresponding join data in subsequent queries from the join tables. When querying the basic information table, filtering conditions such as date range and sending status are also required to ensure that only data within the specified date range and successfully sent is extracted. These filtering conditions also originate from the configuration item information. After obtaining the initial basic data, the join key values for all records in this batch of basic data are extracted. Then, the system uses these related key values as query conditions to batch query the corresponding related data from each related table. It also controls the amount of related data queried in each query based on a preset single query quantity requirement. If the amount of related data is also large, the system can query the related data in batches. For example, it can query collateral information from the supplementary information table and approval records from the management information table. In this way, data scattered across multiple tables is recombined according to business relationships through related fields, maintaining data integrity and consistency. Finally, the obtained basic data and related data are merged. Specifically, for each piece of basic data, a related record with the same related key value is found from the queried related data. Then, the field values of the basic data are combined with the field values of the related data to form a complete database entity. This complete database entity contains all the information of that record. It should be noted that during the merging process, if a piece of basic data does not have a corresponding related record in a certain related table, the value of the corresponding related field is set to null.
[0049] Step S303: Determine whether the query is complete based on the amount of grouped data. If the query is not complete, continue the query based on the updated continuation information to obtain the single database entity until the query is complete. It should be noted that the system determines whether the query is complete by comparing the number of data entries currently queried with the group data volume. Specifically, the system sets a cumulative value for the number of data entries queried in the current group. After each single query, the number of entities obtained in this query is added to this cumulative value. When the cumulative value equals the group data volume, it means that all data corresponding to the current group identifier has been queried, and the query will not continue. If the cumulative value is less than the group data volume, it means that there is still data under that group identifier. The system uses the updated continuation information, i.e., the maximum continuation sequence number obtained in this query, as the starting position for the next query, and continues to execute the single query, repeating the above process until all data of that group identifier has been queried. In another implementation, the system can also infer whether the query is complete by judging whether the number of entities returned by this query is less than the preset single query limit. When the number of target format files generated by this request reaches the preset single request file limit, even if the current group has not been queried, the system will save the continuation information and terminate the current query request, waiting for the next request to continue processing.
[0050] Step S304: Determine the single database entities obtained from the query as a single database entity set.
[0051] It should be noted that the system combines the individual database entities obtained from each query into a single-query database entity set according to the query order. This set contains all database entities actually retrieved in this request for the current group identifier. For example, if this request executes four single queries: the first query retrieves 5000 records, the second 5000 records, the third 5000 records, and the fourth 5000 records, then the single-query database entity set contains 20000 database entities. This set will be used in subsequent target format file conversion steps. It should be noted that "single" in the single-query database entity set refers to the result set returned by each independent database query operation. In practical implementation, the entity set obtained from each query can be converted immediately after each query, without waiting for all queries to complete, thus achieving streaming processing and reducing memory usage.
[0052] For reference Figure 5 , Figure 5 This is a schematic diagram illustrating the specific process of data query. Figure 5In the process, after obtaining the group identifier, it is determined whether there is any continuation information. If so, the business data is queried in ascending order based on the group identifier and the continuation information. If not, the business data is queried in ascending order based on the group identifier. Then, each piece of data is generated into a file. Next, it is determined whether there is any more data for the current group identifier. If there is, it is determined whether the number of files requested in this process exceeds the limit. If the limit is exceeded, the continuation information is preserved and the process ends. If the limit is not exceeded, the process returns to the step of obtaining the group identifier. If there is no more data for the current group identifier, it means that the data for the current group identifier has been processed. The continuation information can be cleared, and it is determined whether the conditions of exceeding the limit for the generated files or completing data processing are met. If the conditions are met, the process ends. If not, the process returns to the step of obtaining the group identifier.
[0053] This embodiment determines the current pending start sequence number of the resume transmission based on the resume transmission information; performs a single query on the business database based on the group identifier, the start sequence number, and the preset single query quantity requirement to obtain single database entities and update the resume transmission information; determines whether the query is complete based on the group data volume; if not, continues querying based on the updated resume transmission information to obtain single database entities until the query is complete; and determines the obtained single database entities as a single database entity set. This embodiment, by querying in stages and updating the resume transmission information in real time, allows the system to process only a portion of the data in a single query, reducing the data volume and memory consumption of a single query, thereby improving the execution efficiency of batch data extraction.
[0054] For example, to help understand the implementation flow of the batch data extraction method obtained by combining this embodiment with the first embodiment described above, please refer to... Figure 6 , Figure 6 A simplified flowchart of a batch data extraction method is provided. Specifically, the steps are as follows: receiving data extraction instructions, quantity statistics, data query, file generation, file packaging and uploading, and file download.
[0055] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the batch data extraction method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0056] This application also provides a batch data extraction device; please refer to... Figure 7 The batch data extraction device includes: The acquisition module 10 is used to acquire a data extraction request, wherein the data extraction request includes a business type and a date range; The statistics module 20 is used to count the amount of group data corresponding to each group identifier based on the data extraction request; The query module 30 is used to obtain resume information and, based on the data extraction request, the group data volume, the resume information, and the group identifier, performs multiple queries on the business database to obtain a single set of database entities. Conversion module 40 is used to convert the single database entity set into multiple target format files; Compression module 50 is used to package the plurality of target format files into a compressed package and upload the compressed package to the storage system.
[0057] This application obtains a data extraction request, which includes a business type and a date range; it calculates the amount of group data corresponding to each group identifier based on the data extraction request; it obtains resume information; and based on the data extraction request, the amount of group data, the resume information, and the group identifier, it performs multiple queries on the business database to obtain a single set of database entities; it converts the single set of database entities into multiple target format files; it packages the multiple target format files into a compressed package and uploads the compressed package to the storage system. This application performs multiple queries on the business database using resume information and group identifiers, with each query only retrieving a portion of the data that meets the preset single query quantity requirement. This ensures that only a portion of the data is processed at a time, avoiding memory resource constraints and query timeouts caused by loading all data at once, thereby improving the efficiency of batch data extraction.
[0058] In one embodiment, the statistics module 20 is further configured to obtain preset configuration item information according to the service type. The configuration item information includes at least a data table name, a group identifier field name, a date field name, a sending status field name, and a continuation sequence number field name. The number of data successfully sent for each group identifier within the date range is calculated based on the date range and the configuration item information, and is used as the number of group data corresponding to each group identifier.
[0059] In one embodiment, the query module 30 is further configured to read resume context information from the cache system, the resume context information containing the maximum resume sequence number of the processed data; if the resume context information is successfully read, the resume context information is used as the resume information; if it does not exist, the resume information is set to empty.
[0060] In one embodiment, the query module 30 is further configured to: determine the current pending start resume sequence number based on the resume information; perform a single query on the business database based on the group identifier, the start resume sequence number, and a preset single query quantity requirement to obtain a single database entity, and update the resume information; determine whether the query is complete based on the group data volume; if the query is not complete, continue querying based on the updated resume information to obtain a single database entity, until the query is complete; and determine the obtained single database entities as a single database entity set.
[0061] In one embodiment, the query module 30 is further configured to obtain the associated table configuration information of the basic information table in the business database, wherein the associated table configuration information includes at least the name of the associated table and the associated fields between the tables; query the initial basic data from the basic information table according to the group identifier, the starting and continuing sequence number and the preset single query quantity requirement; query the corresponding associated data from the associated table according to the associated fields and the preset single query quantity requirement; and determine the basic data and the associated data as a single database entity.
[0062] In one embodiment, the compression module 50 is further configured to obtain a preset number of packages, divide the plurality of target format files into a plurality of file groups according to the number of packages, and compress each file group separately to generate a compressed file group package.
[0063] In one embodiment, the compression module 50 is further configured to receive a download instruction; determine the target service type according to the download instruction; and perform secondary compression on the compressed package of the target service type to obtain the final compressed package.
[0064] The batch data extraction device provided in this application, employing the batch data extraction method in the above embodiments, can solve the technical problem of how to improve the efficiency of batch data extraction. Compared with the prior art, the beneficial effects of the batch data extraction device provided in this application are the same as those of the batch data extraction method provided in the above embodiments, and other technical features in the batch data extraction device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0065] This application provides a batch data extraction device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the batch data extraction method in the first embodiment described above.
[0066] The following is for reference. Figure 8This diagram illustrates a structural schematic of a batch data extraction device suitable for implementing embodiments of this application. The batch data extraction device in embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 8 The illustrated batch data extraction device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0067] like Figure 8 As shown, the batch data retrieval device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read Only Memory) 1002 or a program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the batch data retrieval device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touch screens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the bulk data extraction device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows bulk data extraction devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.
[0068] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0069] The batch data extraction device provided in this application, employing the batch data extraction method described in the above embodiments, can solve the technical problem of how to improve the efficiency of batch data extraction. Compared with the prior art, the beneficial effects of the batch data extraction device provided in this application are the same as those of the batch data extraction method provided in the above embodiments, and other technical features of the batch data extraction device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0070] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0071] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0072] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the batch data extraction method in the above embodiments.
[0073] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0074] The aforementioned computer-readable storage medium may be included in a batch data extraction device; or it may exist independently and not be assembled into a batch data extraction device.
[0075] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by a batch data extraction device, the batch data extraction device performs the following actions: 1. Obtains a data extraction request, the data extraction request including a business type and a date range; 2. Calculates the amount of group data corresponding to each group identifier based on the data extraction request; 3. Obtains resume information, and performs multiple queries on the business database based on the data extraction request, the amount of group data, the resume information, and the group identifier to obtain a single set of database entities; 4. Converts the single set of database entities into multiple target format files; 5. Packages the multiple target format files into a compressed package and uploads the compressed package to the storage system.
[0076] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0077] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0078] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0079] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described batch data extraction method, thereby solving the technical problem of how to improve the efficiency of batch data extraction. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the batch data extraction method provided in the above embodiments, and will not be repeated here.
[0080] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the batch data extraction method described above.
[0081] The computer program product provided in this application can solve the technical problem of how to improve the efficiency of batch data extraction. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the batch data extraction method provided in the above embodiments, and will not be repeated here.
[0082] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for extracting batch data, characterized in that, The method includes: Obtain a data extraction request, which includes a business type and a date range; The amount of group data corresponding to each group identifier is calculated based on the data extraction request. Obtain resume information, and query the business database in several steps based on the data extraction request, the group data volume, the resume information, and the group identifier to obtain the set of database entities retrieved in a single step. Convert the single database entity set into multiple target format files; The multiple target format files are packaged into a compressed file, and the compressed file is uploaded to the storage system.
2. The method as described in claim 1, characterized in that, The step of calculating the amount of group data corresponding to each group identifier based on the data extraction request includes: According to the service type, obtain the preset configuration item information, which includes at least the data table name, group identifier field name, date field name, sending status field name, and resume sequence number field name; The amount of data successfully sent by each group identifier within the date range is calculated based on the date range and the configuration item information, and this amount is taken as the data volume of the group corresponding to each group identifier.
3. The method as described in claim 1, characterized in that, The steps for obtaining resume information include: Read the resume context information from the cache system, the resume context information including the maximum resume sequence number of the processed data; If the resume context information is successfully read, then the resume context information is used as the resume information. If it does not exist, the resume information will be set to empty.
4. The method as described in claim 1, characterized in that, The step of performing multiple queries on the business database based on the data extraction request, the group data volume, the resume information, and the group identifier to obtain the queried single database entity set includes: The current pending start sequence number of the resume transmission is determined based on the resume transmission information; A single query is performed on the business database based on the group identifier, the starting resume sequence number, and the preset single query quantity requirement to obtain a single database entity and update the resume information. The query is completed based on the amount of grouped data. If the query is not completed, the query continues based on the updated continuation information to obtain the single database entity until the query is completed. The single database entities obtained from the query are determined as a single database entity set.
5. The method as described in claim 4, characterized in that, The step of performing a single query on the business database based on the group identifier, the starting and continuing sequence number, and the preset single query quantity requirement to obtain a single database entity includes: Obtain the associated table configuration information of the basic information table in the business database. The associated table configuration information includes at least the name of the associated table and the associated fields between the tables. The initial basic data is retrieved from the basic information table based on the group identifier, the starting and continuing sequence number, and the preset single query quantity requirement. Based on the associated fields and the preset single query quantity requirement, query the corresponding associated data from the associated table; The basic data and the associated data are identified as a single database entity.
6. The method as described in claim 1, characterized in that, The step of packaging the plurality of target format files into a compressed package includes: Obtain the preset number of files to be packaged, and divide the multiple target format files into multiple file groups according to the number of files to be packaged; Compress each file group separately to generate a compressed file group package.
7. The method as described in claim 1, characterized in that, After the step of uploading the compressed package to the storage system, the method further includes: Receive download instructions; The target service type is determined based on the download instruction; The compressed package of the target business type is compressed a second time to obtain the final compressed package.
8. A batch data extraction device, characterized in that, The device includes: The acquisition module is used to acquire data extraction requests, which include business type and date range; The statistics module is used to count the amount of group data corresponding to each group identifier based on the data extraction request; The query module is used to obtain resume information. Based on the data extraction request, the group data volume, the resume information, and the group identifier, the business database is queried in multiple steps to obtain a single set of database entities. The conversion module is used to convert the single database entity set into multiple target format files; The compression module is used to package the multiple target format files into a compressed package and upload the compressed package to the storage system.
9. A batch data extraction device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the batch data extraction method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the batch data extraction method as described in any one of claims 1 to 7.