Target file generation system and device
Through the combined architecture of the big data platform and the data service platform, and by utilizing partitioning and bucketing processing, the problem of low target file generation efficiency is solved, and efficient and fast target file generation is achieved.
Patent Information
- Application Number
- CN202510739887.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-23
AI Technical Summary
The existing target file generation system has insufficient performance when processing big data, resulting in low generation efficiency and difficulty in meeting the requirements of high concurrency and high computing performance.
It adopts a combined architecture of big data platform and data service platform, including data buffer layer, data operation layer, data model layer and data promotion layer. Through partitioning and bucketing processing, combined with script file conversion, it realizes efficient generation of target files.
It improves the efficiency of target file generation, meets the requirements of high concurrency and high computing performance, and ensures the speed and accuracy of data processing.
Smart Images

Figure CN120687514A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a system and device for generating a target file. Background Art
[0002] At present, with the development of diversified businesses, the initial data generated by the businesses is also increasing continuously, and the target file generation system is also required to have high concurrency performance and high computing performance.
[0003] In related technologies, a target file generation system reads initial data in batches from a database. The performance of the system and the efficiency of reporting the initial data depend on the performance of the database used by the system. When the initial data to be processed is large, the performance of the above system is insufficient and the processing time is long, resulting in a technical problem of low efficiency in generating the target file.
[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0005] The embodiments of the present application provide a system and device for generating a target file, so as to at least solve the technical problem of low efficiency in generating the target file.
[0006] According to one aspect of an embodiment of the present application, a target file generation system is provided. The system may include: a big data platform for receiving initial data from a database and, in response to data processing instructions, converting the initial data into an initial file; a data service platform for responding to file integration instructions and converting the initial file into a target file in a target format corresponding to the target file to be generated; and a data file platform for storing the target file.
[0007] Optionally, the data service platform includes: a data buffer layer for extracting initial data from a database; a data operation layer for extracting multiple common data from the initial data, and associating the multiple common data to obtain associated data.
[0008] Optionally, the data service platform also includes: a data model layer, which is used to call a general data model to process the associated data to obtain processed associated data, wherein the general data model is established based on the general requirements of the initial file; a data enhancement layer, which is used to call a target data model to process the processed associated data to obtain an initial file, wherein the target data model is established based on the target requirements of the initial file.
[0009] Optionally, the system also includes: a central scheduling platform, used to send data processing instructions to the big data platform, and used to send file integration instructions to the data service platform.
[0010] Optionally, the big data platform is further used to partition the initial file to obtain partitioned files, and / or to bucket the initial file to obtain bucketed files.
[0011] Optionally, the data service platform is further configured to call a script file to convert the partition file into a target file, and / or to convert the bucket file into a target file.
[0012] According to another aspect of an embodiment of the present application, a device for generating a target file is provided. The device may include: a data reading module for reading initial data from a database; a data analysis module for converting the initial data into an initial file in response to a data processing instruction; and a file integration module for converting the initial file into a target file in accordance with a target format corresponding to the target file to be generated in response to the file integration instruction.
[0013] Optionally, the data reading module includes: a data instruction set for storing rules for reading initial data from a database; a parameter configuration set for controlling the speed of reading the initial data; and an instruction editor for editing reading instructions for reading the initial data.
[0014] Optionally, the device further includes: a data storage module for storing at least initial data, initial files, and target files; and an instruction scheduling and analysis module for scheduling data processing instructions and file integration instructions.
[0015] Optionally, the instruction scheduling and analysis module includes: an instruction generation unit for generating data processing instructions and file integration instructions; an instruction sending unit for sending data processing instructions to the data analysis module and sending file integration instructions to the file integration module; an instruction parsing unit for parsing data processing instructions and file integration instructions; and an instruction monitoring unit for monitoring the execution of data processing instructions and file integration instructions.
[0016] In an embodiment of the present application, the big data platform is used to receive initial data from a database and, in response to data processing instructions, convert the initial data into an initial file; the data service platform is used to respond to file integration instructions and convert the initial file into a target file according to a target format corresponding to the target file to be generated; the data file platform is used to store the target file, thereby achieving the technical effect of improving the efficiency of generating the target file, and thereby solving the technical problem of low efficiency in generating the target file. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1 is a schematic diagram of a target file generation system according to an embodiment of the present application;
[0019] Figure 2 is a schematic diagram of a device for generating a target file according to an embodiment of the present application;
[0020] Figure 3 It is a schematic diagram of a data file generation system according to the related art;
[0021] Figure 4 is a schematic diagram of a data file generation system according to an embodiment of the present application;
[0022] Figure 5 is a schematic diagram of a data architecture of a data file generation system according to an embodiment of the present application;
[0023] Figure 6 is a schematic diagram of a data reporting device according to an embodiment of the present application;
[0024] Figure 7 is a schematic block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0027] Figure 1 is a schematic diagram of a target file generation system according to an embodiment of the present application, such as Figure 1As shown, the target file generation system 100 may include: a big data platform 102 , a data service platform 104 and a data file platform 106 .
[0028] The big data platform 102 is used to receive initial data from the database and, in response to data processing instructions, convert the initial data into an initial file.
[0029] Optionally, the database stores initial data. The initial data may be business-related data and may also be referred to as business data, raw data, or target data. The initial data may also be referred to as data files or raw data files. The big data platform 102 preprocesses the initial data, generates initial files, and synchronizes and stores the initial files.
[0030] Optionally, the business data in the database is synchronized to the big data platform 102. The big data platform receives data processing instructions from the unified instruction scheduling platform, performs screening, merging, data logical operations and other processing on the target data, generates an initial file containing the target data, stores and synchronizes it to the data service platform.
[0031] The data service platform 104 is configured to respond to the file integration instruction and convert the initial file into a target file according to a target format corresponding to the target file to be generated.
[0032] Alternatively, the target file may also be referred to as a target data file. The data service platform 104 integrates and processes the initial files generated in the big data platform 102 to obtain the target file.
[0033] Optionally, the data service platform 104 receives file integration instructions from the unified instruction scheduling platform, and classifies, merges, compresses, modifies the name of the initial files processed by the big data platform according to the data file format (target format) and content requirements to form a target data file, and finally transfers the target data file to the data file platform to complete the generation of the target data file.
[0034] The data file platform 106 is used to store target files.
[0035] In an embodiment of the present application, the big data platform 102 is used to receive the initial data of the database and, in response to data processing instructions, convert the initial data into an initial file; the data service platform 104 is used to respond to file integration instructions and convert the initial file into a target file according to the target format corresponding to the target file to be generated; the data file platform 106 is used to store the target file, thereby achieving the technical effect of improving the efficiency of generating the target file, and thereby solving the technical problem of low efficiency in generating the target file.
[0036] The above embodiments of the present application are further described below.
[0037] In some embodiments of the present application, the data service platform includes: a data buffer layer for extracting initial data from a database; a data operation layer for extracting multiple common data from the initial data and associating the multiple common data to obtain associated data.
[0038] Alternatively, common data can be referred to as target data, and associated data can be the data obtained by performing operations such as data association and classification on the common data. The data buffer layer is the first layer of the big data platform and can be used to connect to external databases and store initial data read from external sources. The data operation layer is the second layer of the big data platform and can be used to store data extracted, converted, and loaded from the data buffer layer. This data is related to the target data.
[0039] Optionally, the big data platform reads data from an external database and stores it in the data buffer layer, enabling the loading of external data into the big data platform. The data operation layer extracts target class data from the data buffer layer and performs operations such as data association and classification to improve the speed and efficiency of subsequent target data association operations.
[0040] In some embodiments of the present application, the data service platform also includes: a data model layer, which is used to call a general data model to process the associated data to obtain the processed associated data, wherein the general data model is established based on the general requirements of the initial file; a data enhancement layer, which is used to call a target data model to process the processed associated data to obtain the initial file, wherein the target data model is established based on the target requirements of the initial file.
[0041] Optionally, a general data model can be established based on the common requirements of data files and can also be referred to as a general model. A target data model can be established based on the specific requirements of data files and can also be referred to as a specific model. The data model layer is the third layer of the big data platform. It establishes a general model (general data model) based on the common requirements of data files to perform computational analysis on the data and uniformly complete basic data processing. The data enhancement layer is the fourth layer of the big data platform. It establishes a specific model (target data model) based on the specific requirements of data files to process the data and form data files with special requirements.
[0042] Optionally, the data model layer reads target data from the data operation layer and performs unified calculations based on the model to complete data processing of common data. The data enhancement layer reads data from the data operation layer and performs further calculations to complete specific data processing requirements and initially generate data files. The data model layer and data enhancement layer can save data processing time, reduce repetitive processing operations, and improve the speed and accuracy of file generation.
[0043] In some embodiments of the present application, the system further includes: a central scheduling platform for sending data processing instructions to the big data platform, and for sending file integration instructions to the data service platform.
[0044] Optionally, the central dispatching platform may be a unified instruction dispatching platform, which performs unified dispatching of operations such as synchronization, storage, calculation, and service calls of data between modules in the system.
[0045] In some embodiments of the present application, the big data platform is also used to partition the initial file to obtain partition files, and / or to bucket the initial file to obtain bucket files.
[0046] Optionally, after obtaining business data from the database, extracting the data through the big data platform, and integrating and processing the data to generate the initial data files, to improve data reading and storage efficiency, the data files can be partitioned and bucketed before data synchronization and storage. The data files are further integrated through file merging, classification, renaming, and compression and packaging to form the required files. Finally, the integrated data files are stored for subsequent operations.
[0047] In some embodiments of the present application, the data service platform is further used to call a script file to convert a partition file into a target file, and / or to convert a bucket file into a target file.
[0048] Optionally, after obtaining data from the database, performing preliminary processing on the data, synchronizing it to the big data platform, and the big data platform storing and processing the report and generating partition bucket files, a data service platform client can be established to call the script file to integrate, rename and compress the partition files and bucket files of the big data platform so that the partition files and bucket files meet the file requirements.
[0049] In an embodiment of the present application, the big data platform is used to receive initial data from a database and, in response to data processing instructions, convert the initial data into an initial file; the data service platform is used to respond to file integration instructions and convert the initial file into a target file according to a target format corresponding to the target file to be generated; the data file platform is used to store the target file, thereby achieving the technical effect of improving the efficiency of generating the target file, and thereby solving the technical problem of low efficiency in generating the target file.
[0050] Figure 2 is a schematic diagram of a device for generating a target file according to an embodiment of the present application, such as Figure 2 As shown, the target file generation device 200 may include: a data reading module 202 , a data analysis module 204 and a file integration module 206 .
[0051] The data reading module 202 is used to read the initial data of the database.
[0052] Optionally, the data reading module 202 can implement the function of reading external data. The data reading module can at least include a data instruction set, a parameter configuration set, and an instruction editor for implementing the writing, compilation, and configuration of data reading instructions.
[0053] The data analysis module 204 is configured to respond to the data processing instruction and convert the initial data into an initial file.
[0054] Optionally, the data analysis module 204 is used to implement functions such as calculation and analysis of reported data. The data analysis module may at least include a task allocation unit and a data processing unit to implement allocation of data calculation tasks and data processing.
[0055] Optionally, the task allocation unit can be used to manage and schedule data analysis tasks to ensure reasonable allocation and load balancing during the data processing process, and to fully utilize the system's computing resources. The task allocation unit can decompose complex data analysis tasks into a series of smaller subtasks to facilitate parallel processing. For example, the data processing task of a large table can be decomposed into multiple small tasks, each of which processes specific rows or columns in the table. The task allocation unit can allocate subtasks to the computing resources that are most suitable for processing them based on the type, data volume and complexity of the subtasks. The task allocation unit can determine the execution order and timing of data processing tasks to optimize the efficiency of the overall process. The task allocation unit can monitor the execution status and performance of tasks and dynamically adjust the allocation strategy according to actual conditions. For example, when it is found that certain tasks are executing slowly, tasks can be reallocated or computing resources can be increased.
[0056] Optionally, the data processing unit can be responsible for performing data analysis tasks, including data cleaning, conversion, calculation and analysis. The data processing unit is used to identify and process errors, anomalies or incomplete information in the data. For example, removing duplicate data, correcting erroneous values or filling missing data. The data processing unit is used to convert data from its original format to a format suitable for analysis. The data processing unit is used to perform mathematical or statistical operations on the data, such as sum, average, median or standard deviation, as well as more complex aggregation and analytical calculations. The data processing unit is used to conduct in-depth analysis of the data, extract valuable business insights or meet specific regulatory requirements. The data processing unit is used to convert the processed and analyzed data into the final output format.
[0057] The file integration module 206 is configured to respond to the file integration instruction and convert the initial file into a target file according to a target format corresponding to the target file to be generated.
[0058] Optionally, the file integration module integrates various scattered reporting files formed based on internal data to form reportable files. The file integration module may at least include a file reading unit, a file operation unit, and a file operation command set to realize the integrated generation of sent files.
[0059] In an embodiment of the present application, a data reading module is used to read the initial data of a database; a data analysis module is used to respond to data processing instructions and convert the initial data into an initial file; a file integration module is used to respond to file integration instructions and convert the initial file into a target file according to the target format corresponding to the target file to be generated, thereby achieving the technical effect of improving the efficiency of generating the target file, and thus solving the technical problem of low efficiency in generating the target file.
[0060] In some embodiments of the present application, the data reading module includes: a data instruction set for storing rules for reading initial data from a database; a parameter configuration set for controlling the speed of reading the initial data; and an instruction editor for editing reading instructions for reading the initial data.
[0061] Optionally, a data instruction set is a set of predefined instructions that instruct the module on how to read and process data from an external data source. The data instruction set may include, but is not limited to, data source connection information (such as database type, address, port, username, and password), data read scope (such as table name, field name), data filtering conditions (such as timestamp), and data processing logic (such as data conversion and data cleaning rules). The data instruction set can be used to enable the data reading module to flexibly respond to different data sources and various data requirements, while simplifying the data reading and processing process through a standardized instruction format.
[0062] Optionally, a parameter configuration set is a set of configurable parameters used by the data reading module to control data reading behavior during operation. These parameters can include specific settings such as data reading frequency, data buffer size, and error handling strategies. Parameter configuration sets allow system administrators or data engineers to adjust the data reading process based on system performance, data characteristics, and business needs to optimize data reading efficiency and reliability. Using parameter configuration sets, the data reading module can more flexibly adapt to various data environments and reading scenarios.
[0063] Optionally, a command editor is a tool or interface for creating, editing, and managing data read commands, providing a user-friendly way to write and modify specific data read commands. Command editors typically feature a visual interface, allowing users to intuitively generate complex queries or other types of data read logic through drag-and-drop methods, form filling, and other methods. Furthermore, command editors offer preview and testing features to help users verify the correctness and efficiency of commands. This is crucial for streamlining the data read command writing process and improving the efficiency of data engineers.
[0064] In some embodiments of the present application, the device further includes: a data storage module for storing at least initial data, initial files, and target files; and an instruction scheduling and analysis module for scheduling data processing instructions and file integration instructions.
[0065] Optionally, the data storage module is used to implement storage functions after external data is read and during data processing. The data storage module may include at least a data storage unit and a storage recording unit, which are used to implement data storage and storage process recording.
[0066] Optionally, the data storage unit can be used to store data in a persistent manner in the system to ensure data integrity and accessibility. The data storage unit is used to store raw data obtained from the database layer or other external data sources in the big data platform. The data storage unit is used to convert data into a storage format suitable for the big data platform when storing data. The data storage unit is used to establish data indexes and perform necessary data optimization measures during the storage process to facilitate subsequent data retrieval and processing, thereby increasing the speed of data operations. The data storage unit is used to regularly back up data and has a data recovery mechanism to ensure that data can be quickly restored in the event of a system failure.
[0067] Optionally, the storage recording unit can be used to record detailed information during the data storage process, including the data storage time, storage location, storage format, and possible data operation history. The storage recording unit is used to record the time, data volume, storage location, and parameter configuration used in the storage process for each data storage, so as to facilitate subsequent data auditing and problem troubleshooting. The storage recording unit is used to record operations performed on data, such as data cleaning, conversion, or calculation, as well as the executors of the operations and the results of the operations, which helps to understand the reasons and processes of data changes. The storage recording unit is used to record data storage efficiency information, such as read and write speeds, cache hit rates, etc., which helps to evaluate the performance of the data storage module and provide a basis for subsequent optimization. The storage recording unit is used to record data access and usage, including who accessed the data, how the data was used, and access time and frequency.
[0068] Optionally, the instruction scheduling and analysis module is used to implement functions such as generation and sending of various instructions in the system. The instruction scheduling and analysis module may at least include an instruction storage unit, an instruction logic parsing unit, an instruction monitoring unit, and an instruction transceiver unit, which are used to implement generation, parsing, sending and receiving, and execution status monitoring of instructions in the system.
[0069] In some embodiments of the present application, the instruction scheduling and analysis module includes: an instruction generation unit for generating data processing instructions and file integration instructions; an instruction sending unit for sending data processing instructions to the data analysis module and sending file integration instructions to the file integration module; an instruction parsing unit for parsing data processing instructions and file integration instructions; and an instruction monitoring unit for monitoring the execution of data processing instructions and file integration instructions.
[0070] Optionally, the instruction parsing unit can be used to parse instructions stored in the instruction storage unit, converting them into specific operations that can be executed by other parts of the system. It can also be used to analyze the logical structure of instructions, such as conditional statements, loop structures, data flow control, and possible dependencies between instructions. By parsing the instruction module, an execution sequence and strategy can be generated, allowing data processing and file generation tasks to proceed according to predetermined logic and processes.
[0071] Optionally, the instruction monitoring unit can monitor the execution status of instructions and system performance, tracking the entire lifecycle of each instruction from generation to completion, including execution time, resource consumption, and error messages, to ensure correct and efficient execution. Furthermore, the instruction monitoring unit provides real-time feedback and alerts, immediately notifying system administrators of any anomalies or delays during instruction execution, allowing for timely intervention and adjustments.
[0072] Optionally, the instruction generation unit may be used to generate a data processing instruction and a file integration instruction. The instruction sending unit may be used to send the data processing instruction to the data analysis module and send the file integration instruction to the file integration module.
[0073] In an embodiment of the present application, a data reading module is used to read the initial data of a database; a data analysis module is used to respond to data processing instructions and convert the initial data into an initial file; a file integration module is used to respond to file integration instructions and convert the initial file into a target file according to the target format corresponding to the target file to be generated, thereby achieving the technical effect of improving the efficiency of generating the target file, and thus solving the technical problem of low efficiency in generating the target file.
[0074] In order to facilitate those skilled in the art to better understand the technical solution of the present application, a specific embodiment is now described.
[0075] Currently, in industries like finance, regulatory authorities require regular reporting of business-related data to stay informed of industry trends and market conditions. With the development of diversified businesses, the amount of data generated continues to increase, placing new demands on the systems that generate the data files being reported, such as high concurrency and computational performance.
[0076] Figure 3 This is a schematic diagram of a data file generation system according to related technology, such as Figure 3 As shown, the data file generation system comprises a database layer 301, a basic service layer 302, a public layer 303, and a data file layer 304. The basic service layer 302 includes a data persistence layer 3021, a data service layer 3022, and a management service layer 3033. The database layer stores raw business data, while the data persistence layer connects to the basic database. The data persistence layer stores data scripts that conform to business logic and manages the database connection pool component of the basic service layer. The data service layer extracts data from the database layer and analyzes and processes it according to business logic. The management service layer is responsible for auxiliary functions such as collecting system operation logs and performing operational analysis. The public layer implements overall data scheduling and corresponding service delivery by transmitting access instructions and requests to the basic service layer. The data file layer stores generated data files.
[0077] However, the above system reads data from the database in batches, and the system performance and data reporting efficiency depend on the performance of the database tools used by the system. When the amount of data processed is large, the above system has risks such as insufficient performance and excessive processing time. It is difficult to effectively meet the requirements of real-time generation of massive data and high-quality data files, and thus there is a technical problem of low efficiency in generating target files.
[0078] In order to solve the above problems, this application proposes a data file generation system based on a big data platform. In business scenarios with massive data volumes, it can achieve the business requirements of efficiently, high-quality and stable generation of large-scale reporting data files. Figure 4 is a schematic diagram of a data file generation system according to an embodiment of the present application, such as Figure 4As shown, the data file generation system includes a database 401, a unified instruction scheduling platform 402, a big data platform 102, a data service platform 104, and a data file platform 106. Specifically, the data service platform 104 includes a scheduling instruction processing layer 1041, a data file service layer 1042, and a data instruction processing layer 1043. Database 401 stores raw data. The big data platform 102 is responsible for preprocessing the data, generating raw data files, and synchronizing and storing the files. The data service platform 104 integrates and processes the raw data files generated by the big data platform. The unified instruction scheduling platform 402 coordinates operations such as data synchronization, storage, calculation, and service calls between modules within the system. The data file platform 106 stores the generated data files.
[0079] In this embodiment, business data in the database is synchronized to the big data platform. The big data platform receives instructions from the unified scheduling platform and processes the target data, including filtering, merging, and performing data logic operations. It generates an initial file containing the target data, stores it, and synchronizes it to the data service platform. The data service platform receives instructions from the unified scheduling platform and, in accordance with data file format and content requirements, classifies, merges, compresses, and renames the files processed by the big data platform to form the target data file. The file is then transferred to the data file platform to complete the data file generation.
[0080] The big data platform 102 includes a data buffer layer 1021, a data operation layer 1022, a data model layer 1023, and a data enhancement layer 1024. The data buffer layer 1021 is the first layer of the big data platform, connecting to external databases and used to store data read externally. The data operation layer 1022 is the second layer of the big data platform, storing data extracted, converted, and loaded from the data buffer layer. This data is related to the target data and prepares for subsequent data processing. The data model layer 1023 is the third layer of the big data platform. At this layer, a general model is established based on the common requirements of data files to perform computational analysis on the data, and basic data processing is completed in a unified manner. The data enhancement layer 1024 is the fourth layer of the big data platform. At this layer, a specific model is established based on the specific requirements of the data files to process the data and form data files with special requirements.
[0081] The big data platform reads data from external databases and stores it in the data buffer layer, enabling the loading of external data into the big data platform. The data operation layer extracts target data from the data buffer layer and performs operations such as data association and classification, improving the speed and efficiency of subsequent target data association operations. The data model layer reads target data from the data operation layer and performs unified calculations based on the model to complete data processing for common data. The data model layer reads data from the data operation layer and performs further calculations to complete specific data processing requirements and initially generate data files. These two designs save data processing time, reduce repetitive processing operations, and improve the speed and accuracy of file generation.
[0082] Figure 5 is a schematic diagram of a data architecture of a data file generation system according to an embodiment of the present application, such as Figure 5 As shown in Figure 1, the data architecture includes data acquisition, data processing, and data file integration. Data acquisition involves obtaining the source data required for reporting. Data processing involves analyzing and processing the data to generate preliminary data files. Data file integration involves processing the data files according to the file generation requirements to generate data files that meet the requirements.
[0083] First, business data is obtained from the database. The system extracts the data through the big data platform, integrates and processes the data, and generates initial files (data submission and initial file generation). To improve data reading and storage efficiency, the data files are partitioned and bucketed before data synchronization and storage. The data files are further integrated, through file merging, classification, renaming, and packaging, to form the required data files. Finally, the integrated data files are stored for subsequent operations.
[0084] After obtaining data from the database, it undergoes preliminary processing and is then synchronized to the big data platform. The big data platform stores and processes the reports, generating partition and bucket files. Simultaneously, a data service platform client is established, which calls scripts to consolidate, rename, and compress the partition and bucket files on the big data platform to meet file requirements.
[0085] Figure 6 is a schematic diagram of a data reporting device according to an embodiment of the present application, such as Figure 6 As shown, the data reporting device 600 includes a data reading module 202 , a data storage module 601 , a data analysis module 204 , an instruction scheduling analysis module 602 , and a file integration module 206 .
[0086] The data reading module 202 implements the function of reading external data, and at least includes a data instruction set, a parameter configuration set, and an instruction editor, which is used to implement the writing, compilation, and configuration of data reading instructions.
[0087] The data storage module 601 completes the storage function after external data is read and the storage function during data processing, and at least includes a data storage unit and a storage recording unit, which are used to realize data storage and storage process recording.
[0088] The data analysis module 204 implements functions such as calculation and analysis of reported data, and includes at least a task allocation unit and a data processing unit for implementing allocation of data calculation tasks and data processing.
[0089] The instruction scheduling and analysis module 602 realizes the generation and sending of various instructions in the system, and includes at least an instruction storage unit, an instruction logic parsing unit, an instruction monitoring unit, and an instruction transceiver unit, which are used to realize the generation, parsing, sending and receiving, and execution status monitoring of instructions in the system.
[0090] The file integration module 206 integrates various scattered reporting files formed after the system internal data to form a reportable file, which at least includes a file reading unit, a file operation unit, and a file operation command set; and is used to realize the integrated generation of sending files.
[0091] In this embodiment, a unified instruction scheduling platform, big data platform, data service platform, and data file platform work in close coordination. Using a big data platform, the data architecture consists of a data buffer layer, a data operation layer, a data model layer, and a data enhancement layer. The data operation layer performs operations such as data association and classification, improving the speed and efficiency of subsequent target data association operations. The data model layer performs unified calculations based on the model, completing data processing for common data. The system's data architecture includes three parts: data acquisition, data processing, and data file integration. The data processing part extracts data from the big data platform, integrates and processes the data, and generates initial data files. To improve data reading and storage efficiency, data files are partitioned and bucketed before data synchronization and storage. The data file integration part establishes a data service platform client and calls a script file to integrate, rename, and compress the partitioned and bucketed files of the big data platform to meet file requirements. The data reporting device includes a data reading module, a data storage module, a data analysis module, an instruction scheduling analysis module, and a file integration module.
[0092] This embodiment uses a unified scheduling platform to achieve centralized task management, unified monitoring, and efficient operations and maintenance, reducing operations and maintenance costs, improving development efficiency, and effectively ensuring data accuracy and consistency. The big data platform's four-tier architecture, with layered decoupling and common unified computing, reduces operations and maintenance, facilitates rapid development, and rapidly generates data files. Leveraging the high concurrency and computing power of the big data platform, data processing processes such as data reading, calculation, and processing are transformed into file processing tasks within multi-level partitioned tables of big data. This simplifies the processing flow and file processing tasks of data processing, improves the overall file processing efficiency of the system, and enables the efficient and stable generation of large volumes of data files.
[0093] In an embodiment of the present application, the big data platform is used to receive initial data from a database and, in response to data processing instructions, convert the initial data into an initial file; the data service platform is used to respond to file integration instructions and convert the initial file into a target file according to a target format corresponding to the target file to be generated; the data file platform is used to store the target file, thereby achieving the technical effect of improving the efficiency of generating the target file, and thereby solving the technical problem of low efficiency in generating the target file.
[0094] According to an embodiment of the present application, an electronic device is provided, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any of the above-mentioned voltage value determination methods, devices, storage media and processors.
[0095] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0096] Figure 7 is a schematic block diagram of an electronic device 700 according to an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0097] like Figure 7As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0098] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0099] The computing unit 701 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above. A computer software program is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute in any other appropriate manner (e.g., by means of firmware).
[0100] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0101] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0102] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0103] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0104] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0105] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0106] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0107] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0108] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0109] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0110] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0111] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0112] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A target file generation system, characterized in that: include: The big data platform is configured to receive initial data from the database and, in response to data processing instructions, convert the initial data into an initial file; The data service platform is used to respond to the file integration instruction and convert the initial file into a target file according to the target format corresponding to the target file to be generated; The data file platform is used to store the target file.
2. The system according to claim 1, wherein: The data service platform includes: A data buffer layer, configured to extract the initial data from the database; The data operation layer is used to extract a plurality of common data from the initial data and associate the plurality of common data to obtain associated data.
3. The system according to claim 2, characterized in that The data service platform also includes: A data model layer, configured to process the associated data by calling a general data model to obtain processed associated data, wherein the general data model is established based on the general requirements of the initial file; The data enhancement layer is used to call the target data model to process the processed associated data to obtain the initial file, wherein the target data model is established based on the target requirements of the initial file.
4. The system according to claim 1, wherein: The system further comprises: The central scheduling platform is used to send the data processing instructions to the big data platform and to send the file integration instructions to the data service platform.
5. The system according to claim 1, wherein: The big data platform is further used to partition the initial file to obtain partition files, and / or to bucket the initial file to obtain bucket files.
6. The system according to claim 5, characterized in that The data service platform is further configured to call a script file to convert the partition file into the target file, and / or to convert the bucket file into the target file.
7. A device for generating a target file, characterized in that: include: Data reading module, used to read the initial data of the database; A data analysis module, configured to respond to a data processing instruction and convert the initial data into an initial file; The file integration module is used to respond to the file integration instruction and convert the initial file into a target file according to the target format corresponding to the target file to be generated.
8. The device according to claim 7, characterized in that The data reading module includes: a data instruction set, configured to store rules for reading the initial data from the database; A parameter configuration set for controlling a speed of reading the initial data; An instruction editor is used to edit a read instruction for reading the initial data.
9. The device according to claim 7, characterized in that The device further comprises: A data storage module, configured to store at least the initial data, the initial file, and the target file; The instruction scheduling analysis module is used to schedule the data processing instructions and the file integration instructions.
10. The device according to claim 9, characterized in that The instruction scheduling analysis module includes: An instruction generating unit, configured to generate the data processing instruction and the file integration instruction; an instruction sending unit, configured to send the data processing instruction to the data analysis module and send the file integration instruction to the file integration module; An instruction parsing unit, configured to parse the data processing instruction and the file integration instruction; The instruction monitoring unit is used to monitor the execution status of the data processing instruction and the execution status of the file integration instruction.