Multi-source data storage method, system, device and computer readable storage medium

By pre-configuring data processing strategies, the problem of low storage efficiency caused by inconsistent data formats from multiple sources is solved, achieving unified data format and fast and accurate storage.

CN119668498BActive Publication Date: 2026-04-07DONGFENG COMML VEHICLE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing technologies, inconsistent data formats from multiple sources lead to low storage efficiency, and big data platforms struggle to store data from various data sources quickly and accurately.

Method used

Multiple data processing strategies are pre-configured to convert source files into target files through data extraction, association, and format conversion, including the standardization and custom filling of data item names and contents.

Benefits of technology

It improves the efficiency of multi-source data storage, ensures data format consistency, and enables fast and accurate data storage to the big data platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119668498B_ABST
    Figure CN119668498B_ABST
Patent Text Reader

Abstract

A multi-source data storage method, system, device, and computer-readable storage medium, belonging to the field of big data storage, includes pre-configured multiple data processing strategies. These strategies are used to extract data from at least one source file with a source format, perform data association and data format conversion to obtain a target file with a target format. Data extraction includes extracting at least one data item and its data content. Data association includes associating related data items and their data content based on identical data items. Data format conversion includes unifying the names of data items, unifying the data units of data content, and / or customizing the data content during data association. After receiving source files from multiple data sources, the data is processed according to the data processing strategies before storage. This application improves data storage efficiency by configuring data processing strategies to unify the format of multi-source data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data storage, specifically to a multi-source data storage method, system, device, and computer-readable storage medium. Background Technology

[0002] Current big data storage faces challenges due to the increasing variety of data sources and their diverse data formats. This makes it difficult to standardize data formats and ensure that all data is stored on the big data platform quickly and accurately. Furthermore, for some data sources, reference data needs to be stored alongside the data on the big data platform during storage.

[0003] Currently, before storing multi-source data on a big data platform, it generally relies on manual methods to convert the data format to a specified format. If the data volume is small, manual methods can be used to adjust the data format of each data source and extract useful information for storage. However, problems will arise once the data volume becomes large. Summary of the Invention

[0004] This application provides a multi-source data storage method, system, device, and computer-readable storage medium, which can solve the technical problem of low storage efficiency caused by different multi-source data formats in the prior art.

[0005] In a first aspect, embodiments of this application provide a multi-source data storage method, the method comprising:

[0006] Multiple data processing strategies are pre-configured. These strategies are used to extract data from at least one source file with a source format, perform data association and data format conversion, and obtain a target file with a target format.

[0007] Both the source file and the target file include multiple data items and the data content corresponding to each data item; the data extraction includes extracting at least one data item and its data content; the data association includes associating data items and their data content together based on the same data item; the data format conversion includes unifying the names of data items, unifying the data units of data content, and / or customizing the data content when performing data association;

[0008] After receiving source files from multiple data sources, each source file is processed according to the preset data processing strategy to obtain the target file and store it.

[0009] In conjunction with the first aspect, in one implementation, the source format includes xlsx format, json format, csv format, txt format, and xlsx format.

[0010] In conjunction with the first aspect, in one embodiment, the method includes:

[0011] The data processing strategy is formed by providing a configuration template for users to fill in content.

[0012] In conjunction with the first aspect, in one embodiment, the method includes:

[0013] After receiving source files from multiple data sources and removing interfering data from each source file, each source file is processed according to the preset data processing strategy to obtain the target file and store it.

[0014] In conjunction with the first aspect, in one implementation, the data items are categorized into data source identity data items and data source operation data items.

[0015] In conjunction with the first aspect, in one implementation, the data association includes associating the data content of a data source identity data item with the data content of a data source operation data item associated with that data source identity data item, based on the same data source identity data item; and

[0016] The data association includes running data items based on the same data source, and associating the data content of the data items running from the same data source with other data items running from the same data source and their data content.

[0017] In conjunction with the first aspect, in one implementation, the data extraction includes extracting at least one data item and all its data content; and

[0018] The data extraction includes extracting at least one data item and a portion of its data content.

[0019] Secondly, embodiments of this application provide a multi-source data storage system, the system comprising:

[0020] A configuration module is used to pre-configure multiple data processing strategies. These strategies extract data from at least one source file with a source format, then perform data association and data format conversion to obtain a target file with a target format. Both the source file and the target file include multiple data items and corresponding data content for each data item. Data extraction includes extracting at least one data item and its data content. Data association includes associating data items and their data content based on the same data item. Data format conversion includes unifying the names of data items, unifying the data units of the data content, and / or customizing the data content during data association.

[0021] The data processing module is used to receive source files from multiple data sources, process each source file according to the preset data processing strategy, and obtain and store the target file.

[0022] Thirdly, embodiments of this application provide a multi-source data storage device, which includes a processor, a memory, and a multi-source data storage program stored in the memory and executable by the processor, wherein when the multi-source data storage program is executed by the processor, it implements the steps of the multi-source data storage method.

[0023] Fourthly, embodiments of this application provide a computer-readable storage medium storing a multi-source data storage program, wherein when the multi-source data storage program is executed by a processor, it implements the steps of the multi-source data storage method.

[0024] The beneficial effects of the technical solutions provided in this application include:

[0025] By pre-configuring multiple data processing strategies, data is extracted from at least one source file with a source format, followed by data association and data format conversion to obtain a target file with a target format. Data extraction includes extracting at least one data item and its data content. Data association includes associating data items and their data content based on the same data item. Data format conversion includes unifying the names of data items, unifying the data units of data content, and / or customizing the data content during data association. This solves the technical problem of low storage efficiency caused by different data formats from multiple sources in related technologies. Attached Figure Description

[0026] Figure 1 This is a flowchart illustrating an embodiment of the multi-source data storage method of this application;

[0027] Figure 2 This is a flowchart illustrating a specific embodiment of the multi-source data storage method of this application;

[0028] Figure 3 This is a functional module diagram of an embodiment of the multi-source data storage system of this application;

[0029] Figure 4 This is a schematic diagram of the hardware structure of the multi-source data storage device involved in the embodiments of this application. Detailed Implementation

[0030] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0031] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0032] In a first aspect, embodiments of this application provide a multi-source data storage method.

[0033] In one embodiment, reference is made to Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the multi-source data storage method of this application. Figure 1 As shown, multi-source data storage methods include:

[0034] Step S1: Pre-configure multiple data processing strategies. These strategies are used to extract data from at least one source file with a source format, perform data association and data format conversion, to obtain a target file with a target format. Both the source and target files include multiple data items and the corresponding data content for each data item. Data extraction includes extracting at least one data item and its data content. Data association includes associating data items and their data content together based on the same data item. Data format conversion includes unifying the names of data items, unifying the data units of the data content, and / or customizing the data content during data association.

[0035] Step S2: After receiving source files from multiple data sources, process each source file according to a preset data processing strategy to obtain the target file and store it.

[0036] In this embodiment, multiple data processing strategies are pre-configured to extract data from at least one source file with a source format, perform data association and data format conversion to obtain a target file with a target format. Data extraction includes extracting at least one data item and its data content. Data association includes associating data items and their data content together based on the same data item. Data format conversion includes unifying the names of data items, unifying the data units of data content, and / or customizing the data content during data association. This solves the technical problem of low storage efficiency caused by different data formats from multiple sources in related technologies.

[0037] Furthermore, in one embodiment, the source format includes xlsx format, json format, csv format, txt format, and xlsx format.

[0038] In this embodiment, in addition to the formats listed above, other file formats are also applicable as long as they can satisfy the requirement that data items and their data content can be extracted from them and that the data items can be associated with data items in other source files.

[0039] Furthermore, in one embodiment, the above method includes:

[0040] By providing configuration templates for users to fill in content, the aforementioned data processing strategy can be formed.

[0041] In this embodiment, a configuration template is provided. The configuration template can be in tabular form or other forms. Auxiliary information is pre-filled in, including prompts for data extraction operations from the source file, as well as prompts for data association and data format conversion operations.

[0042] In the first embodiment, taking the first source file as a JSON file and the second source file as an XLSX file as an example, the pre-configured data processing strategies for these two types of source files are shown in Table 1 below. Table 1 is a configuration template for the JSON-XLSX data processing strategy.

[0043] Table 1 shows the configuration template for the json-xlsx data processing strategy.

[0044]

[0045] The first column of Table 1 contains auxiliary information, while the second and third columns contain the specific details of the data processing strategy.

[0046] According to Table 1, the json-xlsx data processing strategies include:

[0047] The data extraction operation involves extracting all data items and their complete content from a JSON file, extracting operational data from an XLSX file, extracting the SIM card number, chassis number, and mileage (in km) data items and their content from the operational data, extracting data items from the first row of the header row of the XLSX file, and extracting the data content of each data item starting from the sixth row of the data rows in the XLSX file. Data extraction from the JSON file is performed using UTF-8 encoding, and data extraction from the XLSX file is performed using GBK encoding.

[0048] The data association operation uses sim as the association field to associate the data content of sim in the JSON file, as well as the data items and their data content of the data source associated with sim, with the data content of sim in the XLSX file, as well as the data items and their data content of the data source associated with sim. One sim and all its related data constitute an output field, and all output fields constitute the output file, namely result_20240801.csv. Then, result_20240801.csv is stored in the big data platform.

[0049] The data format conversion operation, that is, when extracting data from an xlsx file, if certain speeds lack specific data content, then empty values ​​will be filled with 0.

[0050] Specifically, the JSON file is as follows:

[0051] {"sim":["21344","21555"],

[0052] "21344":{"Chassis Number":"E2345222",

[0053] "vin":"1111111121344",

[0054] "Market of Origin": "Long-distance transportation"

[0055] "21555":{"Chassis Number":"D2343243",

[0056] "vin":"2222222221555",

[0057] "Market of Origin": "Medium and Long-Distance Transportation"

[0058] }

[0059] The xlsx file is as follows:

[0060]

[0061] The target files after processing the JSON and XLSX files according to the data processing strategy are as follows:

[0062]

[0063]

[0064] Furthermore, in one embodiment, the above method includes:

[0065] After receiving source files from multiple data sources and removing interfering data from each source file, each source file is processed according to the preset data processing strategy to obtain the target file and store it.

[0066] In this embodiment, there is a piece of useless interference data in the source file output from the data source. This data does not need to be stored in the big data platform. Therefore, by pre-extracting the interference data, the effective data can be extracted. Taking the xlsx file in the first embodiment as an example, its second to fifth lines are interference data, which can be directly removed. By removing interference data, the efficiency and accuracy of extracting data items and data content in the later stage can be improved.

[0067] Furthermore, in one embodiment, the aforementioned data items are categorized into data source identity data items and data source operation data items.

[0068] In this embodiment, the source file generated by the data source contains static information and dynamic information. The static information is the data source identity data item, which is used to characterize the identity information of the data source (for example, SIM and chassis number are two different levels of data source identity data items. One chassis number can correspond to multiple SIMs. When associating data later, it can be associated based on SIM or based on a combination of SIM and chassis number). The dynamic information is the data source operation data item, which is used to ensure the equipment operation data collected by the data source during the operation of the data source or during the operation of the device in which the data source is located.

[0069] When processing source files of different formats generated from multiple data sources according to the data processing strategy, data association can be performed based on the identity data items of the same data source in each source file, or data association can be performed based on the running data items of the same data source in each source file.

[0070] In other words, the aforementioned data association includes associating the data content of the identity data item from the same data source with the data content of the associated data source's operational data item.

[0071] The aforementioned data association includes running data items based on the same data source, associating the data content of the data item running from the same data source with other data items running from the same data source and their data content.

[0072] Furthermore, in one embodiment, the above data extraction includes extracting at least one data item and its entire data content.

[0073] The above data extraction includes extracting at least one data item and a portion of its data content.

[0074] In this embodiment, when extracting data, all data content under a certain data item in one source file can be associated with all data content under a certain data item in another source file, or a portion of data content under a certain data item in one source file can be associated with a portion of data content under a certain data item in another source file, or all data content under a certain data item in one source file can be associated with a portion of data content under a certain data item in another source file.

[0075] In the second embodiment, taking the first source file as a txt file and the second source file as a csv file as an example, the pre-configured data processing strategies for these two types of source files are shown in Table 2 below. Table 2 is the configuration template for the txt-csv data processing strategy.

[0076] Table 2. Configuration template for txt-csv data processing strategy

[0077]

[0078]

[0079] The first column of Table 2 contains auxiliary information, while the second and third columns contain the specific details of the data processing strategy.

[0080] Based on the information in Table 2, the txt-csv data processing strategies include:

[0081] The data extraction operation involves extracting all data items and their complete content from a TXT file, and extracting the SIM, chassis number, VIN, and market number, along with their content, from a CSV file. Specifically, it extracts data items from the first row of the header row of the CSV file, and then extracts the content of each data item starting from the second row of the data rows. Data extraction from the TXT file uses UTF-8 encoding, and data extraction from the CSV file uses GBK encoding.

[0082] The data association operation uses sim as the association field to associate the data content of sim in the txt file, as well as the data items and their data content of the data source associated with sim, with the data content of sim in the csv file, as well as the data items and their data content of the data source associated with sim. One sim and all its related data constitute an output field, and all output fields constitute the output file, namely result_20240801.csv. Then, result_20240801.csv is stored in the big data platform.

[0083] The data format conversion operation involves extracting all GPS times associated with the same SIM from the CSV file, averaging them, and using the average value as the data content of the GPS times associated with that SIM in the target file.

[0084] Specifically, the txt file is as follows:

[0085] sim, gpstime, height, speed, mile

[0086] 21344,2024-01-01 01:12:12,120,20,2000

[0087] 21344, 2024-01-01 01:12:12, 120, 30, 2000

[0088] 21344,2024-01-01 01:12:12,120,28,2000

[0089] 21555, 2024-01-01 01:12:12,12,30,1130

[0090] 21555, 2024-01-01 01:12:12, 122, 30, 1200

[0091] 21555,2024-01-01 01:12:12,13,30,1150.

[0092] The CSV file is as follows:

[0093]

[0094] The target files after processing the txt and csv files according to the data processing strategy are as follows:

[0095]

[0096] In the third embodiment, taking the first source file as an xlsx file, the second source file as an xlsx file, the third source file as an xlsx file, and the fourth source file as an xlsx file as examples, the pre-configured data processing strategies for these four types of source files are shown in Table 3 below. Table 3 is the configuration template for the xlsx-xlsx-xlsx-xlsx data processing strategy.

[0097] Table 3 shows the configuration template for the xlsx-xlsx-xlsx-xlsx data processing strategy.

[0098]

[0099] The first column of Table 3 contains auxiliary information, while the second to fifth columns contain the specific details of the data processing strategy.

[0100] According to Table 3, the xlsx-xlsx-xlsx-xlsx data processing strategies include:

[0101] The data extraction operation involves retrieving the vehicle speed data (from line 51 onwards) from Data1.xlsx (the first source file, Sheet1); the speed data (from line 50 onwards) from Data2.xlsx (the second source file, Sheet1); the speed data (from line 52 onwards) from Data3.xlsx (the third source file, Sheet1); and the GPS speed data (from line 53 onwards) from Data4.xlsx (the fourth source file, Sheet1). Vehicle speed data, speed data, speed data, and GPS speed data are essentially the same type of data. However, due to different operators or other reasons, the data source assigns different names to this type of data when outputting it, thus categorizing it under different data items.

[0102] The data association operation uses speed as the association field to associate the data content of the data items with the same speed in the four xlsx files. A speed and all its related data constitute an output field, and all output fields constitute the output file, namely result_20240808.xlsx. Then, result_20240808.xlsx is stored in the big data platform.

[0103] The data format conversion operation involves renaming all the acquired data columns to speed when extracting data from an XLSX file, and then concatenating them to form a single output field.

[0104] Furthermore, process log information is recorded, and the file storage path is displayed upon successful execution.

[0105] If the storage of a target file to the big data platform fails, check the log file for error messages. Modify the configuration template and target file according to the error messages and then store the file again.

[0106] Specifically, the data file 1.xlsx is as follows:

[0107] line number … Speed 50 … … 51 0.02 85 52 0.04 84

[0108] The data file 2.xlsx is as follows:

[0109] line number … Speed 49 … … 50 0.03 83 51 0.06 84

[0110] The data file 3.xlsx is as follows:

[0111]

[0112]

[0113] The data file 4.xlsx is as follows:

[0114] line number … Speed 52 … … 53 0.01 85 54 0.02 83

[0115] The target files after processing the four xlsx files according to the data processing strategy are as follows:

[0116] Speed 85 84 83 84 85 84 85 83

[0117] In the fourth embodiment, taking the first source file as an xlsx file, the second source file as a csv file, and the third source file as an xls file as examples, the pre-configured data processing strategies for these three types of source files are shown in Table 4 below. Table 4 is the configuration template for the xlsx-csv-xls data processing strategy.

[0118] Table 4. Configuration template for xlsx-csv-xls data processing strategies

[0119]

[0120]

[0121] The first column of Table 4 contains auxiliary information, while the second to fourth columns contain the specific details of the data processing strategy.

[0122] According to Table 4, the xlsx-csv-xls data processing strategies include:

[0123] The data extraction operation involves retrieving the SIM, DPH, speed, and mile data from line 2 onwards in the first source file, Sheet1 (data file 1.xlsx). It also retrieves the SIM card number, chassis number, and mileage (in km) data from line 6 onwards in the CSV file, and the SIM, DPH, speed (in km / h), and Mile (in km) data from line 2 onwards in the XLS file. The data extraction method for the XLSx file is UTF-8 encoding, the data extraction method for the CSV file is GBK encoding, and the data extraction method for the XLS file is UTF-8 encoding.

[0124] The data association operation involves concatenating all the acquired data. One output field contains sim, dph, speed, and mile. All output fields constitute the output file, namely result_20240809.csv, which is then stored in the big data platform.

[0125] Data format conversion operations, i.e., when extracting data from a CSV file, fill empty values ​​with 0 if specific data content is missing for certain speeds. When extracting data from an XLS file, convert the unit of Mile from km to m.

[0126] Specifically, the xlsx file is as follows:

[0127]

[0128] The CSV file is as follows:

[0129]

[0130] The xls file is as follows:

[0131]

[0132] The target files after processing the xlsx, csv, and xls files according to the data processing strategy are as follows:

[0133]

[0134] In the fifth embodiment, taking the first source file as an xlsx file, the second source file as an xlsx file, and the third source file as an xlsx file as examples, the pre-configured data processing strategies for these three types of source files are shown in Table 5 below. Table 5 is the configuration template for the xlsx-xlsx-xlsx data processing strategy.

[0135] Table 5 Configuration Template for xlsx-xlsx-xlsx Data Processing Strategy

[0136]

[0137] The first column of Table 5 contains auxiliary information, while the second to fourth columns contain the specific details of the data processing strategy.

[0138] According to Table 5, the xlsx-xlsx-xlsx data processing strategies include:

[0139] The data extraction operation involves retrieving the gpstime, dph, and speed (in km / h) data from Sheet2 in the first source file (data1.xlsx), the gpstime, dph, and Mile (in km) data from Sheet3 in the second source file (data2.xlsx), and the gpstime, dph, and Height (in m) data from Sheet4 in the third source file (data3.xlsx).

[0140] The data association operation involves associating the acquired data using gpstime and dph, then deleting redundant gpstime and dph columns (deleting common columns), and finally renaming all columns to sim, dph, speed (in km / h), mile (in m), and height (in m). Each output field contains sim, dph, speed (in km / h), mile (in m), and height (in m). All output fields constitute the output file, namely result_20240810.csv, which is then stored in the big data platform.

[0141] Specifically, the data file 1.xlsx is as follows:

[0142]

[0143] The data file 2.xlsx is as follows:

[0144]

[0145] The data file 3.xlsx is as follows:

[0146]

[0147]

[0148] The target files after processing the three xlsx files according to the data processing strategy are as follows:

[0149]

[0150] In summary, this invention primarily addresses the challenge of big data platforms accommodating diverse and irregularly stored files for data integration. Through data cleaning (removing interfering data, converting column names from English to Chinese, and concatenating data), it achieves the standards required for data integration, reducing the workload of data processing personnel. It allows for template-based configuration of multiple files with non-standard data storage, integrating the data through data cleaning, association, filling, and dimensionality reduction. It allows for one-time configuration and multiple uses. Users can specify sheets, starting rows and columns, and data to extract, and can configure multiple integration templates each time, customizing the output results.

[0151] Reference Figure 2As shown, various types of data files that need to be cleaned are collected. The integrated data files are uniformly encoded, the information of the result files (i.e., the target files) is defined, the data source file encoding required for generating the result files is specified, the rules for data extraction are defined, the required and output field information is specified, and the data processing methods are specified, forming a configuration template file. The configuration template file information is parsed, and based on the configuration information of each result file, data is extracted from the specified data source files according to the specified pagination, header node, tail node, and field information. The extracted data is then filtered, filled, correlated, transformed, sorted, and merged according to the specified processing methods, generating the result files according to the defined result file names. The generated result files are then stored in the big data platform either manually imported locally or via scripts.

[0152] Secondly, embodiments of this application also provide a multi-source data storage system.

[0153] In one embodiment, reference is made to Figure 3 , Figure 3 This is a functional module diagram of an embodiment of the multi-source data storage system of this application. Figure 3 As shown, the multi-source data storage system includes:

[0154] Configuration module 1 is used to pre-configure multiple data processing strategies. These strategies extract data from at least one source file with a source format, then perform data association and data format conversion to obtain a target file with a target format. Both the source file and the target file include multiple data items and corresponding data content for each data item. Data extraction includes extracting at least one data item and its data content. Data association includes associating data items and their data content based on the same data item. Data format conversion includes unifying the names of data items, unifying the data units of the data content, and / or customizing the data content during data association.

[0155] The data processing module 2 is used to receive source files from multiple data sources, process each source file according to the preset data processing strategy, and obtain and store the target file.

[0156] The functions of each module in the above-mentioned multi-source data storage system correspond to the steps in the above-mentioned multi-source data storage method embodiments, and their functions and implementation processes will not be described in detail here.

[0157] Thirdly, embodiments of this application provide a multi-source data storage device, which can be a personal computer (PC), laptop computer, server, or other device with data processing capabilities.

[0158] Reference Figure 4 , Figure 4 This is a schematic diagram of the hardware structure of a multi-source data storage device involved in an embodiment of this application. In this embodiment, the multi-source data storage device may include a processor, a memory, a communication interface, and a communication bus.

[0159] The communication bus can be of any type and is used to interconnect the processor, memory, and communication interface.

[0160] Communication interfaces include input / output (I / O) interfaces, physical interfaces, and logical interfaces used for interconnecting devices within a multi-source data storage device, as well as interfaces used for interconnecting the multi-source data storage device with other devices (such as other computing devices or user equipment). Physical interfaces can be Ethernet interfaces, fiber optic interfaces, ATM interfaces, etc. User equipment can be displays, keyboards, etc.

[0161] Memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.

[0162] The processor can be a general-purpose processor, which can call a multi-source data storage program stored in memory and execute the multi-source data storage method provided in the embodiments of this application. For example, the general-purpose processor can be a central processing unit (CPU). The method executed when the multi-source data storage program is called can be referred to in the various embodiments of the multi-source data storage method of this application, and will not be repeated here.

[0163] Those skilled in the art will understand that Figure 4 The hardware structure shown does not constitute a limitation of this application and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0164] Fourthly, embodiments of this application also provide a computer-readable storage medium.

[0165] This application stores a multi-source data storage program on a computer-readable storage medium, wherein when the multi-source data storage program is executed by a processor, it implements the steps of the multi-source data storage method described above.

[0166] The method implemented when the multi-source data storage program is executed can be referred to in various embodiments of the multi-source data storage method of this application, and will not be repeated here.

[0167] It should be noted that the sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0168] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus. The terms "first," "second," and "third," etc., are used to distinguish different objects, etc., and do not indicate a sequence, nor do they limit "first," "second," and "third" to different types.

[0169] In the description of the embodiments of this application, terms such as "exemplary," "for example," or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary," "for example," or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary," "for example," or "for instance" is intended to present the relevant concepts in a concrete manner.

[0170] In the description of the embodiments of this application, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more.

[0171] In some processes described in the embodiments of this application, multiple operations or steps are included in a specific order. However, it should be understood that these operations or steps may not be executed in the order they appear in the embodiments of this application, or they may be executed in parallel. The sequence number of the operation is only used to distinguish different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed sequentially or in parallel, and these operations or steps may be combined.

[0172] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device to execute the methods described in the various embodiments of this application.

[0173] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A multi-source data storage method, characterized in that, The method includes: Multiple data processing strategies are pre-configured. These strategies are used to extract data from at least one source file with a source format, perform data association and data format conversion, and obtain a target file with a target format. Both the source file and the target file include multiple data items and the data content corresponding to each data item; the data extraction includes extracting at least one data item and its data content; the data association includes associating data items and their data content together based on the same data item; the data format conversion includes unifying the names of data items, unifying the data units of data content, and / or customizing the data content when performing data association; After receiving source files from multiple data sources, each source file is processed according to the preset data processing strategy to obtain the target file and store it. The data items are categorized into data source identity data items and data source operation data items. The data association includes associating the data content of the same data source identity data item with the data content of the associated data source runtime data item; and the associated data source identity data item and its data content. The data association includes running data items based on the same data source, and associating the data content of the data items running from the same data source with other data items running from the same data source and their data content.

2. The multi-source data storage method as described in claim 1, characterized in that, The source formats include xlsx, json, csv, txt, and xlsx.

3. The multi-source data storage method as described in claim 1, characterized in that, The method includes: The data processing strategy is formed by providing a configuration template for users to fill in content.

4. The multi-source data storage method as described in claim 1, characterized in that, The method includes: After receiving source files from multiple data sources and removing interfering data from each source file, each source file is processed according to the preset data processing strategy to obtain the target file and store it.

5. The multi-source data storage method as described in claim 1, characterized in that, The data extraction includes extracting at least one data item and its entire data content; and The data extraction includes extracting at least one data item and a portion of its data content.

6. A multi-source data storage system, characterized in that, The system includes: A configuration module is used to pre-configure multiple data processing strategies. These strategies extract data from at least one source file with a source format, then perform data association and data format conversion to obtain a target file with a target format. Both the source file and the target file include multiple data items and corresponding data content for each data item. Data extraction includes extracting at least one data item and its data content. Data association includes associating data items and their data content based on the same data item. Data format conversion includes unifying the names of data items, unifying the data units of the data content, and / or customizing the data content during data association. The data processing module is used to receive source files from multiple data sources, process each source file according to the preset data processing strategy, and obtain and store the target file. The data items are categorized into data source identity data items and data source operation data items. The data association includes associating the data content of the same data source identity data item with the data content of the associated data source runtime data item; and the associated data source identity data item and its data content. The data association includes running data items based on the same data source, and associating the data content of the data items running from the same data source with other data items running from the same data source and their data content.

7. A multi-source data storage device, characterized in that, The multi-source data storage device includes a processor, a memory, and a multi-source data storage program stored in the memory and executable by the processor, wherein when the multi-source data storage program is executed by the processor, it implements the steps of the multi-source data storage method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a multi-source data storage program, wherein when the multi-source data storage program is executed by a processor, it implements the steps of the multi-source data storage method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-data-source data report processing method and system

    CN117034884A