Multi-source data processing method and system

By using a visual configuration interface and data template in multi-source data processing and determining the reading method in combination with data classification, the problem of complex and inefficient multi-source data processing in the prior art is solved, and efficient and accurate data fusion processing is achieved.

CN120145149APending Publication Date: 2025-06-13HEBEI WANYUE INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510250030.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing multi-source data processing process is complex and inefficient, which is particularly difficult for non-professional personnel.

Method used

The data template is constructed through the visual configuration interface, the data reading method is determined based on the classification of multi-source data, and the data is read and fused from multi-source data in combination with the data template.

Benefits of technology

It improves the efficiency and accuracy of data processing, lowers the professional threshold for data processing, and makes it easy for non-professional personnel to operate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145149A_ABST
    Figure CN120145149A_ABST
Patent Text Reader

Abstract

The invention provides a multi-source data processing method and system, and belongs to the technical field of data development, and the method comprises the following steps: constructing a data template based on configuration information input by a user through a configuration interface; the data template comprises a plurality of fields; determining a data reading mode based on the classification corresponding to the multi-source data; based on the data reading mode and the data template, reading first data corresponding to each field from the multi-source data; and filling the corresponding field with the first data to realize fusion processing of the multi-source data. According to the multi-source data processing method and system provided by the invention, the data processing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure belongs to the technical field of software development, and more specifically, relates to a multi-source data processing method and system. Background Art

[0002] In today's digital age, the amount of data has grown explosively. Enterprises and organizations need to process data from various different data sources, including databases, file systems, network interfaces, etc. Integrating data from different channels, such as in the business field, integrating an enterprise's sales data, customer data, and market research data, and in the medical field, integrating patient medical record data, examination data, and genetic data, can enable decision-makers to obtain more comprehensive and richer information and avoid decision-making mistakes caused by the limitations of a single data source.

[0003] The existing data development process based on multi-source data involves complex SQL statement writing, data format conversion, and proficient use of various data processing tools. The data processing efficiency is low, and it is very difficult for non-professionals. Summary of the Invention

[0004] The purpose of this disclosure is to provide a multi-source data processing method and system to improve data processing efficiency.

[0005] In the first aspect of the embodiments of this disclosure, a multi-source data processing method is provided, including: Constructing a data template based on the configuration information input by the user through a configuration interface; the data template includes multiple fields; Determining a data reading method based on the classification corresponding to the multi-source data; Based on the data reading method and the data template, reading the first data corresponding to each field from the multi-source data; Filling the first data into the corresponding fields to implement the fusion processing of the multi-source data.

[0006] In the second aspect of the embodiments of this disclosure, a multi-source data processing system is provided, including: A configuration information input module for constructing a data template based on the configuration information input by the user through a configuration interface; the data template includes multiple fields; A data classification module for determining a data reading method based on the classification corresponding to the multi-source data; A data reading module for reading the first data corresponding to each field from the multi-source data based on the data reading method and the data template; A data fusion module for filling the first data into the corresponding fields to implement the fusion processing of the multi-source data.

[0007] In a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, where when the processor executes the computer program, the steps of the above multi-source data processing method are implemented.

[0008] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program, where when the computer program is executed by a processor, the steps of the above multi-source data processing method are implemented.

[0009] The beneficial effects of the multi-source data processing method and system provided by the embodiments of the present disclosure are as follows: In the embodiments of the present disclosure, through a visual configuration interface, a user can input configuration information according to their own business needs to construct a corresponding data template. At the same time, considering that different types of multi-source data may have different formats, storage methods, and characteristics, by classifying the multi-source data and determining a corresponding data reading method for each type of data, combining the determined data reading method and the data template, the first data that meets the requirements of each field in the data template can be accurately extracted from the multi-source data; finally, the first data is filled into the corresponding fields of the data template to achieve the fusion of multi-source data.

[0010] Therefore, in this embodiment, a corresponding data reading method is selected according to the classification of multi-source data. At the same time, the use of the data template ensures the accuracy of data extraction. The appropriate data reading method and the accurate data extraction process avoid unnecessary data processing operations and improve the efficiency of data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0012] Figure 1 It is a flowchart of a multi-source data processing method provided by an embodiment of the present disclosure; Figure 2 It is a structural block diagram of a multi-source data processing system provided by an embodiment of the present disclosure; Figure 3 It is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0014] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments in conjunction with the accompanying drawings.

[0015] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a multi-source data processing method provided for an embodiment of the present disclosure. The method includes: S101: Construct a data template based on the configuration information input by the user through the configuration interface; the data template includes multiple fields.

[0016] In this embodiment, the configuration information may include the use of the data, data format requirements, business rules, etc. The data template includes multiple fields, and these fields define the structure and attributes of the data to be processed subsequently, similar to creating a "framework" for the data, providing a unified standard and structure for subsequent data fusion.

[0017] For example, in e-commerce data analysis, a merchant can construct a data template including fields such as product name, sales volume, and positive review rate according to different dimensions such as sales data and user evaluation data, accurately focusing on key business indicators.

[0018] The user can input configuration information on the configuration interface to construct a data template according to their own specific business needs and data processing objectives. This customized way of constructing a data template can meet diverse needs. At the same time, the visual configuration method reduces the professional threshold of data processing.

[0019] S102: Determine the data reading method based on the classification corresponding to the multi-source data.

[0020] In this embodiment, the multi-source data can be loaded by dragging. Considering that different types of data may have different formats, storage methods, and characteristics, by classifying the multi-source data and determining a suitable data reading method for each type of data, targeted data processing can be achieved.

[0021] Specifically, multi-source data can be divided into structured data, semi-structured data, and unstructured data according to the metadata of each source data. Among them, for structured data, strict data models and schemas are usually clearly defined in its metadata, including table structures, field names, data types, primary keys, foreign keys, etc. There are clear relationship definitions between data elements. For example, in a relational database, the metadata will describe in detail the schema of each table. For example, the student table has fields such as student ID, name, and age, and the student ID is the primary key. These information indicate that the data is structured.

[0022] For semi-structured data, its metadata usually contains some self-describing information, but there is no complete and strict schema definition like structured data. For example, in an XML document, its metadata will describe the meanings of some tags and the general hierarchical structure, but there are no strict regulations on the data content within the tags and the order of appearance, etc., with a certain degree of flexibility.

[0023] For unstructured data, there are basically no predefined data models and schemas in its metadata, and it is difficult to obtain clear information about the data structure from it. For example, the metadata of a text may only contain some basic attributes, such as title, author, abstract, etc., and cannot reflect the internal structure and organization method of the data.

[0024] Based on the classification of multi-source data obtained, for structured data (such as data in a relational database), SQL query statements can be used for reading; for semi-structured data (such as XML and JSON data), corresponding parsers can be used for reading; for unstructured data (such as document files, social media content, etc.), text processing technologies need to be used for reading.

[0025] S103: Read the first data corresponding to each field from the multi-source data based on the data reading method and the data template.

[0026] In this embodiment, based on a suitable data reading method, correct access and acquisition of data can be ensured, and based on the data template, it can be determined which specific data items need to be extracted. For example, fields such as "product name", "price", and "sales volume" are defined in the data template. According to the reading methods of different data sources, the first data corresponding to these fields are accurately extracted from multiple data sources. Therefore, by combining the determined data reading method and the data template, the data that meets the requirements of each field in the data template, that is, the first data, can be accurately extracted from the multi-source data.

[0027] S104: Fill the first data into the corresponding fields to achieve the fusion processing of multi-source data.

[0028] In this embodiment, the first data read from multi-source data is filled into the fields corresponding to the data template. Through the unified specification of the data template, multi-source data in different formats is integrated into a unified framework, thus achieving the fusion of multi-source data.

[0029] As can be seen from the above, in this embodiment, through the visual configuration interface, users can input configuration information according to their own business needs to construct the corresponding data template. At the same time, considering that multi-source data of different types may have different formats, storage methods and characteristics, by classifying the multi-source data and determining the corresponding data reading method for each type of data, combined with the determined data reading method and the data template, the first data that meets the requirements of each field in the data template can be accurately extracted from the multi-source data; finally, the first data is filled into the fields corresponding to the data template to achieve the fusion of multi-source data.

[0030] Therefore, in this embodiment, the corresponding data reading method is selected according to the classification of multi-source data. At the same time, the use of the data template ensures the accuracy of data extraction. The appropriate data reading method and the accurate data extraction process avoid unnecessary data processing operations and improve the efficiency of data processing.

[0031] In an embodiment of the present disclosure, the classification corresponding to the multi-source data includes structured data. Determining the data reading method based on the classification corresponding to the multi-source data includes: If the first source data has a corresponding API, the data reading method is determined as the first method; the first method is the method of reading data based on the API; the first source data is any one of the structured data in the multi-source data; If the first source data does not have a corresponding API, the data reading method is determined based on the data volume of the first source data.

[0032] In this embodiment, for structured data with an API, data reading can be achieved through the API interface. Through the API, users can conveniently and quickly obtain data without complex data processing processes and data transmission operations. At the same time, the API interface usually follows unified data formats and protocol standards, making data interaction between different systems more convenient and smooth.

[0033] For structured data without an API, the corresponding data reading method can be determined according to the data volume. Among them, for structured data with a small data volume, a simple and direct reading method can be adopted, so as to quickly obtain data and consume less system resources; for the case of a large data volume, more complex and efficient data reading and processing technologies, such as distributed data processing frameworks, need to be adopted to improve the efficiency and performance of data processing.

[0034] As can be seen from the above, in this embodiment, the API reading method is preferably adopted, and the perfect security mechanism and stable interface of the API can be utilized to ensure the security and stability of data transmission and reading. At the same time, a suitable data reading method is selected according to different data volumes, avoiding resource waste. When the data volume is small, a simple method can be adopted to reduce unnecessary resource overhead; when the data volume is large, an efficient method can be adopted to make full use of system resources and improve processing efficiency.

[0035] In an embodiment of the present disclosure, determining a data reading method based on the data volume of the first source data includes: If the data volume of the first source data is greater than the first quantity, determine the data reading method as the second method; the second method is a method of reading data based on an ETL tool. If the data volume of the first source data is less than or equal to the first quantity, determine the data reading method as the third method; the third method is a method of reading data based on a database query language.

[0036] In this embodiment, when the data volume of the first source data is greater than the first quantity, data can be read using an ETL (Extract - Transform - Load) tool. The ETL tool can efficiently extract data from multiple data sources in parallel. During the extraction process, operations such as cleaning and transforming the data can also be performed, such as handling missing values and error values in the data and converting the data into a unified format.

[0037] When the data volume of the first source data is less than or equal to the first quantity, data can be read using a database query language (such as SQL). The database query language has the characteristics of simplicity and directness. For small - scale data, it can quickly execute query operations to obtain the required data.

[0038] As can be seen from the above, in this embodiment, when the data volume of the first source data is large, the efficient processing ability of the ETL tool can quickly complete data reading and processing tasks; when the data volume of the first source data is small, the database query language can quickly obtain data, thus improving the overall data processing efficiency.

[0039] In an embodiment of the present disclosure, the classification corresponding to the multi - source data includes semi - structured data. Determining a data reading method based on the classification corresponding to the multi - source data includes: Construct a data relationship diagram according to the data dictionary of the second source data; the second source data is any semi - structured data in the multi - source data. If the proportion of the first type of nodes in the data relationship graph is greater than the first threshold, determine the data reading method as the fourth method; the fourth method is a method of reading data based on a parser; the first type of nodes are nodes with a corresponding node degree greater than the second threshold; If the proportion of the first type of nodes in the data relationship graph is less than or equal to the first threshold, determine the data reading method as the fifth method, and the fifth method is a method of reading data based on regular expressions.

[0040] In this embodiment, for semi-structured data with a relatively high degree of data structure complexity (such as the second source data), a parser can be used to read the data. The parser can handle complex data structures and accurately parse out each data element and its relationship by understanding the syntax and semantic rules of the data, ensuring the integrity and accuracy of data reading. For example, when processing semi-structured data in complex XML or JSON formats, the parser can extract data by delving into nested levels according to its structure rules.

[0041] For semi-structured data with a simple data structure, a pattern can be defined through regular expressions to match specific strings or patterns in the semi-structured data, and the data reading task can be quickly completed to extract the required information.

[0042] Among them, when determining the complexity of the semi-structured data, a data relationship graph can be constructed based on the data dictionary of the second source data. The nodes in the data relationship graph represent data types, and the edges represent the relationships between data types. If the degrees of a large number of nodes in the relationship graph are relatively high, that is, connected to multiple other nodes, it indicates that the connections between nodes are close and may form a complex network structure. For example, in a social network, if many users have friend relationships with a large number of other users, a complex social relationship network will be formed.

[0043] Accordingly, a second threshold can be set in advance. If the degree of a certain node in the relationship graph is greater than the second threshold, mark this node as the first type of node, and then count the proportion of the first type of nodes in the relationship graph. If the proportion of the first type of nodes is greater than the first threshold, it indicates that the degrees of a large number of nodes in the relationship graph are relatively high, indicating that the connections between nodes are close and a complex network structure is formed. On the contrary, if the proportion of the first type of nodes is less than or equal to the first threshold, it indicates that the data structure is relatively simple and the associations between nodes are not very complex.

[0044] From the above, it can be concluded that in this embodiment, the appropriate data reading method is automatically selected according to the complexity of the data structure. For data with a simple structure, regular expressions can be used to quickly extract data, saving processing time; for data with a complex structure, a parser can be used for accurate parsing to ensure the efficiency of data reading.

[0045] In an embodiment of the present disclosure, the multi-source data processing method further includes: If the change amount of the proportion of the first type of nodes is greater than the third threshold, adjust the second threshold based on a preset proportional parameter; the preset proportional parameter is a positive number less than 1.

[0046] In this embodiment, the change amount of the proportion of the first type of nodes reflects the change degree of the source data. When the change amount of the proportion of the first type of nodes is greater than the third threshold, it indicates that the change range of the second source data is relatively large. At this time, in order to avoid misjudgment of the data reading method when the proportion of the first type of nodes is small, the second threshold can be adjusted smaller, and the data reading method based on the parser is more adopted to ensure the correct reading of the data.

[0047] Specifically, a proportional parameter between 0 and 1 can be preset, and the proportional parameter is multiplied by the second threshold to obtain the adjusted second threshold.

[0048] It can be concluded from the above that this embodiment can dynamically adjust the second threshold according to the real-time change situation of the data structure, further ensuring the accurate reading of the data.

[0049] In an embodiment of the present disclosure, filling the first data into the corresponding field includes: If there are multiple first data corresponding to the same field, and the data source corresponding to the second data has a priority flag, fill the second data into the same field; the second data is any one of the first data, and the priority flag is determined by looking up a preset mapping relationship; If there are multiple first data corresponding to the same field, and the data sources corresponding to the multiple first data do not have priority flags, perform a weighted sum of the multiple first data to obtain a third data, and fill the third data into the same field.

[0050] In this embodiment, if there is only one first data corresponding to the same field, write the first data into the corresponding field in the data template.

[0051] If there are multiple first data corresponding to the same field, it is necessary to determine the first data finally written into the corresponding field in the data template according to the actual situation of the multi-source data. Specifically, a mapping relationship can be pre-constructed based on common data sources. The mapping relationship can be in the form of a table. Each row in the table includes a data source and the corresponding priority flag. If a data source belongs to an authoritative institution, its corresponding priority flag is "yes", indicating that it has a priority flag; if a data source does not belong to an authoritative institution, its corresponding priority flag is "no", indicating that it does not have a priority flag.

[0052] For example, in the field of academic research, when conducting a literature review or statistical data analysis, researchers prefer to refer to data from authoritative academic journals and well-known research institutions. Therefore, for authoritative institutions in a specific field, their priority identification can be set to "yes", while for other institutions, their priority identification can be set to "no".

[0053] Based on this, it is possible to determine whether the multiple data sources corresponding to the multiple first data have priority identification by looking up the above mapping relationship. If the data source corresponding to a certain first data (denoted as the second data) has priority identification, it indicates that the data source comes from an authoritative institution, and the second data is written into the corresponding field in the data template. If none of the data sources corresponding to the multiple first data have priority identification, a more representative and comprehensive value can be calculated by assigning corresponding weights to each first data and filling the corresponding field in the data template.

[0054] It can be concluded from the above that this embodiment preferentially selects data sources with priority identification for filling the data template, which can reduce the influence of incorrect or low-quality data. When there is no priority identification, the weighted summation method can comprehensively consider the influence of multiple data, making the filled data better reflect the real situation and improving the accuracy and credibility of the data.

[0055] In an embodiment of the present disclosure, the multi-source data processing method further includes: Determining the weights of the multiple first data based on their update times.

[0056] In this embodiment, considering that the update of multi-source data is dynamic, the data update times of different data sources may vary. The closer the update time is, the more accurate the real situation reflected by the data is, and the higher the relevance to the current business scenario. Therefore, determining the weights based on the update time can assign higher weights to the newer data, enabling these data to play a greater role in subsequent data processing, analysis, and decision-making.

[0057] It can be concluded from the above that this embodiment determines the weights based on the update time, and the updated data has greater weights, thereby improving the accuracy and timeliness of the data fusion result and helping decision-makers make correct decisions in a timely manner.

[0058] Corresponding to the multi-source data processing method in the above embodiments, Figure 2 is a structural block diagram of a multi-source data processing system provided by an embodiment of the present disclosure. For the sake of convenience of description, only the parts related to the embodiments of the present disclosure are shown. Refer to Figure 2 , the multi-source data processing system 20 includes: a configuration information input module 21, a data classification module 22, a data reading module 23, and a data fusion module 24. Among them, the configuration information input module 21 is used to construct a data template based on the configuration information input by the user through the configuration interface; the data template includes multiple fields; The data classification module 22 is used to determine the data reading method based on the classification corresponding to the multi-source data; The data reading module 23 is used to read the first data corresponding to each field from the multi-source data based on the data reading method and the data template; The data fusion module 24 is used to fill the first data into the corresponding fields to implement the fusion processing of the multi-source data.

[0059] In an embodiment of the present disclosure, the data classification module 22 is specifically used for: If the first source data has a corresponding API, determine the data reading method as the first method; the first method is the method of reading data based on the API; the first source data is any one of the structured data in the multi-source data; If the first source data does not have a corresponding API, determine the data reading method based on the data volume of the first source data.

[0060] In an embodiment of the present disclosure, the data classification module 22 is specifically further used for: If the data volume of the first source data is greater than the first quantity, determine the data reading method as the second method; the second method is the method of reading data based on the ETL tool; If the data volume of the first source data is less than or equal to the first quantity, determine the data reading method as the third method; the third method is the method of reading data based on the database query language.

[0061] In an embodiment of the present disclosure, the data classification module 22 is specifically used for: Construct a data relationship diagram according to the data dictionary of the second source data; the second source data is any one of the semi-structured data in the multi-source data; If the proportion of the first type of nodes in the data relationship diagram is greater than the first threshold, determine the data reading method as the fourth method; the fourth method is the method of reading data based on the parser; the first type of nodes are the nodes with a corresponding node degree greater than the second threshold; If the proportion of the first type of nodes in the data relationship diagram is less than or equal to the first threshold, determine the data reading method as the fifth method, and the fifth method is the method of reading data based on the regular expression.

[0062] In an embodiment of the present disclosure, the data classification module 22 is specifically further used for: If the change amount of the proportion of the first type of nodes is greater than the third threshold, adjust the second threshold based on the preset proportional parameter; the preset proportional parameter is a positive number less than 1.

[0063] In one embodiment of the present disclosure, the data fusion module 24 is specifically configured to: If there are multiple pieces of first data corresponding to the same field and the data source corresponding to the second data has a priority identifier, fill the second data into the same field; the second data is any one of the first data, and the priority identifier is determined by looking up a preset mapping relationship; If there are multiple pieces of first data corresponding to the same field and none of the data sources corresponding to the multiple pieces of first data have a priority identifier, perform weighted summation on the multiple pieces of first data to obtain third data, and fill the third data into the same field.

[0064] In one embodiment of the present disclosure, the data fusion module 24 is further specifically configured to: Determine the weights of the multiple pieces of first data based on the update times of the multiple pieces of first data.

[0065] Refer to Figure 3 , Figure 3 which is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 3 shown, the electronic device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store a computer program, and the computer program includes program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above-mentioned device embodiments, such as Figure 2 the functions of the modules 21 to 24 shown.

[0066] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0067] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.

[0068] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may further include a non-volatile random access memory. For example, the memory 304 may further store information about the device type.

[0069] In a specific implementation, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present disclosure may implement the implementation manners described in the first embodiment and the second embodiment of the multi-source data processing method provided by the embodiments of the present disclosure, and may also implement the implementation manner of the electronic device described in the embodiments of the present disclosure, which will not be elaborated herein.

[0070] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the methods of the above embodiments are implemented. It may also be completed by instructing relevant hardware through the computer program. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments may be implemented. Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0071] The computer-readable storage medium may be an internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the computer-readable storage medium may further include both an internal storage unit and an external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store the data that has been output or will be output.

[0072] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.

[0073] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0074] In several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces or units, and can also be electrical, mechanical, or other forms of connection.

[0075] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this disclosure.

[0076] In addition, the functional units in each embodiment of this disclosure can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0077] The above is only the specific implementation manner of this disclosure, but the protection scope of this disclosure is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by this disclosure, and these modifications or substitutions should all be covered by the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be subject to the protection scope of the claims.

Claims

1. A multi-source data processing method, characterized in that: include: Building a data template based on configuration information input by a user through a configuration interface; The data template includes a plurality of fields; Determine the data reading method based on the classification corresponding to the multi-source data; Based on the data reading method and the data template, read the first data corresponding to each field from the multi-source data; The first data is filled into corresponding fields to realize fusion processing of the multi-source data.

2. The multi-source data processing method according to claim 1, characterized in that: The classification corresponding to the multi-source data includes structured data. The data reading method is determined based on the classification corresponding to the multi-source data, including: If the first source data has a corresponding API, the data reading method is determined to be a first method; the first method is a method of reading data based on the API; the first source data is any structured data among multiple source data; If the first source data does not have a corresponding API, the data reading method is determined based on the data amount of the first source data.

3. The multi-source data processing method according to claim 2, characterized in that: The determining of the data reading method based on the data amount of the first source data includes: If the amount of the first source data is greater than the first amount, the data reading method is determined to be a second method; the second method is a method of reading data based on an ETL tool; If the data volume of the first source data is less than or equal to the first quantity, the data reading method is determined to be a third method; the third method is a method of reading data based on a database query language.

4. The multi-source data processing method according to claim 1, characterized in that: The classification corresponding to the multi-source data includes semi-structured data. The data reading method is determined based on the classification corresponding to the multi-source data, including: Constructing a data relationship diagram according to a data dictionary of second source data; the second source data is any type of semi-structured data among multi-source data; If the proportion of the first type of nodes in the data relationship graph is greater than the first threshold, the data reading method is determined to be the fourth method; the fourth method is a method of reading data based on a parser; the first type of nodes are nodes whose corresponding node degrees are greater than the second threshold; If the proportion of the first type of nodes in the data relationship graph is less than or equal to the first threshold, the data reading method is determined to be a fifth method, and the fifth method is a method of reading data based on a regular expression.

5. The multi-source data processing method according to claim 4, characterized in that: Also includes: If the change in the proportion of the first type of nodes is greater than the third threshold, adjusting the second threshold based on a preset proportion parameter; The preset ratio parameter is a positive number less than 1.

6. The multi-source data processing method according to claim 1, characterized in that: The step of filling the first data into the corresponding field includes: If there are multiple first data corresponding to the same field, and the data source corresponding to the second data has a priority identifier, the second data is filled into the same field; the second data is any first data, and the priority identifier is determined by searching for a preset mapping relationship; If there are multiple first data corresponding to the same field, and the data sources corresponding to the multiple first data do not have a priority identifier, the multiple first data are weightedly summed to obtain third data, and the third data is filled into the same field.

7. The multi-source data processing method according to claim 6, characterized in that: Also includes: The weights of the plurality of first data are determined based on update times of the plurality of first data.

8. A multi-source data processing system, characterized in that: include: A configuration information input module, used to construct a data template based on the configuration information input by the user through the configuration interface; The data template includes a plurality of fields; A data classification module is used to determine the data reading method based on the classification corresponding to the multi-source data; A data reading module, configured to read first data corresponding to each field from the multi-source data based on the data reading method and the data template; The data fusion module is used to fill the first data into corresponding fields to realize the fusion processing of the multi-source data.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.