Data integration method and device
The proposed data integration method addresses inefficiencies by parallel metadata tagging and rule-based query generation, enhancing efficiency and reducing resource consumption for improved system stability.
Patent Information
- Application Number
- CN201911121056.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-15
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2039-11-15
AI Technical Summary
The prior art has problems such as low processing efficiency, large memory consumption, high CPU consumption, and poor system stability and availability when integrating large data.
By setting tags in parallel to the metadata in multiple data sets and storing the tag information into the distributed search engine, aggregation query statements are generated, and aggregation query is performed based on these statements to achieve data integration.
It improves the processing efficiency of data integration, reduces the memory and CPU consumption of stand-alone servers, and improves the stability and availability of the system.
Smart Images

Figure CN112818026B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a data integration method and apparatus. Background Art
[0002] In the prior art, when integrating multi-source data sets, mainly manual processing methods and specific system processing methods are used. In the manual processing method, operators can use document tools such as Excel to integrate and process multi-source data sets. For a small amount of data, this processing method has a small cost and fast processing speed. However, for a large amount of data, the manual processing method can basically not complete the processing.
[0003] For a large amount of data, currently mainly a specific system is used for integration processing, and its processing process includes: first loading multiple data sets into the memory of a certain server, and then sequentially integrating and processing multiple data sets according to a preset integration rule. For example, assume that there are a total of four data sets data set1 、data set2 、data set3 、data set4 , and the preset integration rule is data set1 ∩data ste2 ∪data set3 -data set4 , then after loading these four data sets into the memory, first take the intersection of the data sets data set1 and data set2 , then take the union with the data set data set3 , and then take the difference set with the data set data set4 .
[0004] In the process of implementing the present invention, the inventors found that the existing data integration solutions based on specific systems have at least the following problems: First, the data integration process can only be executed sequentially according to the integration rule, and the processing process cannot be executed in parallel, resulting in low data integration processing efficiency and unable to evaluate the processing time; Second, when performing data integration, all data sets need to be loaded into the memory of a certain server, resulting in a large memory consumption of the single-machine server; Third, due to the long time required for the data integration process, the CPU of the single-machine server is consumed highly for a long time, thereby affecting the stability and availability of the entire system. Summary of the Invention
[0005] In view of this, the present invention provides a data integration method and apparatus, which can improve the processing efficiency of data integration, reduce the consumption of resources such as the memory and CPU of the single-machine server during the integration process, and improve the stability and availability of the system.
[0006] To achieve the above object, according to one aspect of the present invention, a data integration method is provided.
[0007] The data integration method of the present invention includes: parallelly setting tags for metadata in multiple data sets, and storing the obtained tag information of the metadata in a distributed retrieval engine; wherein, the tag information of the metadata includes the data set identifier where the metadata is located; parsing a pre-set data set integration rule to generate a corresponding aggregation query statement; and performing an aggregation query on the tag information of the metadata based on the aggregation query statement to obtain an aggregation query result.
[0008] Optionally, the step of parallelly setting tags for metadata in multiple data sets and storing the obtained tag information of the metadata in a distributed retrieval engine includes: splitting each of the multiple data sets into multiple data blocks, and parallelly setting tags for the metadata in the multiple data blocks; wherein, the parallelly setting tags for the metadata in the multiple data blocks includes: for the metadata in each data block, determining whether the data set identifier where the metadata is located exists in the tag set corresponding to the metadata; in the case where the determination result is negative, adding the data set identifier where the metadata is located to the tag set corresponding to the metadata, and then performing an insert update operation on the distributed retrieval engine according to the tag set corresponding to the metadata.
[0009] Optionally, the step of parsing a pre-set data set integration rule to generate a corresponding aggregation query statement includes: converting the pre-set data set integration rule into a digital expression; wherein, the digital expression is composed of digital operands and operators; generating a corresponding aggregation query statement according to the digital expression; wherein, the operator includes at least one of the following: an intersection operator, a union operator, or a difference operator.
[0010] Optionally, the distributed retrieval engine includes: an Elastic Search engine.
[0011] To achieve the above object, according to another aspect of the present invention, a data integration device is provided.
[0012] The data integration device of the present invention includes: a tagging module for parallelly setting tags for metadata in multiple data sets and storing the obtained tag information of the metadata in a distributed retrieval engine; wherein, the tag information of the metadata includes the data set identifier where the metadata is located; a generation module for parsing a pre-set data set integration rule to generate a corresponding aggregation query statement; and a query module for performing an aggregation query on the tag information of the metadata based on the aggregation query statement to obtain an aggregation query result.
[0013] Optionally, the tagging module sets tags for metadata in multiple datasets in parallel and stores the tag information of the obtained metadata in the distributed retrieval engine, including: the tagging module splits each of the multiple datasets into multiple data blocks and sets tags for the metadata in the multiple data blocks in parallel; wherein, the tagging module setting tags for the metadata in the multiple data blocks in parallel includes: for the metadata in each data block, the tagging module determines whether the dataset identifier where the metadata is located exists in the tag set corresponding to the metadata; in the case where the determination result is negative, the tagging module adds the dataset identifier where the metadata is located to the tag set corresponding to the metadata, and then performs an insert update operation on the distributed retrieval engine according to the tag set corresponding to the metadata.
[0014] Optionally, the generation module parses a pre-set dataset integration rule to generate a corresponding aggregation query statement, including: the generation module converts the pre-set dataset integration rule into a digital expression; wherein, the digital expression is composed of digital operands and operators; the generation module generates a corresponding aggregation query statement according to the digital expression; wherein, the operators include at least one of the following: an intersection operator, a union operator or a difference operator.
[0015] Optionally, the distributed retrieval engine includes: an Elastic Search engine.
[0016] To achieve the above object, according to another aspect of the present invention, an electronic device is provided.
[0017] The electronic device of the present invention includes: one or more processors; and a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the data integration method of the present invention.
[0018] To achieve the above object, according to yet another aspect of the present invention, a computer-readable medium is provided.
[0019] The computer-readable medium of the present invention stores a computer program thereon, and when the program is executed by a processor, the data integration method of the present invention is implemented.
[0020] One embodiment of the above invention has the following advantages or beneficial effects: By setting tags for metadata in multiple datasets in parallel, storing the obtained tag information of the metadata in a distributed retrieval engine, parsing the pre-set dataset integration rules to generate corresponding aggregate query statements, and performing an aggregate query on the tag information of the metadata based on the aggregate query statements to obtain an aggregate query result, the processing efficiency of data integration can be improved, the consumption of resources such as the memory and CPU of a single machine server during the integration process can be reduced, and the stability and availability of the system can be improved.
[0021] The further effects of the above non-conventional alternative ways will be described below in combination with specific embodiments. Brief Description of the Drawings
[0022] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:
[0023] Figure 1 is a schematic diagram of the main process of the data integration method according to the first embodiment of the present invention;
[0024] Figure 2 is a schematic diagram of the main process of the data integration method according to the second embodiment of the present invention;
[0025] Figure 3 is an exemplary process schematic diagram of parallel metadata tagging according to the second embodiment of the present invention;
[0026] Figure 4 is a schematic diagram of the storage structure of metadata tag information in a distributed retrieval engine according to the second embodiment of the present invention;
[0027] Figure 5 is a schematic diagram of the main modules of the data integration device according to the third embodiment of the present invention;
[0028] Figure 6 is a schematic diagram of the main modules of the data integration system according to the fourth embodiment of the present invention;
[0029] Figure 7 is an exemplary system architecture diagram to which the embodiments of the present invention can be applied;
[0030] Figure 8 is a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present invention. Detailed Description of the Embodiments
[0031] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0032] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0033] Embodiment 1
[0034] Figure 1 is a schematic diagram of the main process of the data integration method according to the first embodiment of the present invention. As Figure 1 shown, the data integration method of the embodiment of the present invention includes:
[0035] Step S101, parallelly set tag information for the metadata in multiple data sets, and store the obtained tag information of the metadata into a distributed retrieval engine.
[0036] Among them, the metadata can be understood as the attribute values that can distinguish each element in the data set or can uniquely represent each element. It can be an attribute value of the data in the data set or a combination of several attribute values of the data in the data set. Taking the user population data set obtained from multiple data sources (such as databases, text files, MQ message queues, etc.) in an e-commerce platform as an example, its metadata can be the user account value (which can be simply referred to as "user PIN").
[0037] In this step, the metadata of multiple data sets to be integrated can be parallelly labeled (the labeling can be understood as setting tags for the metadata) based on a timed worker (a task running framework) tool or a loop thread, etc., to improve the labeling efficiency. For example, assume there are three data sets data set1 、data set2 and data set3 , then a timed woker tool can be started for each data set to parallelly label the metadata of these three data sets. To further improve the labeling efficiency, each of these three data sets can also be split into multiple data blocks, and the metadata in the multiple data blocks can be parallelly labeled.
[0038] Among them, the tag information of the metadata includes the data set identifier where the metadata is located. For example, when labeling the metadata in the data set data set1 , the identifier of this data set (such as "No set1Add it (such as "0" or "1") to the data structure for storing the tag information of the metadata. Among them, the data structure for storing the tag information of the metadata can be an array, a linked list, etc.
[0039] Among them, the distributed retrieval engine may include: Elastic Search engine. The Elastic Search engine, abbreviated as ES, is a search engine based on Lucene.
[0040] Step S102: Parse the pre-set dataset integration rule to generate a corresponding aggregation query statement.
[0041] In one example, the dataset integration rule can be represented by an infix expression. An infix expression is a common representation method for arithmetic or logical formulas. In an infix expression, the operator is in the middle of the operands in infix form. For example, a dataset integration rule represented by an infix expression is data set1 ∩data ste2 ∪data set3 -data set4 .
[0042] In this step, by parsing the pre-set dataset integration rule, an aggregation query statement corresponding to the dataset integration rule can be obtained. Exemplarily, assume that the distributed retrieval engine is an Elastic Search engine, and the dataset integration rule is data set1 ∩data ste2 ∪data set3 -data set4 , then the parsed corresponding aggregation query statement can be expressed as:
[0043]
[0044] Step S103: Perform an aggregation query on the tag information of the metadata based on the aggregation query statement to obtain an aggregation query result.
[0045] In the embodiments of the present invention, by parallelly setting tags for the metadata in multiple datasets, storing the obtained tag information of the metadata in a distributed retrieval engine, parsing the pre-set dataset integration rule to generate a corresponding aggregation query statement, and performing an aggregation query on the tag information of the metadata based on the aggregation query statement to obtain an aggregation query result, the processing efficiency of data integration can be improved, the consumption of resources such as the memory and CPU of a single machine server during the integration processing can be reduced, and the stability and availability of the system can be improved.
[0046] Embodiment 2
[0047] Figure 2 It is a schematic diagram of the main process of the data integration method according to the second embodiment of the present invention. As Figure 2 shown, the data integration method of the embodiment of the present invention includes:
[0048] Step S201: Extract multiple data sets from different data sources based on pre-configured key data extraction rules.
[0049] Specifically, different forms of key data extraction rules (or metadata extraction rules) can be set for different data sources to obtain multiple data sets from different data sources. For example, for a data source such as a text file, the key data extraction rule can be represented by a regular expression; for a data source such as a database, the key data extraction rule can be represented by one or more filtering fields in the database.
[0050] Step S202: Set tags for the metadata in multiple data sets in parallel, and store the tag information of the obtained metadata in a distributed retrieval engine.
[0051] Among them, the metadata can be understood as the attribute values that can distinguish each element in the data set or can uniquely represent each element. It can be an attribute value of the data in the data set or a combination of several attribute values of the data in the data set. Taking the user population data set obtained from multiple data sources (such as databases, text files, MQ message queues, etc.) in an e-commerce platform as an example, its metadata can be the user account value (which can be simply referred to as "user PIN").
[0052] In this step, the metadata of multiple data sets to be integrated can be marked in parallel based on a timed worker (a task running framework) tool or a loop thread, etc. (the marking can be understood as setting tags for the metadata) to improve the marking efficiency. For example, assume there are three data sets data set1 、data set2 and data set3 , then a timed woker tool can be started for each data set to mark the metadata of these three data sets in parallel.
[0053] To further improve the marking efficiency, each of the multiple data sets can also be split into multiple data blocks, and the metadata in the multiple data blocks can be marked in parallel.
[0054] Among them, the tag information of the metadata includes the data set identifier where the metadata is located. For example, when marking the metadata in the data set data set1 , the identifier of this data set (such as "No set1Add it (such as "0" or "1", etc.) to the tag set corresponding to the metadata. When tagging the metadata in the dataset data set2 , the identifier of the dataset (such as "No set2 " or "2", etc.) can be added to the tag set corresponding to the metadata. Among them, the tag set corresponding to the metadata refers to the data structure for storing the tag information of the metadata, and this data structure can be an array, a linked list, etc.
[0055] In this step, tag the metadata in multiple datasets in parallel, and store the tag information of the tagged metadata in the distributed retrieval engine. Finally, the aggregated metadata tag information can be obtained (as Figure 4 shown). For example, assume that there is the same metadata (such as the user account value "1213") in dataset data set1 and dataset data set2 . Then, after tagging, the following metadata tag information is stored in the distributed retrieval engine: [1, 2], and this metadata tag information indicates that this metadata is both in dataset data set1 and in dataset data set2 .
[0056] Step S203: Convert the preset dataset integration rule into a digital expression.
[0057] In one example, the dataset integration rule can be represented by an infix expression. An infix expression is a general representation method for arithmetic or logical formulas. In an infix expression, the operator is in the middle of the operands in infix form. For example, a dataset integration rule represented by an infix expression is data set1 ∩data ste2 ∪data set3 -data set4 .
[0058] Among them, the digital expression is composed of digitalized operands and operators in digital form, and the operators include at least one of the following: intersection operator, union operator, or difference operator.
[0059] For example, assume that the dataset integration rule is "data set1 ∩data ste2 ∪data set3 -data set4 ". Then, through this step, it can be converted to "1∩2∪3-4".
[0060] Step S204: Generate an aggregation query statement corresponding to the digital expression.
[0061] In this step, the digital expression can be scanned. When an operator is encountered during the scanning, a query clause can be constructed based on the operator and the operands. Then, all the constructed query clauses are nested and encapsulated to obtain the corresponding aggregation query statement.
[0062] For example, assume that the distributed retrieval engine is the Elastic Search engine, and the digital expression obtained through step S203 is "1 ∩ 2 ∪ 3 - 4". Then, when the intersection operator "∩" is encountered during the scanning, the following query clause can be constructed based on the intersection operator and the operands "1" and "2" on both sides: "filter": {"item": {"tag": [1, 2]}}; when the union operator "∪" is encountered during the scanning, the following query clause can be constructed based on the union operator and the right operand "3": "should": {"item": {"tag": 3}}; when the difference operator "-" is encountered during the scanning, the following query clause can be constructed based on the difference operator and the right operand "4": "must_not": {"item": {"tag": "4"}}; Then, the above query clauses are nested and encapsulated with the bool clause to obtain the corresponding aggregation query statement.
[0063] Step S205, perform an aggregation query on the tag information of the metadata based on the aggregation query statement to obtain an aggregation query result.
[0064] In the embodiment of the present invention, by parallelly labeling the metadata in multiple data sets, the labeling efficiency is improved; by storing the tag information of the metadata obtained by labeling into a distributed retrieval engine, and by parsing the pre-set data set integration rules to generate the corresponding aggregation query statement, and further, performing an aggregation query on the tag information of the metadata based on the aggregation query statement, the processing efficiency of data integration can be improved, the consumption of resources such as the memory and CPU of a single machine server during the integration processing can be reduced, and the stability and availability of the system can be improved.
[0065] Figure 3 is an exemplary flowchart of parallel metadata labeling according to the second embodiment of the present invention. After splitting the data set into multiple data blocks, the metadata in the multiple data blocks can be parallelly labeled based on Figure 3 the shown process. As Figure 3 shown, the exemplary process of parallelly labeling the metadata in the data block includes:
[0066] Step S301, for the metadata in the data block, obtain the corresponding tag set of the metadata in the distributed retrieval engine.
[0067] Step S302: Determine whether there is a tag set corresponding to the metadata. If there is a tag set corresponding to the metadata, step S303 can be executed; otherwise, step S304 can be executed.
[0068] In this step, if a tag set corresponding to the metadata is obtained from the distributed retrieval engine, it is determined that there is a tag set corresponding to the metadata; if a tag set corresponding to the metadata cannot be obtained from the distributed retrieval engine, it is determined that there is no tag set corresponding to the metadata.
[0069] Step S303: Determine whether the dataset identifier where the metadata is located exists in the tag set. If the dataset identifier where the metadata is located does not exist in the tag set, step S305 can be executed; otherwise, step S307 can be executed.
[0070] Step S304: Create a tag set corresponding to the metadata.
[0071] Step S305: Add the dataset identifier where the metadata is located to the tag set. After step S305, step S306 can be executed.
[0072] For example, when tagging the metadata in data block Block1, assuming that the data block belongs to dataset data set1 , the identifier of the dataset (such as "No set1 " or "1", etc.) can be added to the tag set corresponding to the metadata; when tagging the metadata in data block Block1, assuming that the data block belongs to dataset data set2 , the identifier of the dataset (such as "No set2 " or "2", etc.) can be added to the tag set corresponding to the metadata.
[0073] Step S306: Perform an insert and update operation on the distributed retrieval engine according to the tag set corresponding to the metadata.
[0074] Exemplarily, when the distributed retrieval engine is an Elastic Search engine, the tag information stored in the tag set corresponding to the metadata can be stored in the Elastic Search engine through an upsert operation.
[0075] Step S307: Obtain the next metadata in the data block. After step S307, step S301 can be executed for the obtained next metadata.
[0076] In the embodiments of the present invention, through the above steps, it is possible to perform parallel tagging on the metadata of multiple data sets, improving the tagging efficiency; further, by performing parallel tagging in a distributed environment, the operating pressure on a single machine server can be effectively reduced.
[0077] Figure 5 It is a schematic diagram of the main modules of the data integration device according to the third embodiment of the present invention. As Figure 5 shown, the data integration device 500 in the embodiments of the present invention includes: a tagging module 501, a generation module 502, and a query module 503.
[0078] The tagging module 501 is used to parallelly set tags for the metadata in multiple data sets and store the obtained tag information of the metadata in a distributed retrieval engine.
[0079] Among them, the metadata can be understood as the attribute values that can distinguish each element in the data set or can uniquely represent each element. It can be an attribute value of the data in the data set or a combination of several attribute values of the data in the data set. Taking the user population data set obtained from multiple data sources (such as databases, text files, MQ message queues, etc.) in an e-commerce platform as an example, its metadata can be the user account value (which can be simply referred to as "user PIN").
[0080] Exemplarily, the tagging module 501 can perform parallel tagging on the metadata of multiple data sets to be integrated (the tagging can be understood as setting tags for the metadata) based on a timed worker (a task running framework) tool or a loop thread, etc., to improve the tagging efficiency. For example, assuming there are three data sets data set1 、data set2 and data set3 , then a timed woker tool can be started for each data set to perform parallel tagging on the metadata of these three data sets. To further improve the tagging efficiency, each of these three data sets can also be split into multiple data blocks, and parallel tagging is performed on the metadata in the multiple data blocks.
[0081] Among them, the tag information of the metadata includes the data set identifier where the metadata is located. For example, when tagging the metadata in the data set data set1 , the identifier of this data set (such as "No set1 " or "1", etc.) can be added to the tag set corresponding to this metadata. Among them, the tag set refers to the data structure used to store the tag information of this metadata, and it can be an array, a linked list, etc.
[0082] Among them, the distributed retrieval engine may include: Elastic Search engine. The Elastic Search engine, abbreviated as ES, is a search engine based on Lucene.
[0083] The generation module 502 is configured to parse a pre-set data set integration rule to generate a corresponding aggregation query statement.
[0084] Exemplarily, the generation module 502 parses a pre-set data set integration rule to generate a corresponding aggregation query statement, including: the generation module 502 converts the pre-set data set integration rule into a digital expression; wherein, the digital expression is composed of digital operands and operators; the generation module 502 generates a corresponding aggregation query statement according to the digital expression. Wherein, the operator includes at least one of the following: intersection operator, union operator or difference operator.
[0085] In one example, the data set integration rule can be represented by an infix expression. The infix expression is a general representation method of arithmetic or logical formulas. In the infix expression, the operator is in the middle of the operands in infix form. For example, a data set integration rule represented by an infix expression is data set1 ∩data ste2 ∪data set3 -data set4 . Further, assuming that the distributed retrieval engine is the Elastic Search engine, the aggregation query statement generated by the generation module 502 corresponding to the data set integration rule "data set1 ∩data ste2 ∪data set3 -data set4 " can be expressed as:
[0086]
[0087]
[0088] The query module 503 is configured to perform an aggregation query on the tag information of the metadata based on the aggregation query statement to obtain an aggregation query result.
[0089] In the embodiments of the present invention, the metadata in multiple data sets is marked in parallel by a marking module, and the tag information of the marked metadata is stored in a distributed retrieval engine. The preset data set integration rules are parsed by a generation module to generate corresponding aggregation query statements. Moreover, an aggregation query is performed on the tag information of the metadata based on the aggregation query statements by a query module to obtain an aggregation query result, which can improve the processing efficiency of data integration, reduce the consumption of resources such as the memory and CPU of a single machine server during the integration processing, and improve the stability and availability of the system.
[0090] Figure 6 is a schematic diagram of the main modules of the data integration system according to the fourth embodiment of the present invention. As Figure 6 shown, the data integration system 600 in the embodiments of the present invention includes: a multi-source data set collection device 601, a data integration device 602, an integration rule configuration terminal 603, a data distribution engine 604, a data output terminal 605, and a distributed retrieval engine 606.
[0091] The multi-source data set collection device 601 is used to extract multiple data sets from different data sources based on pre-configured key data extraction rules. Exemplarily, the data sources may be text files, databases, or MQ (Message Queue) message queues.
[0092] Specifically, for different data sources, different forms of key data extraction rules (or metadata extraction rules) can be set in the multi-source data set collection device 601 to obtain multiple data sets from different data sources. For example, for a text file data source, the key data extraction rules can be represented by regular expressions; for a database data source, the key data extraction rules can be represented by one or more filtering fields in the database.
[0093] The integration rule configuration terminal 603 is used to configure data set integration rules. In the integration rule configuration terminal 603, a set of function operation controls can be designed on the interaction interface. For example, function operation controls for adding, editing, deleting, and querying integration rules can be designed to facilitate the flexible configuration and management of integration rules by system users. Among them, the integration rules may include at least one of the following operators: an intersection operator (represented by the mathematical symbol ∩), a union operator (represented by the mathematical symbol ∪), and a difference operator (represented by the mathematical symbol -). Exemplarily, if the data set integration rule is configured as data set1 ∩data set2 , it means to take the same data in data set data set1 and data set data set2 ; if the data set integration rule is configured as data set1 ∪dataset2 which represents taking the union of the data set data set1 and the data set data set2 ; if the data set integration rule configuration is data set1 -data set2 which represents taking the difference set of the data set data set1 and the data set data set2 to obtain a set composed of the data that exists in the data set data set1 but does not exist in the data set data set2 .
[0094] In specific implementation, after configuring the data set integration rule through the integration rule configuration terminal 603, the configured data set integration rule can be persistently stored. For example, the configured data set integration rule can be stored in a database or a centralized cache.
[0095] The data integration device 602 is used to set tags for the metadata in multiple data sets in parallel and store the obtained tag information of the metadata in the distributed retrieval engine 606; parse the pre-set data set integration rule to generate a corresponding aggregation query statement; perform an aggregation query on the tag information of the metadata based on the aggregation query statement to obtain an aggregation query result.
[0096] The distributed retrieval engine 606 can adopt the Elastic Search engine. The Elastic Search engine, abbreviated as ES, is a search engine based on Lucene. The ES engine can achieve real-time query of a large amount of data, which provides a basis for the integration of multiple data sets in the embodiments of the present invention.
[0097] The data distribution engine 604 is used to distribute the aggregation query result, that is, the integrated data. For example, the integrated data can be distributed to a certain text file, a certain database, a certain MQ message queue, and a network cache (such as a redis cluster).
[0098] The data output terminal 605 is used to output and display the integrated data.
[0099] The embodiments of the present invention construct a complete data integration system, which can not only realize the autonomous configuration of multi-source data set collection rules and data set integration rules, has wide applicability and ease of use, but also can effectively avoid the disadvantages existing in data integration in the prior art, improve the integration efficiency of multi-source data sets, reduce the consumption of the memory and CPU of a single machine server for data integration, and improve the stability and availability of the entire system.
[0100] Figure 7FIG. 700 shows an exemplary system architecture to which the data integration method or data integration apparatus according to embodiments of the present invention can be applied.
[0101] As Figure 7 shown, the system architecture 700 may include terminal devices 701, 702, 703, a network 704, and a server 705. The network 704 is used to provide a medium for communication links between the terminal devices 701, 702, 703 and the server 705. The network 704 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0102] Users may use the terminal devices 701, 702, 703 to interact with the server 705 through the network 704 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 701, 702, 703, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0103] The terminal devices 701, 702, 703 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0104] The server 705 may be a server providing various services, such as a background management server that supports data integration requests initiated by users using the terminal devices 701, 702, 703. The background management server may analyze and process data such as received data integration requests, and feedback the processing results (such as data integration results) to the terminal devices.
[0105] It should be noted that the data integration method provided by the embodiments of the present invention is generally executed by the server 705. Correspondingly, the data integration apparatus is generally disposed in the server 705.
[0106] It should be understood that Figure 7 the numbers of terminal devices, networks, and servers in
[0107] are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. Figure 8 FIG. 800 shows a schematic structural diagram of a computer system of an electronic device suitable for implementing embodiments of the present invention. Figure 8 The shown computer system is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0108] As Figure 8As shown, computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the system 800 are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0109] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage section 808 as needed.
[0110] Specifically, according to an embodiment disclosed by the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment disclosed by the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the central processing unit (CPU) 801, the above functions defined in the system of the present invention are executed.
[0111] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0113] The modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a marking module, a generating module, and a query module. Among them, the names of these modules do not constitute a limitation to the module itself in some cases. For example, the marking module can also be described as "a module for parallel marking of metadata in multiple data sets".
[0114] As another aspect, the present invention also provides a computer-readable medium, which can be included in the device described in the above embodiments; or can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the device performs the following processes: parallelly set labels for the metadata in multiple data sets, and store the obtained label information of the metadata into a distributed retrieval engine; wherein, the label information of the metadata includes the data set identifier where the metadata is located; parse the pre-set data set integration rules to generate corresponding aggregation query statements; perform an aggregation query on the label information of the metadata based on the aggregation query statements to obtain an aggregation query result.
[0115] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data integration method, characterized in that The method includes: Parallelly setting labels for metadata in multiple datasets, and storing the label information of the obtained metadata into a distributed retrieval engine; wherein, the label information of the metadata includes the dataset identifier where the metadata is located; Parsing a preset dataset integration rule to generate a corresponding aggregation query statement; Performing an aggregation query on the label information of the metadata based on the aggregation query statement to obtain an aggregation query result; The step of parallelly setting labels for metadata in multiple datasets and storing the label information of the obtained metadata into a distributed retrieval engine includes: Splitting each of the multiple datasets into multiple data blocks, and parallelly setting labels for the metadata in the multiple data blocks; wherein, the parallelly setting labels for the metadata in the multiple data blocks includes: for the metadata in each data block, determining whether the dataset identifier where the metadata is located exists in the label set corresponding to the metadata; in the case where the determination result is negative, adding the dataset identifier where the metadata is located to the label set corresponding to the metadata, and then performing an insert update operation on the distributed retrieval engine according to the label set corresponding to the metadata.
2. The method according to claim 1, wherein The step of parsing a preset dataset integration rule to generate a corresponding aggregation query statement includes: Converting a preset dataset integration rule into a digital expression; wherein, the digital expression is composed of digital operands and operators; generating a corresponding aggregation query statement according to the digital expression; wherein, the operators include at least one of the following: an intersection operator, a union operator, or a difference operator.
3. The method according to claim 1, wherein The distributed retrieval engine includes: an ElasticSearch engine.
4. A data integration device, characterized in that, The device includes: A tagging module for parallelly setting labels for metadata in multiple datasets and storing the label information of the obtained metadata into a distributed retrieval engine; wherein, the label information of the metadata includes the dataset identifier where the metadata is located; A generating module for parsing a preset dataset integration rule to generate a corresponding aggregation query statement; A query module for performing an aggregation query on the label information of the metadata based on the aggregation query statement to obtain an aggregation query result; The tagging module parallelly setting labels for metadata in multiple datasets and storing the label information of the obtained metadata into a distributed retrieval engine includes: The marking module splits each of the multiple data sets into multiple data blocks and sets labels for the metadata in the multiple data blocks in parallel; wherein, the marking module setting labels for the metadata in the multiple data blocks in parallel includes: for the metadata in each data block, the marking module determines whether the data set identifier where the metadata is located exists in the label set corresponding to the metadata; in the case where the determination result is negative, the marking module adds the data set identifier where the metadata is located to the label set corresponding to the metadata, and then performs an insert update operation on the distributed retrieval engine according to the label set corresponding to the metadata.
5. The device according to claim 4, characterized in that, The generation module parses the pre-set data set integration rule to generate a corresponding aggregation query statement, including: The generation module converts the pre-set data set integration rule into a digital expression; wherein, the digital expression is composed of digital operands and operators; the generation module generates a corresponding aggregation query statement according to the digital expression; wherein, the operators include at least one of the following: intersection operator, union operator or difference operator.
6. The device according to claim 4, characterized in that, The distributed retrieval engine includes: ElasticSearch engine.
7. An electronic device, characterized in that, Including: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 3.
8. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Method and device for label production based on electricity marketing data
CN107145586A