A multi-source data fusion method and device, electronic equipment and storage medium
By performing quality assessment and deduplication on multi-source point of interest data, the problems of insufficient non-spatial attribute matching and reliance on human experience for model weights in existing technologies are solved, achieving higher quality and more objective data fusion results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-11
- Publication Date
- 2026-04-07
AI Technical Summary
Existing multi-source point of interest data fusion methods, while achieving high accuracy in spatial geographic location relationships, lack matching calculations for other non-spatial attributes. Furthermore, the model weight threshold settings rely on human experience, resulting in fusion results being significantly influenced by subjectivity.
By acquiring data from source and parent databases, quality assessment and deduplication are performed. The results of the quality assessment and deduplication are then used for data fusion to improve data quality and avoid the impact of low-quality data.
It improves the quality of data after multi-source data fusion, reduces subjective influence, and enhances the objectivity and accuracy of data fusion.
Smart Images

Figure CN117195149B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data fusion, and in particular to a multi-source data fusion method and device, electronic equipment and storage medium. BACKGROUND
[0002] Multi-source data fusion technology refers to a technology of using relevant means to comprehensively integrate all information obtained through investigation and analysis, uniformly evaluating the information, and finally obtaining unified information. The purpose of developing this technology is to comprehensively integrate various data information, extract the characteristics of different data sources, and then extract unified information that is better and richer than single data.
[0003] At present, in a multi-source point of interest (POI) data fusion method based on single common attribute constraint, such as a method of correcting the position of original data through the topological relationship between two-dimensional building surface data and POI, although a relatively accurate spatial geographic position relationship can be obtained, the method lacks matching calculation of other non-spatial attributes, and the quality of the fused data set is poor. In a POI data fusion method based on attribute weighting, a threshold is determined through expert scoring and experience value, so as to obtain a regression weight, and finally NLP technology is used to perform text semantic analysis on POI attributes. Although this method considers the influence of separate processing of different POI attributes, due to the obvious differentiation of POI characteristics in various regions across the country, the setting of model weight threshold depends on artificial experience value, and the setting of weighted weight has always been a difficulty of this method, which finally leads to a large degree of subjective influence on the fusion result. SUMMARY
[0004] The present application provides a multi-source data fusion method, device, electronic equipment and storage medium to solve the above technical problems.
[0005] According to an aspect of the present application, a multi-source data fusion method is provided, comprising:
[0006] obtaining source data in at least one source database and mother database data in a mother database;
[0007] performing quality evaluation on the source data and the mother database data to obtain a quality evaluation result;
[0008] performing duplicate detection processing on any source data and any mother database data to obtain a duplicate detection result;
[0009] performing data fusion on the source data and the mother database data based on the duplicate detection result and the quality evaluation result.
[0010] According to another aspect of the present application, a multi-source data fusion device is provided, comprising:
[0011] A data acquisition model is used to acquire source data from at least one source database and parent database data from a parent database.
[0012] The quality assessment module is used to assess the quality of the source data and the parent database data, and obtain the quality assessment results.
[0013] The deduplication module is used to perform deduplication processing on any of the source data and any of the parent database data to obtain the deduplication result.
[0014] The data fusion module is used to perform data fusion on the source data and the parent database data based on the deduplication result and the quality assessment result.
[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the multi-source data fusion method according to any embodiment of the present invention.
[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the multi-source data fusion method according to any embodiment of the present invention.
[0020] The technical solution of this invention involves acquiring source data from at least one source database and parent database data from a parent database; performing quality assessment on the source data and parent database data to obtain quality assessment results; performing deduplication processing on any source data and any parent database data to obtain deduplication results; and fusing the source data and parent database data based on the deduplication results and quality assessment results. This approach increases the consideration of the quality of source data and parent database data, avoids the impact of low-quality data on the fused data, and improves the data quality of the fused parent database data.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of a multi-source data fusion method provided in Embodiment 1 of the present invention;
[0024] Figure 2 This is a schematic diagram of the name resolution structure provided in Embodiment 1 of the present invention;
[0025] Figure 3 This is a schematic diagram of the structure of a multi-source data fusion device provided in Embodiment 2 of the present invention;
[0026] Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] Example 1
[0030] Figure 1This is a flowchart of a multi-source data fusion method provided in Embodiment 1 of the present invention. This embodiment is applicable to the fusion of multi-source POI data. The method can be executed by a multi-source data fusion device, which can be implemented in hardware and / or software and can be configured in a multi-source data fusion equipment. Figure 1 As shown, the method includes:
[0031] S110. Obtain source data from at least one source database and parent database data from at least one parent database.
[0032] In this context, the source database refers to a database storing POI data from different sources. There can be multiple source databases, each corresponding to a specific source of data. The parent database is one of the source databases, and the merged POI data is stored in the parent database. In this embodiment, the data in both the source databases and the parent database are POI data. POI data includes multiple attribute fields, including one or more of the following: name, name alias, address, address alias, spatial location, type, and phone number.
[0033] S120. Perform a quality assessment on the source data and the parent database data to obtain the quality assessment result.
[0034] In this embodiment, after obtaining POI data, the quality of the POI data obtained from the source database and the parent database is first evaluated to obtain the quality evaluation results of each POI data.
[0035] Based on the above embodiments, optionally, the step of performing quality assessment on the source data and the parent database data to obtain quality assessment results includes: assigning quality scores to each attribute field of the point of interest data to obtain quality scores for each attribute field; inputting the quality scores into a quality classification model for quality assessment to obtain quality assessment results; wherein, the quality assessment results include overall quality assessment results and attribute field quality assessment results, the overall quality assessment result being the quality assessment result for the entire POI data, and the attribute field quality assessment results being the quality assessment result for a single attribute field.
[0036] In this embodiment, the quality score of POI data is obtained by scoring the data based on four attribute fields: address, name, type, and spatial location. The quality score is then used to evaluate the quality of the POI data.
[0037] The quality score for the POI address field includes address integrity score, address accuracy score, and address precision score. In this embodiment, the quality score of the address field is determined based on these three scores.
[0038] 1) The calculation method for address integrity score is as follows:
[0039] Research suggests that a complete POI address field consists of the following fragments:
[0040] "Province + City + District / County + Street / Town + Village + Road (Main Road / Secondary Road) + POI + Building + Unit + Floor + Room Number + Auxiliary Locator Keywords + Distance"
[0041] In the process of scoring the address completeness of the address field, geocoding tools are used to perform address segmentation and semantic recognition on the address field to obtain semantic fragments, and then the address accuracy score is determined based on the semantic fragments. It should be noted that this address accuracy score is used to characterize the accuracy level of the address, that is, the level of detail of the address. The more detailed the address, the higher the completeness of the address. This invention uses the address accuracy score as the address completeness score.
[0042] Specifically, the base accuracy score of an address is determined based on the finest level of precision described by the address. For example, if the semantic fragments after address splitting are only precise to the "province" level, then referring to Table 1, the base score is 20; if the semantic fragments after address splitting are precise to the "door address" level, then referring to Table 1, the base score is 100. For example:
[0043] Information Road, Haidian District, Beijing (Precision level: "Road", Score: 60);
[0044] No. 9, Jia, Xinxi Road, Beijing (Precision level: "address", score: 100);
[0045] Kuike Technology Building, Haidian District (Precision Level: "Business Building", Score: 90);
[0046] 300 meters east of the intersection of Xinxi Road and Shangdi Fifth Street (Precision level: "Intersection Offset", score: 70).
[0047] Address accuracy score expresses the most detailed level of accuracy in address description, and the scoring configuration is shown in Table 1.
[0048] Table 1 Address Accuracy Scoring Table
[0049]
[0050] The address accuracy score is calculated as follows:
[0051] a) Refer to Table 1 to score the address accuracy and obtain the basic accuracy score;
[0052] b) Analyze the last item of the address segmentation and adjust it based on the basic accuracy score: 1. If the category of the last segment after address segmentation is "address", adjust the score (if the basic accuracy score is >= 50, change it to 100; if the basic accuracy score is between 30 and 49, change it to 50); 2. If the category of the last segment after address segmentation is not "address", adjust the score (if the basic accuracy score is >= 90, change it to 80; if the basic accuracy score is < 90, subtract 10 points); 3. In particular, if the content of the last segment of the word segmentation is "inside", it is considered to have no offset, and the score is adjusted (if the basic accuracy score is >= 50, change it to 100; if the basic accuracy score is between 30 and 49, change it to 50). For example, 1. If the basic accuracy score is >= 50 and the last item is a door address, the score becomes 100. If the basic accuracy score is 30-49 and it is a door address, this address is a misidentification of a national highway or provincial highway (G7, S102). The range of such roads is very large, so the score is 50. 2. If the basic accuracy score is >= 90 and the last item is not a door address, then the basic range is a small POI, and the range after offset is similar to that of a large POI, so the score is 80. The scores of other categories are reduced by one level, so the corresponding score is reduced by 10 points. 3. If the last item of the word segmentation is "inside", such as "within the super community on Dongsan Road, Chaoyang District, Beijing", then the score adjustment rule is the same as 1.
[0053] 2) The address accuracy score is calculated as follows:
[0054] Spatial parsing is performed on the address field to calculate the spatial latitude and longitude corresponding to the address. Then, based on the spatial distance algorithm, the spatial distance between the spatial latitude and longitude and the POI latitude and longitude is calculated to determine the consistency between the address and the spatial location. The address accuracy is scored based on the degree of consistency between the address and the spatial location.
[0055] 3) The address accuracy score is calculated as follows:
[0056] The accuracy of the address determination based on the POI's administrative division is assessed. If the address contradicts the administrative division, the POI record is deleted.
[0057] In this embodiment, the method for determining the quality score of the POI spatial location field is as follows: based on the building extent data, determine the inclusion relationship (inclusion / non-inclusion) between the POI and the building extent data, calculate the distance between the POI data and the nearest building extent data, and perform a weighted calculation on the inclusion and distance indicators to obtain the POI coordinate quality score.
[0058] In some embodiments, the spatial coordinates of the POI can be used to calculate the fence entry, determine the accuracy of the POI's spatial location, and delete POI data that fail to enter the fence.
[0059] In this embodiment, the quality scoring method for the POI name field is as follows: the POI name field is scored according to the POI name resolution model to obtain the quality score of the POI name field.
[0060] In some embodiments, the accuracy of the name can be determined based on the POI address. If the address and name contradict each other, the POI data is deleted.
[0061] In this embodiment, the quality scoring method for POI type fields is as follows: the completeness of the POI type field is scored according to the classification table to obtain the POI type quality score. It can be understood that the granularity of POI classification is directly proportional to the score; the more detailed the classification, the higher the quality score.
[0062] In this embodiment, the quality scores of each attribute field of the POI data obtained by the above method are used as model inputs. These scores are fed into a quality assessment model to perform overall quality assessment and quality assessment of individual attribute fields on the POI data, resulting in overall quality assessment results and quality assessment results for each attribute field. The categories of the quality assessment results can be set by those skilled in the art according to their needs, and are not limited here; for example, the quality assessment results can be categorized into three quality classes: high, medium, and low.
[0063] In some embodiments, optionally, after obtaining the quality assessment results, the method further includes: filtering the source data based on the quality assessment results. Specifically, for any source data, the source data can be filtered based on the category of the overall quality assessment results of the source data. For example, if the category of the overall quality assessment results of the source data is low, then the source data is deleted, thereby retaining the source data with high and medium quality categories.
[0064] S130. For any of the source data, perform deduplication processing on the source data and the parent database data to obtain the deduplication result.
[0065] Based on the above embodiments, optionally, the step of performing deduplication processing on the source data and the parent database data for any of the source data to obtain a deduplication result includes: extracting attribute features of each attribute field in the source data and the parent database data; for any of the source data, determining the similarity score between each attribute field of the source data and each parent database data based on the attribute features of the source data and the attribute fields of each parent database data; generating feature engineering based on the similarity score, and inputting the feature engineering into the deduplication classification model for deduplication classification to obtain a deduplication result; wherein, the deduplication result is a binary classification result, and the deduplication result includes duplicate and non-duplicate.
[0066] In this context, attribute features refer to features used to evaluate the similarity score of a certain attribute field in the source data and the parent database data. In this embodiment, for any attribute field in the source data and the parent database data, the attribute features of that attribute field can be extracted, and then the similarity score of that attribute field in the source data and the parent database data can be calculated based on the attribute features. Further, for any source data, feature engineering corresponding to each parent database data can be generated based on the similarity scores of each attribute field in the source data and each parent database data in the parent database. The feature engineering corresponding to each parent database data is then sequentially input into the deduplication classification model to perform deduplication classification on the source data and each parent database data in the parent database, obtaining the deduplication result. The deduplication result is a binary classification result, including duplicate and non-duplicate categories. If the deduplication result is duplicate, it indicates that there is parent database data in the parent database that duplicates the points of interest of the source data, and this parent database data is recorded. If the deduplication result is non-duplicate, it indicates that there is no parent database data in the parent database that duplicates the points of interest of the source data. The deduplication classification model includes, but is not limited to, machine learning models, and is not limited here.
[0067] In this embodiment, the similarity scores for each attribute field are determined using different methods. For example, the methods for determining the similarity scores for each attribute field are as follows:
[0068] The method for determining name similarity scores is as follows: Lexical, syntactic, and semantic analyses are performed on the name fields in the source data and the parent database data to obtain lexical, syntactic, and semantic analysis results. These results are then integrated to obtain the POI name parsing results. Feature embedding is then performed on these POI name parsing results to obtain three name similarity features: suffix words, core words, and positional words from the source and parent database data. Name similarity is calculated based on these two POI data sets to obtain a name similarity score. It should be noted that when determining the name similarity score, the similarity between the source data name and the parent database name, as well as the similarity between the source data name and the aliases of the parent database name, must be calculated simultaneously. The alias database consists of names from source data that were similar to those in previous POI fusion processes but were not merged into the parent database data. For example, Figure 2 This is a schematic diagram of the name resolution structure provided in Embodiment 1 of the present invention, as shown below. Figure 2 As shown, the name is "Shanghai Kuike Technology Building Parking Lot" as an example.
[0069] The address similarity score is determined as follows: Address fields are parsed using a named body recognition model, and then address similarity is calculated using an Albert model to obtain the address similarity score. The named body recognition model includes, but is not limited to, the BERT-BiLSTM-CRF model; no specific limitation is made here. It should be noted that when determining the address similarity score, the similarity between the source data address and the parent database data address, as well as the similarity between the source data address and the alias of the parent database address, must be calculated simultaneously. The address alias library consists of addresses from source data that were similar to the parent database data addresses during previous POI fusion processes but were not fused into the parent database data.
[0070] The spatial location similarity score is determined as follows: The Euclidean distance between two Points of Interest (POIs) is calculated based on their spatial coordinates from the source data and the parent database. Then, the spatial location similarity between the two POIs is determined based on the Euclidean distance and the search radius corresponding to the POI type, thus obtaining the spatial location similarity score. The search radius corresponding to the POI category is manually compiled and iteratively determined.
[0071] The type similarity scoring method involves converting the types of both the source data and the parent database data into a unified POI type system, and then calculating the similarity based on a type similarity algorithm to obtain a type similarity score. The type similarity algorithm can be one of the following two:
[0072] 1) Calculate the type similarity between POIs using Dijkstra's shortest path algorithm;
[0073] 2) Based on the POI type system (Level 1, Level 2, Level 3, Level 4), manually formulated rules are used to calculate the type similarity between POIs. Specifically, if one POI is type 1011701 Food Service - Chinese Food - Hot Pot Restaurant - Sichuan Hot Pot, with the highest level of detail (Level 4, Level 1: Food Service, Level 2: Chinese Food, Level 3: Hot Pot Restaurant, Level 4: Sichuan Hot Pot), and another POI is type 1010000 Food Service - Chinese Food, with only Level 2 detail, then although they are similar at Level 2, they cannot be judged to be similar at the sub-category level, and the similarity score is 50; if one POI is type 1011701 Food Service - Chinese Food - Hot Pot Restaurant - Sichuan Hot Pot, with the highest level of detail (Level 4, Level 4, Level 5, Level 6, Level 7, Level 8, Level 9, Level 1011701 Level 1011701 Level 1011701 Food Service - Chinese Food - Hot Pot Restaurant - Sichuan Hot Pot, with the highest level of detail (Level 1, Level 2, Level 3, Level 4, Level 1011701 ... If a POI of type 1011701 Food Service - Chinese Food - Hot Pot Restaurant - Sichuan Hot Pot, and its level of detail is the highest level (Level 4), then its similarity score is 100. If one POI of type 1011701 Food Service - Chinese Food - Hot Pot Restaurant - Sichuan Hot Pot, and its level of detail is the highest level (Level 4), and another POI of type 1020307 Food Service - Foreign Food - Western Food - British Food, then they are similar in the first level category, but dissimilar in the second level category and below, and thus receive a very low similarity score. If the two POIs are different from the first level, then their similarity score is 0.
[0074] The method for determining the telephone similarity score is as follows: the telephone attributes of POIs in the source database and the parent database are compared based on manual rules to obtain the telephone number similarity score.
[0075] S140. Based on the deduplication result and the quality assessment result, perform data fusion on the source data and the parent database data.
[0076] Based on the above embodiments, optionally, the quality assessment result includes a quality category; the data fusion of the source data and the parent database data based on the deduplication result and the quality assessment result includes: if the deduplication result is duplicated, then the quality category of each attribute field of the source data and the parent database data is compared one by one; if the quality category of the attribute field of the source data is higher than the quality category of the corresponding attribute field of the parent database data, then the attribute field of the source data is used to replace the corresponding attribute field of the parent database data; if the deduplication result is non-duplicated, then it is determined whether the quality category of the source data meets the data addition conditions; if the data addition conditions are met, then the source data is added to the parent database.
[0077] In this embodiment, if the deduplication result is duplicated, it indicates that there is parent database data in the parent database that duplicates the points of interest of the source data. The source data and parent database data with duplicate points of interest are then fused. Specifically, the data fusion method is as follows: the quality assessment results of each attribute field in the source data and parent database data are compared one by one. If the quality category of the attribute field in the source data is higher than the quality category of the corresponding attribute field in the parent database data, then the attribute field in the source data replaces the corresponding attribute field in the parent database data. This embodiment achieves data fusion of each attribute field by comparing the quality categories of the attribute fields in the source data and parent database data, and replacing the corresponding attribute field in the parent database data with the high-quality attribute field from the source data. This improves the quality of the parent database data while retaining the high-quality attribute fields in the original parent database data.
[0078] In some embodiments, due to the poor quality of the parent database data, after deduplication of the source data and the parent database data, it is found that there are multiple parent database data that overlap with the points of interest in the source data. Data fusion can then be performed on the source data and the multiple parent database data. Specifically, one parent database data is selected as the master parent database data, and the source data and other parent database data are sequentially fused with the master parent database data. It should be noted that the data fusion between two POI data points follows the same method as described above. When fusion between parent database data points, non-master parent database data is considered as source data, which will not be elaborated further here.
[0079] In this embodiment, if the deduplication result is "no repetition," it indicates that there is no data in the parent database that duplicates the points of interest of the source data. The system then determines whether the quality category of the source data meets the data addition criteria. If the criteria are met, the source data is added to the parent database. The data addition criteria can be that the quality category of the source data's quality assessment result is "high." In other words, for source data that is not present in the parent database and has an overall quality assessment result indicating a "high" quality category, the source data is directly added to the parent database to supplement it.
[0080] Based on the above embodiments, optionally, the method further includes: when the deduplication result is duplicate, if the address or name of the source data is different from the address or name of the parent database data, then the address or name not stored in the parent database is stored in the alias database.
[0081] In this embodiment, when the deduplication result indicates duplicates, data fusion can be performed on the source data and the parent database data. However, when the quality category of the address or name in the source data is lower than that of the address or name in the parent database data, data fusion will not be performed on the address or name. Instead, such addresses or names can be stored as aliases for names and addresses in the parent database data in an alias library. The alias library can be understood as a subsidiary library of the parent database, used to store address aliases and name aliases for the parent database data.
[0082] The technical solution of this embodiment involves acquiring source data from at least one source database and parent database data from a parent database; performing quality assessment on the source data and parent database data to obtain quality assessment results; performing deduplication processing on any source data and any parent database data to obtain deduplication results; and fusing the source data and parent database data based on the deduplication results and quality assessment results. This approach increases the consideration of the quality of source data and parent database data, avoids the impact of low-quality data on the fused data, and improves the data quality of the fused parent database data.
[0083] Example 2
[0084] Figure 3 This is a schematic diagram of the structure of a multi-source data fusion device provided in Embodiment 2 of the present invention. Figure 3 As shown, the device includes:
[0085] Data acquisition module 310 is used to acquire source data from at least one source database and parent database data from at least one parent database;
[0086] The quality assessment module 320 is used to assess the quality of the source data and the parent database data to obtain the quality assessment result.
[0087] The deduplication module 330 is used to perform deduplication processing on any of the source data and the parent database data to obtain a deduplication result.
[0088] The data fusion module 340 is used to perform data fusion on the source data and the parent database data based on the deduplication result and the quality assessment result.
[0089] The technical solution of this embodiment involves acquiring source data from at least one source database and parent database data from a parent database; performing quality assessment on the source data and parent database data to obtain quality assessment results; performing deduplication processing on any source data and any parent database data to obtain deduplication results; and fusing the source data and parent database data based on the deduplication results and quality assessment results. This approach increases the consideration of the quality of source data and parent database data, avoids the impact of low-quality data on the fused data, and improves the data quality of the fused parent database data.
[0090] Based on the above embodiments, optionally, both the source data and the parent database data are point-of-interest (POI) data, and the POI data includes multiple attribute fields; the attribute fields include one or more of the following: name, name alias, address, address alias, spatial location, type, and telephone number.
[0091] Based on the above embodiments, optionally, the quality assessment module 320 is specifically used to perform quality scoring on each attribute field of the point of interest data to obtain the quality score of each attribute field; input the quality score into the quality classification model for quality assessment to obtain the quality assessment result; wherein, the quality assessment result includes the overall quality assessment result and the quality assessment result of the attribute fields.
[0092] Based on the above embodiments, optionally, the deduplication module 330 is specifically used to extract the attribute features of each attribute field in the source data and the parent database data; for any source data, a similarity score is determined between the source data and each attribute field of each parent database data based on the attribute features of the source data and the attribute fields of each parent database data; a feature engineering is generated based on the similarity score, and the feature engineering is input into the deduplication classification model for deduplication classification to obtain a deduplication result; wherein, the deduplication result is a binary classification result, and the deduplication result includes duplicate and non-duplicate.
[0093] Based on the above embodiments, optionally, the quality scoring result includes a quality category; the data fusion module 340 is specifically used to compare the quality categories of each attribute field of the source data and the parent database data one by one if the deduplication result is duplicate; if the quality category of the attribute field of the source data is higher than the quality category of the corresponding attribute field of the parent database data, the attribute field in the source data replaces the corresponding attribute field in the parent database data; if the deduplication result is non-duplicate, it is determined whether the quality category of the source data meets the data addition conditions; if the data addition conditions are met, the source data is added to the parent database.
[0094] Optionally, based on the above embodiments, the device further includes an alias management module, which is used to store the address or name not stored in the parent database in the alias database if the address or name of the source data is different from the address or name of the parent database data when the deduplication result is duplicate.
[0095] Optionally, based on the above embodiments, the device further includes a source data filtering module, used to filter the source data according to the quality assessment results after obtaining the quality assessment results.
[0096] The multi-source data fusion apparatus provided in the embodiments of the present invention can execute the multi-source data fusion method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0097] Example 3
[0098] Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0099] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0100] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0101] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as multi-source data fusion methods.
[0102] In some embodiments, the multi-source data fusion method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the multi-source data fusion method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the multi-source data fusion method by any other suitable means (e.g., by means of firmware).
[0103] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0104] Computer programs for implementing the multi-source data fusion method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0105] Example 4
[0106] Embodiment 4 of the present invention also provides a computer-readable storage medium storing computer instructions for causing a processor to execute a multi-source data fusion method, the method comprising:
[0107] Obtain source data from at least one source database and parent database data from a parent database; perform quality assessment on the source data and the parent database data to obtain a quality assessment result; for any source data, perform deduplication processing on the source data and the parent database data to obtain a deduplication result; perform data fusion on the source data and the parent database data based on the deduplication result and the quality assessment result.
[0108] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0109] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0110] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0111] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0112] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0113] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A multi-source data fusion method, characterized in that, include: Retrieve source data from at least one source database and parent database data from at least one parent database; The source data and the parent database data are subjected to quality assessment to obtain quality assessment results; For any of the source data, perform deduplication processing on the source data and the parent database data to obtain the deduplication result; Based on the deduplication results and the quality assessment results, the source data and the parent database data are fused together. Both the source data and the parent database data are point-of-interest (POI) data, which include multiple attribute fields. The attribute fields include one or more of the following: name, name alias, address, address alias, spatial location, type, and phone number. The quality assessment of the source data and the parent database data to obtain the quality assessment results includes: For any of the aforementioned point of interest data, a quality score is assigned to each attribute field of the point of interest data to obtain the quality score of each attribute field. The quality scores of each attribute field of the point of interest data are input into a quality classification model for quality assessment to obtain the quality assessment results of the point of interest data; wherein, the quality assessment results include the overall quality assessment results and the quality assessment results of the attribute fields; The quality assessment results include quality categories; the data fusion of the source data and the parent database data based on the deduplication results and the quality assessment results includes: If the deduplication result is duplicate, the quality category of each attribute field of the source data and the parent database data is compared one by one. If the quality category of the attribute field of the source data is higher than the quality category of the corresponding attribute field of the parent database data, the attribute field of the source data is used to replace the corresponding attribute field of the parent database data. If the deduplication result is non-repetition, then determine whether the quality category of the source data meets the data addition conditions. If the data addition conditions are met, then add the source data to the parent database. Among them, attribute features refer to the features used to evaluate the similarity score of any attribute field in the source data and the parent database data; Specifically, for any of the aforementioned point-of-interest (POI) data, a quality score is assigned to each attribute field of the POI data, resulting in a quality score for each attribute field, including: The quality of the point of interest data is assessed based on the address field, the name field, the type field, and the spatial location field. The quality assessment of the point of interest data based on the address field includes: The quality of the point of interest data is scored based on address completeness, address accuracy, and address precision. The quality assessment of the point of interest data based on the name field includes: The name field of the point of interest data is scored according to the name resolution model of the point of interest data to obtain the quality score of the name field of the point of interest data. The quality assessment of the point of interest data based on the type field includes: The completeness of the interest point data type fields is scored according to the classification table to obtain the quality score of the interest point data type. The quality assessment of the point of interest data based on the spatial location field includes: Based on the building extent data, it is determined whether the point of interest data contains the building extent data, and the distance between the point of interest data and the nearest building extent data is calculated. The two indicators of inclusion and distance are weighted to obtain the coordinate quality score of the point of interest data. The quality scoring of the point of interest data based on address completeness includes: Geocoding tools are used to perform address segmentation and semantic recognition on the address field to obtain semantic fragments, and then address accuracy scores are determined based on the semantic fragments. The quality scoring of the point of interest data based on address accuracy includes: Spatial parsing is performed on the address field to calculate the spatial latitude and longitude corresponding to the address. Then, based on the spatial distance algorithm, the spatial distance between the spatial latitude and longitude and the latitude and longitude of the point of interest data is calculated to determine the consistency between the address and the spatial location. Finally, the address accuracy is scored based on the degree of consistency between the address and the spatial location. The quality scoring of the point of interest data based on address accuracy includes: The accuracy of the address is determined based on the administrative division of the point of interest data. If the address contradicts the administrative division, the point of interest data is deleted. The accuracy score for determining the address based on semantic fragments includes: The address accuracy is scored based on the address accuracy scoring table to obtain a basic accuracy score; We analyzed the last item of the address segmentation and made adjustments based on the basic accuracy score.
2. The multi-source data fusion method according to claim 1, characterized in that, For any of the source data, the process of performing deduplication processing between the source data and the parent database data to obtain a deduplication result includes: Extract the attribute features of each attribute field from the source data and the parent database data; For any of the source data, a similarity score between the source data and each attribute field of each parent database is determined based on the attribute features of the source data and the attribute fields of each parent database. Feature engineering is generated based on the similarity score, and the feature engineering is input into the deduplication classification model for deduplication classification to obtain the deduplication result; wherein, the deduplication result is a binary classification result, and the deduplication result includes duplicate and non-duplicate.
3. The multi-source data fusion method according to claim 1, characterized in that, The method further includes: When the deduplication result is duplicate, if the address or name of the source data is different from the address or name of the parent database data, then the address or name not stored in the parent database is stored in the alias database.
4. The multi-source data fusion method according to claim 1, characterized in that, After obtaining the quality assessment results, the method further includes: The source data is filtered based on the quality assessment results.
5. A multi-source data fusion device, characterized in that, include: A data acquisition model is used to acquire source data from at least one source database and parent database data from a parent database. The quality assessment module is used to assess the quality of the source data and the parent database data, and obtain the quality assessment results. The deduplication module is used to perform deduplication processing on any of the source data and any of the parent database data to obtain the deduplication result. The data fusion module is used to perform data fusion on the source data and the parent database data based on the deduplication result and the quality assessment result; Both the source data and the parent database data are point-of-interest (POI) data, which include multiple attribute fields. The attribute fields include one or more of the following: name, name alias, address, address alias, spatial location, type, and phone number. Specifically, the quality assessment module is used to: score the quality of each attribute field of the point of interest data to obtain the quality score of each attribute field; input the quality score into the quality classification model for quality assessment to obtain the quality assessment result; wherein the quality assessment result includes the overall quality assessment result and the quality assessment result of the attribute fields; The quality assessment results include quality categories; Specifically, the data fusion module is used to: if the deduplication result is duplicate, compare the quality category of each attribute field of the source data and the parent database data one by one; if the quality category of the attribute field of the source data is higher than the quality category of the corresponding attribute field of the parent database data, replace the corresponding attribute field of the parent database data with the attribute field of the source data; if the deduplication result is non-duplicate, determine whether the quality category of the source data meets the data addition conditions; if the data addition conditions are met, add the source data to the parent database. Among them, attribute features refer to the features used to evaluate the similarity score of any attribute field in the source data and the parent database data; The quality assessment module is specifically used for: The quality of the point of interest data is assessed based on the address field, the name field, the type field, and the spatial location field. The quality assessment of the point of interest data based on the address field includes: The quality of the point of interest data is scored based on address completeness, address accuracy, and address precision. The quality assessment of the point of interest data based on the name field includes: The name field of the point of interest data is scored according to the name resolution model of the point of interest data to obtain the quality score of the name field of the point of interest data. The quality assessment of the point of interest data based on the type field includes: The completeness of the interest point data type fields is scored according to the classification table to obtain the quality score of the interest point data type. The quality assessment of the point of interest data based on the spatial location field includes: Based on the building extent data, it is determined whether the point of interest data contains the building extent data, and the distance between the point of interest data and the nearest building extent data is calculated. The two indicators of inclusion and distance are weighted to obtain the coordinate quality score of the point of interest data. The quality scoring of the point of interest data based on address completeness includes: Geocoding tools are used to perform address segmentation and semantic recognition on the address field to obtain semantic fragments, and then address accuracy scores are determined based on the semantic fragments. The quality scoring of the point of interest data based on address accuracy includes: Spatial parsing is performed on the address field to calculate the spatial latitude and longitude corresponding to the address. Then, based on the spatial distance algorithm, the spatial distance between the spatial latitude and longitude and the latitude and longitude of the point of interest data is calculated to determine the consistency between the address and the spatial location. Finally, the address accuracy is scored based on the degree of consistency between the address and the spatial location. The quality scoring of the point of interest data based on address accuracy includes: The accuracy of the address is determined based on the administrative division of the point of interest data. If the address contradicts the administrative division, the point of interest data is deleted. The accuracy score for determining the address based on semantic fragments includes: The address accuracy is scored based on the address accuracy scoring table to obtain a basic accuracy score; We analyzed the last item of the address segmentation and made adjustments based on the basic accuracy score.
6. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the multi-source data fusion method according to any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the multi-source data fusion method according to any one of claims 1-4.
Citation Information
Patent Citations
Multi-source POI fusion method and device based on NLP technology and readable storage medium
CN114201480A
De-duplication fusion method based on multi-source POI data and related device
CN115017426A