A multi-source data fusion method based on AI model technology

Through the multi-source data fusion method based on AI model, different identity information databases are processed in a targeted manner, inconsistent information is identified and corrected, and unique information association links are formed, which solves the problems of unique information lack and processing complexity in data integration, and realizes efficient and secure data processing and analysis.

CN119441542BActive Publication Date: 2025-08-22SUZHOU SHENLE TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411541166.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-08-22
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

The prior art fails to effectively integrate unique information from different data sources in data fusion, resulting in insufficient data integrity and utilization, and the unified incremental update processing method does not apply to the characteristics of different data sources, which increases processing complexity and overhead.

Method used

Through AI model technology, multiple types of identity information databases are retrieved separately, data performance characteristics are extracted for adaptation preprocessing, inconsistent information is identified and corrected, usage priority is determined, data updates and formats are unified, and unique information association links are formed.

Benefits of technology

It realizes targeted and flexible data processing, improves processing efficiency, reduces complexity and cost, ensures data consistency and quality, and enhances the relevance and application value of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119441542B_ABST
    Figure CN119441542B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of data fusion technology, and specifically relates to a multi-source data fusion method based on AI model technology. By extracting data performance characteristics of data in various identity information databases before data fusion, an adaptive preprocessing method is selected to achieve targeted and flexible preprocessing of data in different data sources, while improving processing efficiency and reducing the difficulty and cost of the processing process. In the data fusion process, the data is first classified into common information and unique information, thereby identifying and correcting inconsistent information items in the common information, and correlating and integrating unique information with common information to form a unique information correlation link, thereby achieving classified integration of information in different data sources. It can not only ensure the consistency of common information in different data sources, but also strengthen the correlation of information in different data sources, making the application value of the fused data higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data fusion technology, and specifically relates to a multi-source data fusion method based on AI model technology. Background Art

[0002] With the advent of the big data era, data plays an increasingly important role in decision-making and execution. However, data from a single field often cannot provide sufficient information to support decision-making. Therefore, integrating data from multiple fields can provide more comprehensive and accurate decision-making.

[0003] There are also solutions to deal with data fusion in the existing technology. For example, a cross-domain data fusion method disclosed in the Chinese invention patent application with publication number CN116662371A is a method for integrating and managing data from different data sources after parsing and standardizing them to generate a new data set. The new data set is then processed and analyzed and presented to the user. This invention aims to solve the problems of data consistency and conflict when integrating and managing data from different data sources. This integration is mainly aimed at the common information in different data sources. Since different data sources often come from different data collection sources, there may be data inconsistencies on the same information, which requires resolving inconsistencies during integration. However, this invention ignores the fact that there is not only common information but also unique information between data from different data sources. The lack of targeted integration of unique information in each data source may result in the lack of certain data source-specific information in the integrated data, thereby affecting the integrity of the overall data. It is also easy to fail to fully utilize available data resources and affect the effectiveness of data management.

[0004] Furthermore, preprocessing data from different data sources before fusion is paramount. This is done to improve data quality by cleaning the data and addressing outliers. However, existing data fusion processes typically employ a unified incremental update approach to improve processing efficiency. This fixed approach fails to account for the differences in data characteristics across different data sources. For data sources with smaller volumes or infrequent updates, incremental updates may be impractical. This approach not only fails to improve processing efficiency but also increases processing complexity and overhead. Consequently, this unified incremental update approach has limitations and cannot adapt to the data characteristics of different data sources. It can also pose data inconsistency risks. Summary of the Invention

[0005] To this end, the present invention provides a multi-source data fusion method based on AI model technology for multi-source identity information data fusion for various identity information databases generated by social operations, which effectively solves the problems mentioned in the background technology.

[0006] The purpose of the present invention can be achieved through the following technical solutions: A multi-source data fusion method based on AI model technology, comprising the following steps: S1, respectively calling up a first-class identity information database, a second-class identity information database and a third-class identity information database, and extracting data performance characteristics in the database, wherein the data performance characteristics include data occupied space, data storage format and data type, thereby selecting an adaptive preprocessing method for preprocessing.

[0007] S2. Determine identity information as an identifier, and then extract the identity data corresponding to the identifier and the person information related to the identity data from each database. Thus, match the identity data extracted from the first-category identity information database, the second-category identity information database, and the third-category identity information database, and then retrieve the related information of the same person in different databases, and classify the related information into common information and unique information.

[0008] S3. Compare and map the common information of the same person in different databases to identify whether there is inconsistency. If there is inconsistency, output the inconsistent information items.

[0009] S4. Analyze the data update timeliness and data quality of the first-category identity information database, the second-category identity information database, and the third-category identity information database, respectively, and thereby determine the usage priority of the first-category identity information database, the second-category identity information database, and the third-category identity information database.

[0010] S5. Correct the information items with inconsistent output based on the usage priorities of the first-category identity information database, the second-category identity information database, and the third-category identity information database.

[0011] S6. Standardize the presentation format of the common information revised by the same person.

[0012] S7. Associate and integrate the unique information of the same person in different databases to form a unique information association link.

[0013] S8. Link the common information and unique information of the same person in different databases that have been corrected and presented in a unified format to form a data set.

[0014] Combining all the above technical solutions, the positive effects of the present invention are as follows: 1. The present invention extracts data representation characteristics from data in various identity information databases, thereby selecting adaptive preprocessing methods for data in different databases, thereby achieving targeted and flexible preprocessing of data in different data sources, with better applicability. While improving processing efficiency, it reduces the difficulty and cost of the processing process, and at the same time minimizes the data consistency risks caused by preprocessing, which helps to achieve more accurate, efficient and secure data processing and analysis.

[0015] 2. In the process of fusing data from various identity information databases, the present invention first classifies the data into common information and unique information, thereby identifying and correcting inconsistent information items in the common information, and associating and integrating the unique information with the common information to form a unique information association link, thereby realizing the classified integration of information from different data sources. This not only ensures the consistency of common information in different data sources and ensures the stability of data quality, but also strengthens the correlation of information in different data sources based on the unique information association link, making the application value of the fused data higher, helping to form a more comprehensive and integrated information perspective, and providing more accurate information support for decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The present invention is further described with reference to the accompanying drawings. However, the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative effort.

[0017] Figure 1 The present invention is a flowchart of the steps for implementing the method. DETAILED DESCRIPTION

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0019] See also Figure 1 As shown, the present invention proposes a multi-source data fusion method based on AI model technology, comprising the following steps: S1, respectively retrieving a first-class identity information database, a second-class identity information database and a third-class identity information database, wherein different databases correspond to one data source, and each identity information database is established for different data sources, and extracting data performance characteristics from the database, wherein the data performance characteristics include data occupied space, data storage format and data type, thereby selecting an adaptive preprocessing method for preprocessing.

[0020] Applied to the above embodiment, data occupied space refers to the physical space or storage capacity required to store data, and data storage format refers to organizing data in a specific structure and specification to facilitate storage, retrieval and processing in a computer system, including but not limited to text files, Excel files, XML, binary formats, etc.

[0021] Further applied to the above embodiment, data types include dynamic data and static data. In one example, the data type can be identified by determining the nature of the data based on its source. If the data source is data from a sensor, log record, or other source, the data is often dynamic data because it records real-time events and changes. If the data source is data from a fixed report, database snapshot, or other source, the data is likely static data because it typically represents a snapshot at a fixed point in time.

[0022] In another example, data tags can be used to identify whether the data is dynamic or static. For example, a data source may provide a timestamp to indicate when the data was generated. If the timestamp changes frequently, it indicates that the data is dynamic.

[0023] As another example, the data structure can be used to help determine the nature of the data. For example, if the data contains time series or real-time monitoring data, it may be dynamic data. If the data contains a fixed list, it may be static data.

[0024] In the specific implementation of data type recognition, combined judgment can be performed based on the combination of the above examples to improve the accuracy of recognition.

[0025] It should be noted that the reason for using data space occupation, data storage format and data type as the data representation characteristics in the data source is that these factors directly affect the data processing method, efficiency and demand for system resources. By analyzing and understanding these characteristics, we can better select and optimize data processing methods and improve data processing efficiency and quality.

[0026] As a preferred implementation of the above scheme, the adaptive preprocessing method is selected according to the following process: the data occupied space is extracted from the data performance characteristics of the first-class identity information database, the second-class identity information database and the third-class identity information database, and the data occupied space is subtracted from the set occupied space threshold, and then the difference result is divided by the occupied space threshold to obtain the data scale corresponding to the first-class identity information database, the second-class identity information database and the third-class identity information database.

[0027] The occupied space threshold mentioned above is an initial setting, and the purpose of setting it is to assist in the quantitative calculation of data scale.

[0028] The storage format and data type of each data are extracted from the data presentation characteristics of the first-category identity information database, the second-category identity information database, and the third-category identity information database, where the data types include dynamic data and static data. The extracted storage formats and data types are then deduplicated, and the number of storage formats and data types after deduplication is summarized. At the same time, data with the same storage format are classified, and data with the same data type are classified. The amount of data classified for each storage format and the amount of data classified for each data type are counted.

[0029] The data distribution diversity is calculated based on the number of storage formats after deduplication, the number of data types, the amount of data classified by each storage format, and the amount of data classified by each data type. The calculation formula is: , where 、 Respectively represent the number of storage formats and data types after deduplication, Indicates the total amount of data in the database. 、 Respectively represent the maximum amount of data and the minimum amount of data in each storage format classification. 、 They respectively represent the maximum and minimum amounts of data in each data type classification. The fewer duplicate storage formats and data types in the data source, the more diverse the data distribution, which to a certain extent represents the complexity of the data.

[0030] The data scale and data distribution diversity corresponding to the first-class identity information database, the second-class identity information database and the third-class identity information database are weighted averaged. For example, the weight values ​​corresponding to the data scale and the data distribution diversity in the weighted average calculation are 0.4 and 0.6, and the special preprocessing requirement index corresponding to the database is obtained, and it is compared with the set limit value. If the special preprocessing requirement index corresponding to a database is less than the set limit value, ordinary processing is selected as the adaptive preprocessing method corresponding to the database, where ordinary preprocessing refers to a common and general data preprocessing method. Otherwise, incremental processing is selected as the adaptive preprocessing method corresponding to the database.

[0031] It's important to understand that the larger the data size and the more diverse the data distribution, the more specialized processing is needed. This is because large-scale data often comes with more data quality issues, such as missing values, outliers, and duplicate data. Diverse data distributions can lead to more complex data quality issues. Specialized preprocessing techniques can help identify and address these issues, ensuring that data quality meets analysis and modeling requirements. Furthermore, processing large-scale data requires more computing and storage resources, and standard processing may limit efficiency. Specialized preprocessing techniques can help reduce data dimensionality, perform feature selection, and reduce noise, thereby reducing data complexity and improving processing efficiency. Conversely, for smaller data with a more uniform distribution, standard preprocessing methods often meet requirements while offering higher processing efficiency and avoiding the need for complex specialized processing. However, specialized preprocessing methods may overfit to specific data distributions or problem characteristics, leading to poor performance on other datasets. For smaller data with a more uniform distribution, standard preprocessing methods are more likely to avoid overfitting and have better generalization capabilities.

[0032] It's also important to understand that choosing incremental processing during special preprocessing can enable real-time data processing. Incremental processing allows for timely processing of newly arrived data, enabling the system to respond promptly. However, for large datasets, loading all data at once for processing can consume significant computing resources and memory space. Incremental processing processes data in batches, reducing computing resource demands and memory usage, thereby utilizing resources more efficiently. Furthermore, for dynamic data in data sources, which may be generated in real time, incremental processing can reduce data processing latency. Incremental processing processes data immediately after it is generated, without having to wait for all data to be collected. Finally, incremental processing can often be designed to be fault-tolerant and resilient. Even if errors or interruptions occur during processing, the system can restart the incremental processing process to resume processing where it left off, ensuring data is not lost.

[0033] It should be added that data preprocessing includes operations such as data cleaning and data noise reduction. After selecting the appropriate data preprocessing method for each database, the data in each database is imported into the intelligent model of the corresponding preprocessing method for preprocessing to improve data processing efficiency and accuracy.

[0034] S2. Determine identity information as an identifier, and then extract the identity data corresponding to the identifier and the person information related to the identity data from each database. Thus, match the identity data extracted from the first-category identity information database, the second-category identity information database, and the third-category identity information database, and then retrieve the related information of the same person in different databases, and classify the related information into common information and unique information.

[0035] It's important to note that identity information is used as an identifier because it's typically unique and unchangeable. By using identity information as a correlation identifier, we ensure that data across different databases can be accurately matched and linked, allowing us to extract all relevant information about the same person from different data sources based on their identity information.

[0036] Preferably, the classification of the relevant information into common information and unique information is specifically implemented as follows: the relevant information of the same person in different databases is compared. If there is some relevant information that only exists in one database, the relevant information is classified as unique information; otherwise, the relevant information is classified as common information.

[0037] It should be understood that although various identity information databases are established from different data sources, since they involve personal identity, there will be overlaps, thereby generating common information. At the same time, various identity information databases will involve different fields due to differences in data sources, and therefore will have unique information that is different from other databases.

[0038] S3. Compare and map the common information of the same person in different databases to identify whether there is inconsistency. If there is inconsistency, output the inconsistent information items.

[0039] Specifically, in the above scheme, the following process is used to identify whether there is inconsistency: each piece of common information of the same person in different databases is compared. If a certain piece of information is the same in different databases, the information is identified as consistent; otherwise, the information is identified as inconsistent.

[0040] S4. Analyze the data update timeliness and data quality of the first-category identity information database, the second-category identity information database, and the third-category identity information database, respectively, and thereby determine the usage priority of the first-category identity information database, the second-category identity information database, and the third-category identity information database.

[0041] Optimally, the data update timeliness analysis is as follows: retrieve the historical update records corresponding to the first-category identity information database, the second-category identity information database, and the third-category identity information database respectively, and extract the updated data volume, update time, and the total data volume existing in the database at the time of update from the records.

[0042] The updated data volume in each historical update record is compared with the total data volume in the database at the time of update to obtain the data update ratio value corresponding to each historical update record, and the average is calculated to obtain the average data update ratio value.

[0043] Compare the update time in each historical update record with adjacent records to obtain the average update interval duration.

[0044] The average data update ratio value is combined with the average update interval length using the formula Get data update timeliness , Indicates the average update interval. Indicates the time between the update time of the first record and the update time of the last record in the historical update record. Represents the average data update ratio value. The larger the data update ratio value, the shorter the average update interval and the greater the data update timeliness.

[0045] It should be understood that the reason why the data update ratio and update interval duration are used as analysis indicators when analyzing the data update timeliness of a data source is that the data update ratio can intuitively reflect the amount of data change brought about by each data update, thereby evaluating the freshness of the data, and the update interval duration can tell the frequency of data updates, further helping to evaluate the timeliness of the data. By combining these two, the timeliness of the data source can be comprehensively evaluated.

[0046] For further optimization, the data quality analysis refers to the following process: performing missing data detection on the data in the first-category identity information database, the second-category identity information database, and the third-category identity information database, and calculating the proportion of missing data.

[0047] In specific optimization implementations, missing values ​​detection can be visualized using data visualization tools such as scatter plots, box plots, and heat maps to visually identify missing values ​​in the data. Missing values ​​are usually marked with specific symbols or colors to make them easy to identify in visual charts.

[0048] As another example, during the data preprocessing stage, missing values ​​may be marked as specific placeholders, thereby identifying missing values ​​by checking whether these specific placeholders exist in the data.

[0049] The issuing agencies corresponding to each piece of data in the first-category identity information database, the second-category identity information database, and the third-category identity information database are obtained respectively, where the issuing agencies include government-related agencies and third-party agencies. The data issued by the same issuing agency are classified accordingly, and the proportion of data release corresponding to government-related agencies is calculated. The larger the proportion of data release corresponding to government-related agencies, the higher the authority and reliability of the data. This is because government-related agencies are usually the main producers and managers of data, and have authoritative data sources and data management mechanisms.

[0050] It's important to understand that the reason third-party organizations publish data in Category 1, Category 2, and Category 3 identity information databases is because government departments may lack the specialized technology and resources to effectively manage and publish large amounts of data. Therefore, the government may entrust third-party organizations with data publishing.

[0051] Import the missing data percentage values ​​corresponding to the first-class identity information database, the second-class identity information database, and the third-class identity information database and the data release percentage values ​​corresponding to the government-related agencies into the data quality analysis formula Get the data quality coefficients corresponding to the first-class identity information database, the second-class identity information database, and the third-class identity information database , where 、 They respectively represent the proportion of missing data and the proportion of data released by government-related agencies.

[0052] Furthermore, the usage priority of the first-class identity information database, the second-class identity information database and the third-class identity information database is determined by referring to the following process: the data update timeliness and data quality coefficient corresponding to the first-class identity information database, the second-class identity information database and the third-class identity information database are weighted averaged to obtain the usage priority. For example, the weight values ​​of the data update timeliness and the data quality coefficient in the weighted average can be 0.3 and 0.7.

[0053] Arrange the first, second, and third identity information databases in descending order of usage priority to obtain the usage priorities of the first, second, and third identity information databases.

[0054] S5. Based on the usage priority of the first-class identity information database, the second-class identity information database, and the third-class identity information database, the output inconsistent information items are corrected. The specific correction process is: (1) extracting the values ​​of the inconsistent information items in the first-class identity information database, the second-class identity information database, and the third-class identity information database. If the values ​​of the inconsistent information items in the three databases are all different, execute (2); otherwise, execute (3).

[0055] (2) The information item is modified using the value in the database corresponding to the highest usage priority, where the database corresponding to the highest usage priority is the database that ranks first in the descending order of usage priority.

[0056] (3) Obtain the usage priority of the database with the same value for the inconsistent information item and compare it with the highest usage priority. If there is a highest usage priority, the information item is corrected with the value. If there is no highest usage priority, the information item is corrected with the value in the database corresponding to the highest usage priority.

[0057] When correcting inconsistent information items in the first, second and third category identity information databases, the present invention analyzes the usage priorities of different databases and makes corrections accordingly. On the one hand, it can ensure that information from high-quality data sources has priority processing and correction, thereby improving the credibility of information correction. On the other hand, it can improve the efficiency of information correction, avoid the waste of time and resources caused by blind correction, and at the same time prevent the spread of erroneous information, ensure that erroneous information will not further affect data quality, and protect the overall data quality level.

[0058] S6. The common information of the same person after revision is presented in a unified format. The specific implementation is as follows: the common information of the same person in different databases is classified into revised common information and unrevised common information according to whether it is revised or not.

[0059] For the modified common information, the presentation format of the information in each database and the database adopted for modification are obtained, and then the presentation information of the information is unified using the presentation format in the modified adopted database as the target presentation information.

[0060] For the uncorrected common information, the presentation format of the information in each database is obtained, and then the presentation information of the information is unified by using the presentation format in the database corresponding to the highest usage priority as the target presentation information.

[0061] It's important to understand that when fusing multi-source data from Category 1, Category 2, and Category 3 identity information databases, common information may be presented in different formats because different data sources may employ different data collection methods, resulting in variations in the representation of the same information. For example, some data sources may collect data through manual forms or questionnaires, while others may use automated systems or sensors. These varying collection methods may result in different data formats, making it crucial to standardize the format of this information to ensure data consistency and accuracy.

[0062] S7. Associate and integrate the unique information of the same person in different databases to form a unique information association link. For details, see the following process: formulate association principles, and associate and sort the first-class identity information database, the second-class identity information database, and the third-class identity information database. Then, based on the sorting results, select the unique information in the database ranked first and the unique information in other databases according to the association principles to obtain the unique information in other databases associated with each unique information in each database.

[0063] It should be understood that the principles for associating unique data in the above-mentioned various identity information databases should be formulated in combination with the characteristics of the fields in which the various identity information databases are located.

[0064] The unique information in each database is associated with the unique information in other databases to form a unique information association link.

[0065] Exemplarily, the association order of the first-category identity information database, the second-category identity information database, and the third-category identity information database is set as follows: first, the unique information in the first-category identity information database is associated with the unique information corresponding to the second-category identity information database and the third-category identity information database, respectively; secondly, the unique information in the second-category identity information database is associated with the unique information corresponding to the first-category identity information database and the third-category identity information database, respectively; and finally, the unique information in the third-category identity information database is associated with the unique information corresponding to the first-category identity information database and the second-category identity information database, respectively.

[0066] In the above example, assume that the first-class identity information database is A, and the unique information in the first-class identity information database is a1, a2, and a3 respectively; the second-class identity information database is B, and the unique information in the second-class identity information database is b1 and b2 respectively; the third-class identity information database is C, and the unique information in the third-class identity information database is c1 and c2 respectively. According to the association principle, the unique information in each database is associated with the unique information in other databases, which are a1→b1, a2→c1, a3→c2, b1→c2, b2→a2, c1→b1, and c2→a2 respectively. At this time, the unique information association links are a1→b1→c2, a2→c1→b1, and b1→c2→a2.

[0067] S8. Link the common information and unique information of the same person in different databases that have been corrected and presented in a unified format to form a data set.

[0068] In the process of fusing data in various identity information databases, the present invention first classifies the data into common information and unique information, thereby identifying and correcting inconsistent information items in the common information, and associating and integrating the unique information with the common information to form a unique information association link, thereby realizing the classified integration of information in different data sources. It can not only ensure the consistency of common information in different data sources and ensure the stability of data quality, but also strengthen the correlation of information in different data sources based on the unique information association link, so that the application value of the fused data is higher, which helps to form a more comprehensive and integrated information perspective and provide more accurate information support for decision-making.

[0069] The above content is merely an example and explanation of the structure of the present invention. Those skilled in the art may make various modifications or additions to the described specific embodiments or replace them in a similar manner. As long as they do not deviate from the structure of the invention or exceed the scope defined by the present invention, they should all fall within the scope of protection of the present invention.

Claims

1. A multi-source data fusion method based on AI model technology, characterized in that: The following steps are involved: S1. Retrieve the first-category identity information database, the second-category identity information database, and the third-category identity information database respectively, and extract data performance characteristics from the databases, where the data performance characteristics include data occupied space, data storage format, and data type, and select an adaptive preprocessing method for preprocessing; S2. Determine the identity information as an identifier, and then extract the identity data corresponding to the identifier and the person information related to the identity data from each database. Match the identity data extracted from the first, second, and third identity information databases, and then retrieve the related information of the same person from different databases, and classify the related information into common information and unique information. S3. Compare and map the common information of the same person in different databases to identify whether there is inconsistency. If there is inconsistency, output the inconsistent information items; S4. Analyze the data update timeliness and data quality of the first-category identity information database, the second-category identity information database, and the third-category identity information database, respectively, and thereby determine the priority of use of the first-category identity information database, the second-category identity information database, and the third-category identity information database; S5. Correcting the information items that are inconsistent with the output based on the usage priorities of the first-category identity information database, the second-category identity information database, and the third-category identity information database; S6. Standardize the presentation format of the common information revised by the same person; S7. Associate and integrate the unique information of the same person in different databases to form a unique information association link; S8. Link the common information and unique information of the same person in different databases that have been corrected and presented in a unified format to form a data set.

2. The multi-source data fusion method based on AI model technology according to claim 1, characterized in that: The selection of the adaptation preprocessing method refers to the following process: Extracting the data occupied space from the data performance characteristics of the first-category identity information database, the second-category identity information database, and the third-category identity information database, and subtracting it from the set occupied space threshold, and then dividing the subtraction result by the occupied space threshold to obtain the data scale corresponding to the first-category identity information database, the second-category identity information database, and the third-category identity information database; Extracting the storage format and data type of each piece of data from the data representation characteristics of the first-category identity information database, the second-category identity information database, and the third-category identity information database, where the data types include dynamic data and static data, and then deduplicating the extracted storage formats and data types, summarizing the number of storage formats and data types after deduplication, and classifying data with the same storage format and data with the same data type, and counting the amount of data classified by each storage format and data classified by each data type; The data distribution diversity is calculated based on the number of storage formats after deduplication, the number of data types, the amount of data classified by each storage format, and the amount of data classified by each data type. The calculation formula is: Where M and Y represent the number of storage formats and data types after deduplication, respectively; X represents the total amount of data in the database; Dmax and Dmin represent the maximum and minimum amounts of data in each storage format classification, respectively; Fmax and Fmin represent the maximum and minimum amounts of data in each data type classification, respectively. The data scale and data distribution diversity corresponding to the first-class identity information database, the second-class identity information database and the third-class identity information database are weighted averaged to obtain the special preprocessing requirement index corresponding to the database, and compared with the set limit value. If the special preprocessing requirement index corresponding to a database is less than the set limit value, ordinary processing is selected as the adaptive preprocessing method corresponding to the database; otherwise, incremental processing is selected as the adaptive preprocessing method corresponding to the database.

3. The multi-source data fusion method based on AI model technology according to claim 1, characterized in that: The classification of the information involved into common information and unique information is specifically implemented as follows: Compare the relevant information of the same person in different databases. If there is certain relevant information that only exists in one database, then the relevant information will be classified as unique information; otherwise, the relevant information will be classified as common information.

4. The multi-source data fusion method based on AI model technology according to claim 1, characterized in that: The process of identifying whether there is inconsistency is as follows: Compare each piece of common information of the same person in different databases. If a piece of information is the same in different databases, the information is identified as consistent. Otherwise, the information is identified as inconsistent.

5. The multi-source data fusion method based on AI model technology according to claim 1, characterized in that: The data update timeliness analysis is as follows: Retrieving the historical update records corresponding to the first-category identity information database, the second-category identity information database, and the third-category identity information database, respectively, and extracting the updated data volume, update time, and total data volume in the database at the time of update from the records; Compare the updated data volume in each historical update record with the total data volume in the database at the time of update to obtain the data update ratio value corresponding to each historical update record, and perform average calculation to obtain the average data update ratio value; Compare the update time of each historical update record with adjacent records to obtain the average update interval duration; The average data update ratio value is combined with the average update interval length using the formula The data update timeliness DT is obtained, where Δt represents the average update interval, T represents the time between the update time of the first record and the update time of the last record in the historical update record, and λ represents the average data update ratio.

6. The multi-source data fusion method based on AI model technology according to claim 5, characterized in that: The data quality analysis is as follows: Conduct missing data detection on the first-category identity information database, the second-category identity information database, and the third-category identity information database, and calculate the percentage of missing data; Obtain the publishing agency corresponding to each piece of data in the first, second, and third category identity information databases, respectively. The publishing agencies include government-affiliated agencies and third-party agencies. Classify the data published by the same publishing agency and calculate the proportion of data published by government-affiliated agencies. The missing data proportion values ​​corresponding to the first-category identity information database, the second-category identity information database and the third-category identity information database and the data release proportion values ​​corresponding to the government-affiliated agencies are imported into the data quality analysis formula to obtain the data quality coefficients corresponding to the first-category identity information database, the second-category identity information database and the third-category identity information database.

7. The multi-source data fusion method based on AI model technology according to claim 6, characterized in that: The determination of the priority of use of the first-class identity information database, the second-class identity information database, and the third-class identity information database is as follows: The use priority is obtained by performing a weighted average calculation of the data update timeliness and data quality coefficient corresponding to the first-class identity information database, the second-class identity information database, and the third-class identity information database; Arrange the first, second, and third identity information databases in descending order of usage priority to obtain the usage priorities of the first, second, and third identity information databases.

8. The multi-source data fusion method based on AI model technology according to claim 1, characterized in that: The correction of the output inconsistent information items refers to the following process: (1) extracting the values ​​of the inconsistent information items in the first-category identity information database, the second-category identity information database, and the third-category identity information database. If the values ​​of the inconsistent information items in the three databases are all different, then execute (2); otherwise, execute (3); (2) modifying the information item with the value in the database corresponding to the highest priority; (3) Obtain the usage priority of the database with the same value of the inconsistent information item and compare it with the highest usage priority. If there is a highest usage priority, the information item is corrected with the value; if there is no highest usage priority, the information item is corrected with the value in the database corresponding to the highest usage priority.

9. The multi-source data fusion method based on AI model technology according to claim 1, characterized in that: The unified presentation format of the common information corrected by the same person is implemented as follows: The common information of the same person in different databases is classified into revised common information and unrevised common information according to whether it has been revised; For the revised common information, the presentation format of the information in each database and the database adopted for the revision are obtained, and then the presentation information of the information is unified using the presentation format in the revised adopted database as the target presentation information; For the uncorrected common information, the presentation format of the information in each database is obtained, and then the presentation information of the information is unified by using the presentation format in the database corresponding to the highest usage priority as the target presentation information.

10. The multi-source data fusion method based on AI model technology according to claim 1, characterized in that: The process of associating and integrating the unique information of the same person in different databases is as follows: Establishing a correlation principle, and correlating and ranking the first-class identity information database, the second-class identity information database, and the third-class identity information database, and then selecting the unique information in the database ranked first based on the ranking result and the unique information in the other databases according to the correlation principle to obtain the unique information in the other databases that is correlated with each piece of unique information in each database; The unique information in each database is associated with the unique information in other databases to form a unique information association link.

Citation Information

Patent Citations

  • Cross-domain data fusion method

    CN116662371A

  • Multi-source heterogeneous data fusion method and system capable of being used for digital twinning

    CN117056867A

  • City data model management system and method based on multi-feature fusion

    CN118035251A