Multi-source data fusion processing method

Through the multi-source data fusion processing method, multi-source data is cleaned, preprocessed and fusion analysis, solving the problems of inconsistent data standards and low data quality in the construction of traditional data platforms, and achieving more accurate and efficient data analysis and decision-making.

CN119989275APending Publication Date: 2025-05-13山东中图软件技术有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510109340.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The lack of unified standards in the construction of traditional data platforms has led to inconsistent data standards, redundant data, and low data quality in the same business scenario. Multi-source data fusion analysis faces challenges such as data quality and consistency, heterogeneity and scale.

Method used

Multi-source data fusion processing method is adopted, and multi-source data is cleaned, preprocessed, and fusion analysis is carried out by constructing feature models 1 and feature models 2. Statistical analysis, data mining, and machine learning technologies are used for in-depth analysis, and the analysis results are visualized through feature-level fusion and decision-making fusion methods.

Benefits of technology

It improves the quality and consistency of multi-source data, enhances the accuracy of data fusion, meets users' multiple goals and needs, and makes analysis results richer, thereby improving the accuracy and efficiency of data analysis and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989275A_ABST
    Figure CN119989275A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source data fusion analysis processing method, which comprises the following steps: S1, data selection: firstly, constructing a feature model I, and collecting mutually corresponding target data and target results from different data sources by the feature model I; s2, data cleaning and preprocessing: the feature model I cleans the collected target data, removes repeated data, removes error data and processes missing values; s3, performing data fusion analysis, constructing a second feature model, and visualizing an analysis result in the form of charts.According to the method, firstly, multi-source data are cleaned and preprocessed, so that the quality and consistency of the data are ensured, the accuracy of subsequent data fusion is improved, and the data fusion efficiency is improved; data fusion and analysis are performed through two methods of feature-level fusion and decision-level fusion, so that various targets and requirements of users are met, analysis results are richer, and the accuracy and efficiency of data analysis and decision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of chemical waste gas treatment equipment, and in particular to a multi-source data fusion processing method. Background Art

[0002] Due to the lack of overall enterprise-level data planning and coordination, the construction of traditional "chimney-style" data platforms has failed to fully follow unified standards. Instead, they have only designed and built relatively independent data collection, aggregation, and data processing mechanisms according to the needs of each business. This has led to inconsistent data standards, data redundancy, and low data quality in the same business scenario.

[0003] With the advent of the big data era, we are faced with massive, multi-source data. How to extract useful information from these data has become an important issue, and multi-source data fusion technology has emerged. Multi-source data fusion technology refers to the technology of using relevant means to integrate all the information obtained through investigation and analysis, and to make a unified evaluation of the information, and finally obtain unified information. The purpose of developing this technology is to integrate various different data information, absorb the characteristics of different data sources, and then extract unified, better and richer information from them than single data, thereby improving the accuracy and efficiency of data analysis and decision-making.

[0004] In practical applications, multi-source data fusion analysis and processing still face many challenges. On the one hand, the quality and consistency of multi-source data is an important challenge. The data quality and format of different data sources may differ, resulting in deviations and errors in the fusion and analysis process. On the other hand, the heterogeneity and scale of multi-source data are also a problem. The data types and distributions of different data sources may vary greatly. At the same time, the amount of data is huge. How to efficiently process and analyze this data is also a problem that needs to be solved. Summary of the invention

[0005] In order to make up for the above shortcomings, the present invention provides a multi-source data fusion processing method.

[0006] The present invention is achieved through the following technical solutions:

[0007] A multi-source data fusion analysis and processing method includes the following steps: S1, data selection, first constructing a feature model 1, determining goals and requirements and inputting them, the feature model 1 collects corresponding target data and target results from different data sources, and the data sources include relational databases, non-relational databases, real-time streaming data, Web data, sensor data, open data, and geographic spatial data; S2, data cleaning and preprocessing, the feature model 1 cleans the collected target data, including removing duplicate data, removing erroneous data, and processing missing values; the target data and target results are formatted and standardized to ensure data quality and consistency; S3, data fusion analysis, constructing a feature model 2, using statistical analysis, data mining, and machine learning technical methods to conduct in-depth analysis on the preprocessed data, identify patterns, associations, and trends in the data, and use feature-level fusion and decision-level fusion methods to visualize the analysis results in the form of charts and texts, respectively, to help users understand the data and analysis results more intuitively.

[0008] Further optimized, the relational database includes MySQL, PostgreSQL, Oracle, and SQL Server; the non-relational database includes document database, key-value database, and column family database; the real-time streaming data includes real-time generated data from sensors, Internet of Things devices, and log files; the Web data includes data acquired through web crawlers; the sensor data includes data from various sensors such as weather stations, surveillance cameras, and industrial sensors; the open data includes data from public data sets, government data, and scientific research data websites; the geospatial data includes map data, geographic tags, and location data from a geographic information system (GIS).

[0009] For further optimization, based on the algorithm of deep neural network, feature model 1 is established to collect and process the target data in steps 1 and 2. The working steps of feature model 1 are: receiving the goals and requirements input by the customer, searching from different data sources to obtain the target data and the corresponding target results; vertically comparing to remove the target data with missing target results; classifying the target data according to the format, and horizontally comparing to remove duplicate data and erroneous data with large differences; converting the format of the remaining target data and target results, and standardizing them.

[0010] Further optimized, the comparison method of the target data processing feature model includes a common word segmentation method, a feature extraction method and a weighted average algorithm.

[0011] Further optimized, based on the deep neural network algorithm, feature model 2 is established for the data fusion analysis in step 3.

[0012] Further optimized, the steps of the feature model 2 for fusion analysis of target data and target results are: receiving the standardized data import in step 2; finding the optimal feature fusion strategy through grid search and Bayesian optimization methods, selecting and extracting useful features from different data sources, integrating them through connection, summation, and weighted summation, and constructing a representative feature set; visually presenting the analysis results in the form of charts and text.

[0013] Further optimized, the steps of the feature model 2 for fusion analysis of target data and target results are: receiving the standardized data import in step 2; classifying the target results from different data sources by classification and comparison, comparing them separately, and then integrating the comparison results by weight allocation to obtain the final decision result; visually presenting the analysis results in the form of charts and text.

[0014] Further optimization also includes step 4, verifying the analysis results by referring to national standards and industry best cases.

[0015] Further optimization includes a data input module for unified access to various types of data sources; a data collection and processing module for building vital sign model one, receiving input goals and requirements, collecting and processing target data; a data fusion and analysis module for building vital sign model two, intelligently fusing and analyzing pre-processed data, and displaying analysis results in various forms.

[0016] The beneficial effects of the present invention are:

[0017] The present invention first cleans and preprocesses multi-source data to ensure the quality and consistency of the data and improve the accuracy of subsequent data fusion. It performs data fusion and analysis through both feature-level fusion and decision-level fusion to meet the user's various goals and needs, enriching the analysis results and thus improving the accuracy and efficiency of data analysis and decision-making.

[0018] The present invention is based on the algorithm of deep neural network, establishes two feature models in two steps, and processes and analyzes multi-source data in two steps, reducing the burden of simultaneously fusing and analyzing huge amounts of data. Feature model 1 searches for data from multiple channels according to user needs and goals, and standardizes the data and provides it to feature model 2. Feature model 2 performs intelligent fusion analysis on the data and presents the analysis results to the user in a visual manner.

[0019] In the feature model 1 of the present invention, only the target data is classified and compared, and the target result is ignored, so the amount of data processing is halved, which greatly reduces the burden of comparison, format conversion and standardization processing, and improves the efficiency and speed of multi-source data processing. Although some data will be processed more, it will not affect the accuracy of the analysis results when the data volume is large. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION

[0021] In order to clearly illustrate the technical features of the present solution, the present invention is described in detail below through specific implementation methods and in combination with the accompanying drawings. The accompanying drawings should not be construed as limiting the present invention.

[0022] As shown in the figure, the present invention provides a multi-source data fusion analysis and processing method, comprising the following steps:

[0023] S1. Data selection: First, build feature model 1 based on the algorithm of deep neural network, determine the goals and requirements and input them. Feature model 1 collects corresponding target data and target results from different data sources, including relational databases, non-relational databases, real-time streaming data, Web data, sensor data, open data, and geospatial data. Collect as much data as possible to make the analysis results richer, thereby improving the accuracy and efficiency of data analysis and decision-making.

[0024] S2. Data cleaning and preprocessing: Feature model 1 cleans the collected target data, including removing duplicate data, removing erroneous data, and processing missing values; converting and standardizing the target data and target results to ensure data quality and consistency. Ensure data quality and consistency to improve the accuracy of subsequent data fusion.

[0025] S3. Data fusion analysis, building feature model 2, using statistical analysis, data mining, and machine learning techniques to conduct in-depth analysis of the pre-processed data, identify patterns, associations, and trends in the data, and use feature-level fusion and decision-level fusion methods to visualize the analysis results in the form of charts and texts, helping users to understand the data and analysis results more intuitively.

[0026] As a preferred implementation, the relational database includes MySQL, PostgreSQL, Oracle, and SQL Server; the non-relational database includes a document database, a key-value database, and a column family database; the real-time streaming data includes real-time generated data from sensors, IoT devices, and log files; the Web data includes data obtained through web crawlers; the sensor data includes data from various sensors such as weather stations, surveillance cameras, and industrial sensors; the open data includes data from public data sets, government data, and scientific research data websites; and the geospatial data includes map data, geographic markers, and location data from a geographic information system (GIS). Collecting more data bases makes the samples richer, thereby improving the accuracy and efficiency of data analysis and decision-making.

[0027] As a preferred implementation, based on a deep neural network algorithm, a feature model 1 is established to process the target data in step 2. The working steps of the feature model 1 are as follows:

[0028] Receive goals and requirements input by customers, and search from different data sources to obtain target data and corresponding target results;

[0029] Longitudinal comparison to remove target data with missing target results;

[0030] Classify the target data according to the format, and perform horizontal comparison to remove duplicate data and erroneous data with large differences;

[0031] The remaining target data and target results are converted into different formats and standardized.

[0032] This process only performs classification and comparison processing based on the target data, ignoring the target results, and the amount of data processing is halved, which greatly reduces the burden of comparison, format conversion and standardization processing, and improves the efficiency and speed of multi-source data processing. It should be noted that although some data will be processed more, it will not affect the accuracy of the analysis results when the data volume is large.

[0033] As a preferred implementation, the comparison method of the target data processing feature model includes a common word segmentation method, a feature extraction method and a weighted average algorithm.

[0034] As a preferred implementation, based on the algorithm of deep neural network, feature model 2 is established for data fusion analysis in step 3. Feature model 2 performs fusion analysis on target data and target results through two methods. The steps of method 1 are:

[0035] Receive the standardized data import in step 2;

[0036] Through grid search and Bayesian optimization methods, we find the optimal feature fusion strategy, select and extract useful features from different data sources, integrate them through connection, summation, and weighted summation, and build a representative feature set.

[0037] The analysis results are presented visually in the form of charts and text.

[0038] The steps of method 2 are:

[0039] Receive the standardized data import in step 2;

[0040] The target results from different data sources are classified and compared by classification and comparison, and then the comparison results are integrated by weight allocation to obtain the final decision result;

[0041] The analysis results are presented visually in the form of charts and text.

[0042] As a preferred implementation, step 4 is also included. When the user receives the analysis result, he can refer to national standards and industry best cases to verify and evaluate the analysis result.

[0043] The above two methods can extract the advantages of different source data and obtain richer and more comprehensive information through feature-level fusion. For example, the fusion of hyperspectral and radar data can overcome the limitations of spectral information and achieve more accurate identification and classification of targets. Through decision-level fusion, the decision results of different data sources can be integrated, with good real-time performance, less impact by sensor registration errors, and strong fault tolerance. For example, in medical diagnosis, the diagnostic results from different medical devices can be fused to improve accuracy and reliability.

[0044] As a preferred implementation, it includes a data input module for unified access of various types of data sources; a data collection and processing module for building a physical sign model 1, receiving input goals and requirements, collecting and processing target data; a data fusion and analysis module for building a physical sign model 2, intelligently fusing and analyzing pre-processed data, and displaying the analysis results in various forms.

[0045] Those of ordinary skill in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be accomplished by instructing the relevant hardware through a computer program. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory, magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory or external cache memory. As an illustration and not limitation, RAM may be in various forms, such as static random access memory or dynamic random access memory, etc.

[0046] The parts not described in detail in the present invention are all known technologies to those skilled in the art. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions of the present invention, which should be included in the scope of the claims of the present invention.

Claims

1. A multi-source data fusion analysis and processing method, characterized in that: The following steps are involved: S1. Data selection. First, construct feature model 1, determine the goals and requirements and input them. Feature model 1 collects corresponding target data and target results from different data sources. The data sources include relational databases, non-relational databases, real-time streaming data, Web data, sensor data, open data, and geospatial data. S2. Data cleaning and preprocessing: Feature model 1 cleans the collected target data, including removing duplicate data, removing erroneous data, and processing missing values; converting and standardizing the target data and target results to ensure data quality and consistency; S3. Data fusion analysis, building feature model 2, using statistical analysis, data mining, and machine learning techniques to conduct in-depth analysis of the pre-processed data, identify patterns, associations, and trends in the data, and use feature-level fusion and decision-level fusion methods to visualize the analysis results in the form of charts and texts, helping users to understand the data and analysis results more intuitively.

2. The multi-source data fusion processing method according to claim 1, characterized in that: The relational databases include MySQL, PostgreSQL, Oracle, and SQL Server; the non-relational databases include document databases, key-value databases, and column family databases; the real-time streaming data include real-time generated data from sensors, IoT devices, and log files; the Web data include data acquired through web crawlers; the sensor data include data from various sensors such as weather stations, surveillance cameras, and industrial sensors; the open data include data from public data sets, government data, and scientific research data websites; the geospatial data include map data, geographic tags, and location data from a geographic information system (GIS).

3. The multi-source data fusion processing method according to claim 1, characterized in that: Based on the algorithm of deep neural network, feature model 1 is established to collect and process the target data in steps 1 and 2. The working steps of feature model 1 are as follows: Receive goals and requirements input by customers, and search from different data sources to obtain target data and corresponding target results; The longitudinal comparison removes target data with missing target results, and target results with missing target data; Classify the target data according to the format, and perform horizontal comparison to remove duplicate data and erroneous data with large differences; The remaining target data and target results are converted into different formats and standardized.

4. The multi-source data fusion processing method according to claim 3 is characterized in that: The comparison method of the target data processing feature model includes a common word segmentation method, a feature extraction method and a weighted average algorithm.

5. The multi-source data fusion processing method according to claim 1, characterized in that: Based on the deep neural network algorithm, feature model 2 is established for the data fusion analysis in step 3.

6. The multi-source data fusion processing method according to claim 5, characterized in that: The steps of the feature model 2 for fusion analysis of target data and target results are as follows: Receive the standardized data import in step 2; Through grid search and Bayesian optimization methods, we find the optimal feature fusion strategy, select and extract useful features from different data sources, integrate them through connection, summation, and weighted summation, and build a representative feature set. The analysis results are presented visually in the form of charts and text.

7. The multi-source data fusion processing method according to claim 5, characterized in that: The steps of the feature model 2 for fusion analysis of target data and target results are as follows: Receive the standardized data import in step 2; The target results from different data sources are classified and compared by classification and comparison, and then the comparison results are integrated by weight allocation to obtain the final decision result; The analysis results are presented visually in the form of charts and text.

8. The multi-source data fusion processing method according to claim 1, characterized in that: It also includes step 4, which verifies the analysis results by referring to national standards and industry best cases.

9. A multi-source data fusion analysis and processing system, characterized by: It includes a data input module for unified access to various types of data sources; a data collection and processing module for building a physical sign model, receiving input goals and requirements, and collecting and processing target data; The data fusion analysis module is used to construct the second vital sign model, intelligently fuse and analyze the preprocessed data, and display the analysis results in various forms.

Citation Information

Cited By

  • Traffic accident data processing and analysis method and device and electronic equipment

    CN120892768A