Multi-source heterogeneous data acquisition system based on artificial intelligence

Through the AI-based multi-source heterogeneous data acquisition system, the problems of data acquisition systems being susceptible to interference from complex environments and insufficient resource scheduling efficiency are solved, efficient data cleaning, evaluation and resource optimization are achieved, and data quality and system stability are improved. It is suitable for industrial Internet of Things and intelligent operation and maintenance.

CN120653891AInactive Publication Date: 2025-09-16HEBEI LESHU HUIQUAN ELECTRONIC TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510794300.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing data acquisition system is easily affected by complex environmental interference, and its resource scheduling efficiency and acquisition processing efficiency are insufficient, resulting in low quality of collected data and difficulty in meeting the needs of multi-source heterogeneous data processing.

Method used

An artificial intelligence-based multi-source heterogeneous data acquisition system is adopted, including a multi-source data monitoring unit, a data cleaning and repair unit, a data quality assessment unit and a resource dynamic scheduling unit. Through a dynamic cleaning rule base, multimodal fault detection, quality quantitative assessment and intelligent resource scheduling, efficient data integration, cleaning, assessment and resource optimization are achieved.

Benefits of technology

It significantly improves the accuracy, reliability and intelligence of the data acquisition system, realizes the intelligence of data feature extraction and risk assessment, and supports efficient data management in scenarios such as industrial Internet of Things and intelligent operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653891A_ABST
    Figure CN120653891A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source heterogeneous data acquisition system based on artificial intelligence, which relates to the technical field of Internet of Things operating systems, and comprises a multi-source data monitoring unit, a data cleaning and repairing unit, a data quality evaluation unit and a resource dynamic scheduling unit, the problems that an existing data acquisition system is prone to being interfered by a complex environment, resource scheduling efficiency and acquisition processing efficiency are insufficient, the acquired data quality is low, and the defect that the multi-source heterogeneous data processing requirement is difficult to meet is overcome. The whole process from the multi-source data monitoring unit, the data cleaning and repairing unit, the data quality assessment unit to the resource dynamic scheduling unit is automatic, manual intervention is reduced, the processing efficiency is improved, and intelligence of data feature extraction and risk assessment is realized through artificial intelligence deep fusion. The accuracy, the reliability and the intelligent level of the data acquisition system are remarkably improved, and technical support is provided for scenes such as industrial Internet of Things and intelligent operation and maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Internet of Things operating systems, and in particular to a multi-source heterogeneous data acquisition system based on artificial intelligence. Background Art

[0002] AI-based multi-source heterogeneous data acquisition systems are widely used in smart cities, industrial internet, financial risk control, healthcare, environmental monitoring, and other fields. These systems aim to acquire data from multiple sources with diverse structures and use AI technology for preliminary processing, integration, and analysis. However, their development still faces many challenges. With the development of the Internet of Things (IoT) and the Industrial Internet, the number of sensor devices has exploded, leading to an increased demand for multi-source heterogeneous data processing. This multi-source heterogeneous data contains multi-dimensional parameters such as time, environment, attributes, and logs. It has complex formats, diverse structures, and high levels of noise interference. Traditional data acquisition systems struggle to meet the demand for efficient processing. Furthermore, sensor devices in complex environments may generate outliers, missing values, or data with inconsistent formats. Log parameters such as device vibration spectra and current waveforms are susceptible to environmental interference, resulting in poor data quality and affecting the accuracy of subsequent analysis and decision-making. Furthermore, traditional systems lack the ability to dynamically manage computing, storage, and network resources, making it difficult to cope with data volume fluctuations and device failures such as CPU overload, disk damage, and network outages. This results in insufficient system stability and high maintenance costs. Therefore, existing data acquisition systems are susceptible to interference from complex environments, have insufficient resource scheduling efficiency and acquisition and processing efficiency, resulting in low quality of collected data and making it difficult to meet the needs of multi-source heterogeneous data processing; In view of the above technical defects, a solution is now proposed. Summary of the Invention

[0003] The purpose of the present invention is to solve the problems of existing data acquisition systems that are susceptible to interference from complex environments, have insufficient resource scheduling efficiency and acquisition processing efficiency, resulting in low quality of collected data, and are unable to meet the needs of multi-source heterogeneous data processing.

[0004] In order to achieve the above object, the present invention adopts the following technical solutions: An artificial intelligence-based multi-source heterogeneous data acquisition system, including a multi-source data monitoring unit, a data cleaning and repair unit, a data quality assessment unit, and a resource dynamic scheduling unit; The multi-source data monitoring unit is used to import multi-source heterogeneous data: it imports the detection data sources of N0 sensor devices through IoT communication and integrates and marks them as multi-source heterogeneous data; The data cleaning and repair unit is used to clean and repair multi-source heterogeneous data. By inputting the working data of sensor devices, a dynamic cleaning rule base is established to clean the multi-source heterogeneous data. The cleaned multi-source heterogeneous data is marked as multi-source valid data. Multi-modal fault detection is performed on sensor devices to generate early warning fault tolerance strategies. The data quality assessment unit is used to evaluate the quality of multi-source valid data: by establishing a historical database of multi-source valid data and setting a data comparison cycle for regular comparison, the accuracy, completeness, and timeliness of the multi-source valid data are analyzed in sequence, and then the data quality level is determined through comprehensive evaluation, and a regional data quality heat map is generated; The resource dynamic scheduling unit is used to optimize the resource scheduling of the acquisition system: by building a digital twin simulation platform, the resource distribution of the acquisition system is monitored in real time, the usage of computing resources, storage resources and network resources is obtained, and fault injection testing is performed to output a dynamic resource scheduling plan.

[0005] Furthermore, the job operation data includes time windows, environmental parameters, attribute parameters, and log parameters; Among them, environmental parameters include temperature and humidity; attribute parameters include sensor device model and installation location; log parameters include working time, vibration spectrum and current waveform; Mark any sensor device as i, integrate the detection data sources of sensor device i and establish a set D: ; Mark any data point in set D as : ; is the time window, , is the starting time point of the time window, is the end time point of the time window; is the environmental parameter vector, , is the temperature, for humidity; is the attribute parameter vector, , is the sensor device model, Provide installation locations for sensor equipment; is the log parameter vector, , For working hours, is the vibration spectrum, is the current waveform; Integrate the detection data sources of N0 sensor devices to establish a multi-source heterogeneous data set: .

[0006] Furthermore, by inputting the working data of sensor equipment, combining time windows, environmental parameters and attribute parameters, a dynamic cleaning rule base is established to clean multi-source heterogeneous data and mark the cleaned multi-source heterogeneous data as multi-source valid data; Set the cleaning rule set Rc, through the attribute parameter vector Perform data classification and configure corresponding static cleaning rules Rc; Through the time window , environmental parameter vector and attribute parameter vector Combined, dynamic cleaning rule Rd is generated: ; Where H is the dynamic rule generation function; Cleaning rules include data preprocessing and outlier correction to obtain valid data from multiple sources; Through the time window , environmental parameter vector and attribute parameter vector For the data, set the corresponding dynamic cleaning rule Rd.

[0007] Furthermore, the specific process of generating an early warning fault tolerance strategy is as follows: Perform multimodal fault detection on sensor devices by analyzing log parameters; is the log parameter vector, , For working hours, is the vibration spectrum, is the current waveform; Vibration spectrum by Fourier transform Perform frequency domain analysis to extract vibration frequency domain features Jvs; vibration frequency domain features Jvs include main frequency amplitude Adom, energy distribution Ef and kurtosis Kf; Current waveform Perform time domain analysis to extract the current waveform characteristics Jcw; the current waveform characteristics Jcw include the peak factor Fj and the effective value Ij; By working hours Calculate and extract the time feature Jwt; the time feature Jwt includes the total running time Tw of the device and the running time ratio Rw; The multi-dimensional feature vector is obtained by splicing the vibration frequency domain feature Jvs, the current waveform feature Jcw and the time feature Jwt. : ; Through multidimensional feature vector Analyze and calculate the mean square error (MSE), evaluate the abnormal coefficient of each feature, and perform weighted summation of the abnormal coefficients of the features to obtain the multimodal fault risk index (Risk) through comprehensive evaluation; Set the evaluation interval of the multimodal fault risk index Risk, determine the equipment risk level by comparing the intervals, and thus generate an early warning fault tolerance strategy.

[0008] Furthermore, by establishing a historical database of valid data from multiple sources and setting a data comparison period for regular comparison, the accuracy, completeness, and timeliness of the valid data from multiple sources can be analyzed in turn. The specific process is as follows: Compare the multi-source valid data obtained at regular intervals with the standard measurement data obtained by high-precision sensors to analyze the degree of deviation between the multi-source valid data and the standard measurement data; The number of indicators of multi-source valid data is marked as M0, any indicator value of multi-source valid data is marked as Xm, and the indicator value of the standard measurement data corresponding to Xm is marked as Ym, so as to obtain the accuracy evaluation coefficient , used to evaluate the accuracy of multi-source valid data; By cleaning multi-source heterogeneous data, the number of missing values ​​Ns is obtained, and the completeness assessment coefficient is calculated. , used to evaluate the completeness of multi-source valid data; Calculate the interval duration by the difference between the collection time node of multi-source valid data and the current time node , and set the interval time The threshold Rt, when the interval length If the value is lower than the threshold Rt, the timeliness of the evaluation data indicator is up to standard. The timeliness evaluation coefficient is calculated by accumulating the number of indicators that meet the timeliness standards of all valid data from multiple sources and marking it as Nt. , used to evaluate the timeliness of multi-source valid data.

[0009] Furthermore, the specific process of generating a regional data quality heat map is as follows: Accuracy evaluation coefficient , completeness assessment coefficient and timeliness evaluation coefficient Perform weighted summation to obtain the data quality index Q; By setting the evaluation interval of the data quality index Q and performing interval comparison, the data quality can be comprehensively evaluated and the data quality level can be determined; The collection area is marked by the installation locations of N0 sensor devices, and the data quality levels of the sensor devices are displayed to generate a regional data quality heat map.

[0010] Furthermore, by building a digital twin simulation platform, the resource distribution of the acquisition system is monitored in real time to obtain the usage of computing resources, storage resources, and network resources. The specific process is as follows: By monitoring CPU usage and GPU usage and performing weighted fusion, we can obtain computing resource usage and analyze computing resource usage. By monitoring the disk remaining space ratio r1 and the disk read and write rate r2, and performing weighted fusion, the storage resource remaining rate is obtained, thereby analyzing the storage resource usage; By monitoring the network bandwidth utilization u1 and network delay time u2 and performing weighted fusion, the network resource utilization is obtained, thereby analyzing the usage of network resources.

[0011] Furthermore, a fault injection test is conducted through the digital twin simulation platform to output a dynamic resource scheduling plan. The specific process is as follows: Fault injection testing includes CPU overload on computing nodes, disk damage on storage devices, and network link interruption. By adjusting the fault type and duration of the sensor equipment, the operating status of the data acquisition system can be observed; By coordinating the allocation of computing resources, storage resources, and network resources, the changes in the data quality index Q are monitored synchronously, and the solution corresponding to the maximum value of the data quality index Q is marked as the dynamic resource scheduling solution.

[0012] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: This invention automates the entire process from data collection, cleaning and repair, quality assessment to resource scheduling, reducing manual intervention and improving processing efficiency. Through deep integration of artificial intelligence, the synergy of dynamic cleaning, multimodal detection, quality quantitative assessment, and intelligent resource scheduling enables intelligent data feature extraction and risk assessment, significantly improving the accuracy, reliability, and intelligence level of the data collection system, and providing technical support for scenarios such as the Industrial Internet of Things and intelligent operation and maintenance. This invention efficiently integrates and cleans multi-source heterogeneous data, combines time windows, environmental parameters, and attribute parameters to generate a dynamic cleaning rule base, achieves data format standardization, time alignment, and missing value repair, and improves data availability. It also constructs a multidimensional feature vector through multimodal fault detection, combines vibration frequency domain features, current waveform features, and time features, calculates a multimodal fault risk index, implements predictive maintenance for equipment failures, and improves risk management efficiency. The present invention implements quantitative analysis of data quality through multi-dimensional quality assessment of data accuracy, completeness, and timeliness, thereby generating a regional data quality heat map and visualizing the data quality level according to the equipment installation location, so as to quickly locate low-quality data areas, support targeted optimization, and improve data management efficiency; by building a digital twin simulation platform, the usage of computing resources, storage resources, and network resources is monitored in real time, abnormal scenarios are simulated through fault injection testing, and the optimal resource scheduling plan is dynamically generated to improve the system resource coordination and allocation capabilities, ensuring that the system can still maintain high data quality in fault scenarios, and through the dynamic allocation of resources and fault handling strategy verification, the stability and robustness of the system are improved, the waste of hardware resources is reduced, and intelligent operation and maintenance are realized. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 shows a schematic diagram of the connection of the system modules of the present invention; Figure 2 A schematic diagram showing the steps of the workflow of the present invention is shown. DETAILED DESCRIPTION

[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0015] Example 1: like Figure 1-Figure 2 As shown, the multi-source heterogeneous data acquisition system based on artificial intelligence includes a multi-source data monitoring unit, a data cleaning and repairing unit, a data quality assessment unit and a resource dynamic scheduling unit, wherein the multi-source data monitoring unit, the data cleaning and repairing unit, the data quality assessment unit and the resource dynamic scheduling unit are communicatively connected; The working steps are as follows: S1, the multi-source data monitoring unit imports multi-source heterogeneous data: imports the detection data sources of N0 sensor devices through IoT communication, and integrates and marks them as multi-source heterogeneous data; Job operation data includes time window, environment parameters, attribute parameters and log parameters; Among them, environmental parameters include temperature and humidity; attribute parameters include sensor device model and installation location; log parameters include working time, vibration spectrum and current waveform; Mark any sensor device as i, integrate the detection data sources of sensor device i and establish a set D: ; Mark any data point in set D as : ; is the time window, , is the starting time point of the time window, is the end time point of the time window; is the environmental parameter vector, , is the temperature, for humidity; is the attribute parameter vector, , is the sensor device model, Provide installation locations for sensor equipment; is the log parameter vector, , For working hours, is the vibration spectrum, is the current waveform; Integrate the detection data sources of N0 sensor devices to establish a multi-source heterogeneous data set: .

[0016] S2, the data cleaning and repair unit cleans and repairs multi-source heterogeneous data: S2-1, by inputting the working data of the sensor equipment, a dynamic cleaning rule base is established to clean the multi-source heterogeneous data, and the cleaned multi-source heterogeneous data is marked as multi-source valid data. By inputting the working data of sensor equipment, combining time windows, environmental parameters and attribute parameters, a dynamic cleaning rule base is established to clean multi-source heterogeneous data and mark the cleaned multi-source heterogeneous data as multi-source valid data; S2-101, through the time window Get data sampling frequency : , For sensor device i in the time window The number of data sampling times; Through the environmental parameter vector and the log parameter vector Analyze historical data and establish parameter impact regression model: , where F is the mapping function from environmental parameters to log parameters, is the error term; Set the cleaning rule set Rc, through the attribute parameter vector Perform data classification and configure the corresponding static cleaning rules Rc: , where G is the mapping function from attribute parameters to static cleaning rules; S2-102, through the time window , environmental parameter vector and attribute parameter vector Combined, dynamic cleaning rule Rd is generated: ; Where H is the dynamic rule generation function, and H includes the time window , environmental parameter vector and attribute parameter vector The numerical conditions satisfied by each indicator, and the corresponding cleaning rules are applied according to different numerical conditions; Among them, the cleaning rules include data preprocessing and outlier correction, so as to obtain valid data from multiple sources; Data preprocessing includes format standardization, time node alignment, and missing value processing. Format standardization methods include Z-score standardization and maximum-minimum value standardization. Time node alignment is to align data with different sampling frequencies to a unified time grid. Missing value processing methods include mean filling and linear interpolation. Outlier correction includes Z-score value comparison and error term Comparison, specifically setting the corresponding threshold, performing anomaly detection through threshold comparison, and thus determining that the data is an outlier. The outlier correction method is the same as the missing value processing method; Through the time window , environmental parameter vector and attribute parameter vector The corresponding dynamic cleaning rule Rd is set for the data. The specific rule setting needs to be combined with the experience of professional technicians and actual application needs. The function of this module is to dynamically determine the use of a cleaning rule through data analysis; Through the efficient integration and cleaning of multi-source heterogeneous data, a dynamic cleaning rule base is generated by combining time windows, environmental parameters (temperature, humidity) and attribute parameters (device model, installation location), which can achieve data format standardization, time alignment and missing value repair, thereby improving data availability.

[0017] S2-2, by performing multi-modal fault detection on sensor devices, an early warning fault-tolerant strategy is generated; S2-201, multi-modal fault detection of sensor devices by analyzing log parameters; is the log parameter vector, , For working hours, is the vibration spectrum, is the current waveform; S2-202, vibration spectrum by Fourier transform Perform frequency domain analysis and extract vibration frequency domain features Jvs; The vibration frequency domain characteristics Jvs include the main frequency amplitude Adom, energy distribution Ef and kurtosis Kf; Among them, the amplitude of the peak frequency in the vibration spectrum is marked as the main frequency amplitude Adom, the energy proportion of each frequency interval is integrated and marked as the energy distribution Ef, and the kurtosis Kf is calculated by normalizing the fourth-order central moment in the frequency domain; Current waveform Perform time domain analysis to extract the current waveform feature Jcw; The current waveform characteristics Jcw include the peak factor Fj and the effective value Ij; Among them, the current waveform The ratio of the peak value to the average value is marked as the peak factor Fj, and the current waveform The absolute value of the mean is marked as the effective value Ij; By working hours Calculate and extract time feature Jwt; The time feature Jwt includes the total running time Tw and the running time ratio Rw of the device; Among them, through the time window Accumulate the equipment running time and obtain the working time The timing distribution diagram is obtained, and the total device running time Tw is calculated cumulatively. The running time proportion Rw is calculated by the ratio of the total device running time Tw to the total window duration. S2-203, obtain a multi-dimensional feature vector by splicing the vibration frequency domain feature Jvs, the current waveform feature Jcw and the time feature Jwt : ; Through multidimensional feature vector Analyze and calculate the mean square error (MSE) to evaluate the abnormal coefficient of each feature. Then, by weighted summing the abnormal coefficients of the features, a comprehensive evaluation is performed to obtain the multimodal fault risk index (Risk), quantifying the comprehensive fault risk level of the multimodal fault. S2-204, set the evaluation interval of the multi-modal fault risk index Risk, determine the equipment risk level by comparing the intervals, and thus generate an early warning fault tolerance strategy; Early warning fault tolerance strategy means formulating corresponding response plans based on risk levels to achieve predictive maintenance to avoid failures; For example, for low-risk levels, the monitoring period is extended; for medium-risk levels, an early warning is triggered and manual inspections are arranged; for high-risk registrations, the system is immediately shut down and a fault work order is generated; By constructing a multidimensional feature vector through multimodal fault detection, the vibration frequency domain features, current waveform features and time features are combined to calculate the multimodal fault risk index, enabling predictive maintenance of equipment failures (such as extending the monitoring period for low-risk faults and immediate shutdown for high-risk faults), thereby reducing downtime losses.

[0018] S3, the data quality assessment unit assesses the quality of valid data from multiple sources: S3-1, by establishing a historical database of valid data from multiple sources and setting a data comparison period for regular comparison, the accuracy, completeness and timeliness of valid data from multiple sources are analyzed in turn; S3-101, comparing the multi-source valid data acquired at regular intervals with the standard measurement data acquired by the high-precision sensor, thereby analyzing the degree of deviation between the multi-source valid data and the standard measurement data; The number of indicators of multi-source valid data is marked as M0, any indicator value of multi-source valid data is marked as Xm, and the indicator value of the standard measurement data corresponding to Xm is marked as Ym, so as to obtain the accuracy evaluation coefficient : ; in, It refers to the deviation coefficient. The threshold value Rm of the deviation coefficient is set. When the deviation coefficient is lower than the threshold Rm, the accuracy of the evaluation data indicator meets the standard. The number of indicators that meet the accuracy standard of all valid data from multiple sources is accumulated and marked as Nm; When the accuracy evaluation coefficient The higher it is, the higher the accuracy of evaluating multi-source valid data; S3-102, by cleaning multi-source heterogeneous data, obtain the number of missing values ​​Ns, and then calculate the completeness assessment coefficient : ; When the completeness evaluation coefficient The higher it is, the higher the completeness of the multi-source valid data is; S3-103, calculate the interval duration by the difference between the collection time node of multi-source valid data and the current time node , and set the interval time The threshold Rt, when the interval length If the value is lower than the threshold Rt, the timeliness of the evaluation data indicator is up to standard. The timeliness evaluation coefficient is calculated by accumulating the number of indicators that meet the timeliness standards of all valid data from multiple sources and marking it as Nt. : ; Time effectiveness evaluation coefficient The higher it is, the higher the timeliness of evaluating multi-source valid data is; S3-2, comprehensively evaluate and determine the data quality level and generate a regional data quality heat map; S3-201, passed the accuracy assessment coefficient , completeness assessment coefficient and timeliness evaluation coefficient Perform weighted summation to obtain the data quality index Q; S3-202, by setting the evaluation interval of the data quality index Q and performing interval comparison, the data quality is comprehensively evaluated and the data quality level is determined; The collection area is marked by the installation locations of N0 sensor devices, and the data quality level of the sensor devices is displayed to generate a regional data quality heat map; Among them, the regional data quality heat map is used to visualize the data quality of each data collection area, making it easier to quickly locate the source of quality problems and carry out targeted treatment; For example: automatically reducing the collection frequency of low-quality data sources, or adaptively adjusting the equipment operating environment in low-quality areas; Through multi-dimensional quality assessment, we sequentially obtain evaluation coefficients for accuracy, completeness, and timeliness, enabling quantitative analysis of data quality. This generates a regional data quality heat map, visualizes data quality levels based on device installation locations, and quickly locates low-quality data areas, supporting targeted optimization and improving data management efficiency.

[0019] S4, the resource dynamic scheduling unit optimizes the resource scheduling of the acquisition system: S4-1, by building a digital twin simulation platform, the resource distribution of the acquisition system is monitored in real time to obtain the usage of computing resources, storage resources, and network resources; S4-101, by monitoring CPU usage and GPU usage and performing weighted fusion, obtains computing resource usage, thereby analyzing computing resource usage; S4-102: By monitoring the disk remaining space ratio r1 and the disk read and write rate r2, and performing weighted fusion, the storage resource remaining rate is obtained to analyze the storage resource usage; S4-103, by monitoring the network bandwidth utilization u1 and the network delay time u2 and performing weighted fusion, the network resource utilization is obtained, thereby analyzing the network resource usage; S4-2, and perform fault injection testing to output a dynamic resource scheduling solution; Conduct fault injection testing through the digital twin simulation platform and output dynamic resource scheduling solutions; Types of fault injection testing include computing node CPU overload, storage device disk damage, and network link interruption; By adjusting the fault type and duration of sensor equipment, the operating status of the data acquisition system can be observed. By coordinating the allocation of computing resources, storage resources, and network resources, the changes in the data quality index Q are synchronously monitored. The solution corresponding to the maximum value of the data quality index Q is marked as the dynamic resource scheduling solution. This allows the rationality of the fault handling solution to be verified in a virtual environment and the efficiency of the strategy evolution to be improved. By building a digital twin simulation platform, we monitor the usage of computing resources (CPU / GPU usage), storage resources (disk space / read / write speed), and network resources (bandwidth utilization / latency) in real time. Through fault injection testing, we simulate abnormal scenarios (such as CPU overload and network interruption), dynamically generate optimal resource scheduling plans, and improve the system's ability to coordinate resource allocation. This ensures that the system can still maintain high data quality in fault scenarios (with the goal of maximizing the data quality index Q). Through dynamic resource allocation and fault handling strategy verification, we improve the stability and robustness of the system, reduce hardware resource waste, and achieve intelligent operation and maintenance.

[0020] In summary, the present invention automates the entire process from data acquisition (multi-source data monitoring unit), cleaning (data cleaning and repair unit), quality assessment (data quality assessment unit) to resource scheduling (resource dynamic scheduling unit), reducing manual intervention and improving processing efficiency. Furthermore, through deep integration of artificial intelligence, using algorithms such as regression models, Fourier transforms, and mean square error (MSE), it achieves intelligent data feature extraction and risk assessment, providing technical support for scenarios such as the Industrial Internet of Things and intelligent operations and maintenance. Through the synergistic effects of dynamic cleaning, multimodal detection, quality quantitative assessment and intelligent resource scheduling, the present invention significantly improves the accuracy, reliability and intelligence level of the data acquisition system. It is suitable for complex scenarios requiring high data quality, such as industrial equipment monitoring, smart grids, and environmental monitoring, and has strong engineering application value.

[0021] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution.

[0022] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0023] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. An artificial intelligence-based multi-source heterogeneous data acquisition system, characterized by: It includes multi-source data monitoring unit, data cleaning and repair unit, data quality assessment unit and resource dynamic scheduling unit; The multi-source data monitoring unit is used to import multi-source heterogeneous data: it imports the detection data sources of N0 sensor devices through IoT communication and integrates and marks them as multi-source heterogeneous data; The data cleaning and repair unit is used to clean and repair multi-source heterogeneous data. By inputting the working data of sensor devices, a dynamic cleaning rule base is established to clean the multi-source heterogeneous data. The cleaned multi-source heterogeneous data is marked as multi-source valid data. Multi-modal fault detection is performed on sensor devices to generate early warning fault tolerance strategies. The data quality assessment unit is used to evaluate the quality of multi-source valid data: by establishing a historical database of multi-source valid data and setting a data comparison cycle for regular comparison, the accuracy, completeness, and timeliness of the multi-source valid data are analyzed in sequence, and then the data quality level is determined through comprehensive evaluation, and a regional data quality heat map is generated; The resource dynamic scheduling unit is used to optimize the resource scheduling of the acquisition system: by building a digital twin simulation platform, the resource distribution of the acquisition system is monitored in real time, the usage of computing resources, storage resources and network resources is obtained, and fault injection testing is performed to output a dynamic resource scheduling plan.

2. The multi-source heterogeneous data acquisition system based on artificial intelligence according to claim 1 is characterized by: Job operation data includes time window, environment parameters, attribute parameters and log parameters; Among them, environmental parameters include temperature and humidity; attribute parameters include sensor device model and installation location; log parameters include working time, vibration spectrum and current waveform; Mark any sensor device as i, integrate the detection data sources of sensor device i and establish a set D: ; Mark any data point in set D as : ; is the time window, , is the starting time point of the time window, is the end time point of the time window; is the environmental parameter vector, , is the temperature, for humidity; is the attribute parameter vector, , is the sensor device model, Provide installation locations for sensor equipment; is the log parameter vector, , For working hours, is the vibration spectrum, is the current waveform; Integrate the detection data sources of N0 sensor devices to establish a multi-source heterogeneous data set: 。 3. The multi-source heterogeneous data acquisition system based on artificial intelligence according to claim 2, characterized in that: By inputting the working data of sensor equipment, combining time windows, environmental parameters and attribute parameters, a dynamic cleaning rule base is established to clean multi-source heterogeneous data and mark the cleaned multi-source heterogeneous data as multi-source valid data; Set the cleaning rule set Rc, through the attribute parameter vector Perform data classification and configure corresponding static cleaning rules Rc; Through the time window , environmental parameter vector and attribute parameter vector Combined, dynamic cleaning rule Rd is generated: ; Where H is the dynamic rule generation function; Cleaning rules include data preprocessing and outlier correction to obtain valid data from multiple sources; Through the time window , environmental parameter vector and attribute parameter vector For the data, set the corresponding dynamic cleaning rule Rd.

4. The multi-source heterogeneous data acquisition system based on artificial intelligence according to claim 3 is characterized by: The specific process of generating early warning fault tolerance strategy is as follows: Perform multimodal fault detection on sensor devices by analyzing log parameters; is the log parameter vector, , For working hours, is the vibration spectrum, is the current waveform; Vibration spectrum by Fourier transform Perform frequency domain analysis to extract vibration frequency domain features Jvs; vibration frequency domain features Jvs include main frequency amplitude Adom, energy distribution Ef and kurtosis Kf; Current waveform Perform time domain analysis to extract the current waveform feature Jcw; The current waveform characteristics Jcw include the peak factor Fj and the effective value Ij; By working hours Calculate and extract time feature Jwt; The time feature Jwt includes the total running time Tw and the running time ratio Rw of the device; The multi-dimensional feature vector is obtained by splicing the vibration frequency domain feature Jvs, the current waveform feature Jcw and the time feature Jwt. : ; Through multidimensional feature vector Analyze and calculate the mean square error (MSE), evaluate the abnormal coefficient of each feature, and perform weighted summation of the abnormal coefficients of the features to obtain the multimodal fault risk index (Risk) through comprehensive evaluation; Set the evaluation interval of the multimodal fault risk index Risk, determine the equipment risk level by comparing the intervals, and thus generate an early warning fault tolerance strategy.

5. The multi-source heterogeneous data acquisition system based on artificial intelligence according to claim 4 is characterized in that: By establishing a historical database of valid data from multiple sources and setting a data comparison period for regular comparison, the accuracy, completeness, and timeliness of valid data from multiple sources can be analyzed in sequence. The specific process is as follows: Compare the multi-source valid data obtained at regular intervals with the standard measurement data obtained by high-precision sensors to analyze the degree of deviation between the multi-source valid data and the standard measurement data; The number of indicators of multi-source valid data is marked as M0, any indicator value of multi-source valid data is marked as Xm, and the indicator value of the standard measurement data corresponding to Xm is marked as Ym, so as to obtain the accuracy evaluation coefficient , used to evaluate the accuracy of multi-source valid data; By cleaning multi-source heterogeneous data, the number of missing values ​​Ns is obtained, and the completeness assessment coefficient is calculated. , used to evaluate the completeness of multi-source valid data; Calculate the interval duration by the difference between the collection time node of multi-source valid data and the current time node , and set the interval time The threshold Rt, when the interval length If the value is lower than the threshold Rt, the timeliness of the evaluation data indicator is up to standard. The timeliness evaluation coefficient is calculated by accumulating the number of indicators that meet the timeliness standards of all valid data from multiple sources and marking it as Nt. , used to evaluate the timeliness of multi-source valid data.

6. The multi-source heterogeneous data acquisition system based on artificial intelligence according to claim 5, characterized in that: The specific process of generating a regional data quality heat map is as follows: Accuracy evaluation coefficient , completeness assessment coefficient and timeliness evaluation coefficient Perform weighted summation to obtain the data quality index Q; By setting the evaluation interval of the data quality index Q and performing interval comparison, the data quality can be comprehensively evaluated and the data quality level can be determined; The collection area is marked by the installation locations of N0 sensor devices, and the data quality levels of the sensor devices are displayed to generate a regional data quality heat map.

7. The artificial intelligence-based multi-source heterogeneous data acquisition system according to claim 6, characterized in that: By building a digital twin simulation platform, we can monitor the resource distribution of the acquisition system in real time and obtain the usage of computing resources, storage resources, and network resources. The specific process is as follows: By monitoring CPU usage and GPU usage and performing weighted fusion, we can obtain computing resource usage and analyze computing resource usage. By monitoring the disk remaining space ratio r1 and the disk read and write rate r2, and performing weighted fusion, the storage resource remaining rate is obtained, thereby analyzing the storage resource usage; By monitoring the network bandwidth utilization u1 and network delay time u2 and performing weighted fusion, the network resource utilization is obtained, thereby analyzing the usage of network resources.

8. The artificial intelligence-based multi-source heterogeneous data acquisition system according to claim 7, characterized in that: Fault injection testing is performed on the digital twin simulation platform to output a dynamic resource scheduling solution. The specific process is as follows: Fault injection testing includes CPU overload on computing nodes, disk damage on storage devices, and network link interruption. By adjusting the fault type and duration of the sensor equipment, the operating status of the data acquisition system can be observed; By coordinating the allocation of computing resources, storage resources, and network resources, the changes in the data quality index Q are monitored synchronously, and the solution corresponding to the maximum value of the data quality index Q is marked as the dynamic resource scheduling solution.

Citation Information

Cited By

  • Intelligent water quality prediction method and system based on multi-frequency adaptive adjustment

    CN120908404A

  • Automobile extended insurance multi-source heterogeneous data cleaning evaluation system

    CN120929453A