Multi-source data analysis system and method based on artificial intelligence
Through a multi-source data analysis system based on artificial intelligence, the problems of inefficient data acquisition and lack of directionality in the existing technology are solved, and more efficient data preprocessing and abnormal detection are achieved, and more accurate data analysis and decision-making are supported.
Patent Information
- Application Number
- CN202510668213.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-23
AI Technical Summary
In the multi-source data analysis, the existing technology has inefficient data acquisition efficiency, complex access methods, lack of unified and efficient integration methods, and lack of directionality in response to abnormal data, so the most appropriate measures cannot be taken for different abnormal situations.
A multi-source data analysis system based on artificial intelligence is adopted to obtain structured data and metadata from different data sources, perform pre-processing, extract statistical features, set thresholds to judge abnormalities, and build a decision tree classification model based on multi-dimensional features. Classify according to the abnormal index and update the model parameters incrementally, and finally set response rules based on the abnormal classification results.
Improve data quality and consistency, reduce errors and noise in the data, achieve more accurate and comprehensive abnormal detection, can promptly detect various abnormal situations in the entire process, and support more effective data analysis and decision-making.
Smart Images

Figure CN120197071A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and particularly to a multi-source data analysis system and method based on artificial intelligence. Background Art
[0002] In today's digital age, the analysis and processing of multi-source data are crucial for decision-making, business optimization, and risk prevention and control in various industries. Multi-source data covers data in various different formats and sources, such as from databases, formatted text files, standard API interfaces, etc. During the transmission process of this data, the transmission rate and loss rate can affect the integrity and timeliness of the data. Anomaly detection aims to identify data points that are significantly different from most of the data, and these abnormal data may reflect important information such as system failures, abnormal business operations, or data errors. By effectively processing the abnormal data, measures can be taken in a timely manner to correct the problems, ensure the stable operation of the system, and improve business efficiency.
[0003] There are many limitations in the existing technology for multi-source data analysis. In the data acquisition stage, the access methods of different data sources are complex and lack a unified and efficient integration method, resulting in low data acquisition efficiency and easy errors. In addition, the existing technology lacks directionality in the response to abnormal data. When abnormal data is detected, differentiated processing strategies are often not formulated according to factors such as the type and degree of the anomaly, and a unified general processing method is mostly used, and the most appropriate measures cannot be taken for different abnormal situations. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-source data analysis system and method based on artificial intelligence to solve the problems raised in the existing technology.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A multi-source data analysis method based on artificial intelligence, the method includes the following steps: Step 1, obtain structured data and metadata from different data sources, monitor the transmission rate and loss rate, and generate a transmission health report; Step 2, preprocess the metadata, and obtain an integrated data set through conversion; Step 3, extract statistical features, set thresholds to judge abnormal situations, summarize abnormal information to establish an index system, and generate an abnormal report; Step 4, construct a decision tree classification model based on multi-dimensional features, classify according to the anomaly index, and incrementally update the model parameters; Step 5, use the classification model to judge the abnormal situation of new data, set response rules according to the abnormal classification results, push feedback messages and generate logs.
[0006] In step 1, for database data: The staff pre-configures the standard database connection protocol (such as JDBC / ODBC), identifies the data table structure based on the standard database connection protocol according to the data source type, and reads the data; For formatted text files: Use a file parser (such as parsing rules for CSV, JSON, XML) to automatically parse the file content and identify each data field in the file; For standard API interfaces: Build an API call module according to the RESTful protocol, automatically initiate a request, and parse the returned JSON / XML data; Obtain structured data and metadata in a unified format; The metadata description includes field name, data type, collection time t, and data volume m; Among them, m is a positive integer representing the number of metadata records; The field is denoted as X j ; Among them, j = 1, 2,..., n, used to distinguish different fields; n is a positive integer representing the number of metadata description fields; The record is represented as x i ; Among them, i = 1, 2,..., m, used to distinguish different records; Monitor the transmission rate v(t) and the loss rate L: Use a sliding window to calculate the average transmission rate: v bar = 1 / w · Σ w i=1 v(t i ); Among them, w represents the window length; Calculate the packet loss rate, the number of lost packets is represented as N loss , the total number of packets is represented as N total , the loss rate calculation formula is: L = N loss / N total ; Generate a transmission health report H, including: the value range of v bar : and determine whether it reaches the predetermined threshold v min ; The loss rate L: whether it exceeds the set upper limit L max ; In step 2, for metadata, perform missing value processing: For the missing values in the field X j , use the mean filling formula to fill them; Obtain the data matrix after missing value processing and the filling log; For the data matrix after missing value processing, use the Z-score method to detect noise: For each field, calculate the mean and standard deviation, and further calculate the Z-score value; When the Z-score value is the preset threshold, consider this value as abnormal noise data; Obtain the data matrix after noise filtering and the noise anomaly report; Perform format standardization on the data matrix after noise filtering: Convert the date and time data using a fixed format (such as the ISO8601 standard); the original date is represented as d raw , and after conversion, it is denoted as d std =f date (d raw ); where f date represents the date conversion function corresponding to the fixed format; obtain a data matrix with unified format; Based on the data matrix with unified format, use Min - Max normalization to convert the numerical data items to the interval [0, 1]; obtain the normalized data matrix; For the normalized data matrix, perform data matching and fusion: Use the common field timestamp t to perform inner join or outer join; generate the integrated dataset; In step 3, extract statistical features: For each standardized field X j ', calculate the mean μ j , standard deviation σ j , maximum value max(X j ') and minimum value min(X j '); obtain the index set I = {μ j , σ j , max(X j '), min(X j ')}; Set a threshold to judge anomalies: When the standardized data record x ij ' is not within the range of [μ j - k·σ j , μ j + k·σ j , it is regarded as an anomaly; where k is a positive integer representing the preset coefficient; For data transmission anomalies: Use the transmission rate v(t) as the signal, and perform mean filtering or median filtering to judge the mutation points in the time series; calculate the first - order difference Δv(t i ) = v(t i ) - v(t i-1 ), when |Δv(ti)| > γ, it is regarded as a transmission anomaly point; where γ is the set threshold; obtain the set A trans of transmission anomaly time points; For source data anomalies: Use the Isolation Forest to calculate the anomaly score s(x i ); the anomaly score is determined based on the distance of the data in the feature space. When s(x i ) > δ, it is judged as an anomaly; where δ represents the preset threshold; obtain the set A source of source data anomaly records and the corresponding anomaly levels; For data conversion anomalies: Compare the converted data X’ with the expected value X* described by the original metadata, and calculate the semantic consistency index: S = 1 - ||X’ - X*|| / ||X*||; When S < λ, it is judged as a data conversion anomaly; where λ is the set minimum alignment degree; Mark and generate a set A of conversion anomaly records conv ; The expected value represents the data state or value range expected according to the definitions, constraints, business rules, etc. of the original metadata; Summarize various anomaly scores and records, and establish an anomaly index system E = {A trans , A source , A conv}; Generate an anomaly index E(x i ) for each data record xi, and the calculation formula is: E(x i ) = w1·F(x i ∈ A trans ) + w2·F(x i ∈ A source ) + w3·F(x i ∈ A conv ); Among them, F is an index function, which takes the value of 1 when the condition in the parentheses holds, and 0 otherwise; w1, w2, and w3 are the weights of each type of anomaly, 0 < w1, w2, w3 < 1 and w1 + w2 + w3 = 1; Generate an anomaly report and classification labels.
[0007] In step 4, construct a classification model based on multi-dimensional features: Using the branch structure of the decision tree, compare the anomaly index E(x i ) of the input data with the preset category thresholds τ1, τ2: When E(x i ) < τ1, the category is normal; when τ1 ≤ E(x i ) < τ2, the category is warning; when E(x i ) ≥ τ2, the category is abnormal; Among them, τ1 and τ2 are preset category thresholds, and τ1 < τ2; Establish an anomaly classification model M, and generate an anomaly index for each new data; Compare the prediction result of model M with the actual feedback, adjust the parameters and classification thresholds, and realize the incremental update of the model.
[0008] In step 5, for new data, after preprocessing, determine its anomaly level through model M; Set response rules according to the anomaly classification results: When E(x i ) ≥ τ2, automatically trigger an alarm and retransmit the data; When E(x i) When it is in the range of [τ1, τ2), it is marked for manual review; the feedback message is pushed and an automatic processing log is generated.
[0009] A multi-source data analysis system based on artificial intelligence, which includes a data access module, a data preprocessing module, an intelligent analysis module, a pattern recognition module, and a result output module; The data access module is used to obtain structured data and metadata from different data sources, monitor the transmission rate and loss rate, and generate a transmission health report; the data preprocessing module is used to preprocess the metadata and obtain an integrated data set through conversion; the intelligent analysis module is used to extract statistical features, set thresholds to judge abnormal situations, summarize abnormal information to establish an index system, and generate an abnormal report; the pattern recognition module is used to build a decision tree classification model based on multi-dimensional features, classify according to the abnormal index, and incrementally update the model parameters; the result output module is used to judge the abnormal situation of new data according to the classification model, set response rules according to the abnormal classification result, push feedback messages and generate logs.
[0010] The data access module includes a database unit, a file parsing unit, an API call unit, and a transmission monitoring unit; The database unit is used to connect to the database through a protocol and extract structured data; the file parsing unit is used to parse files and identify fields and data content; the API call unit is used to call interfaces according to the RESTful protocol and parse JSON / XML responses; the transmission monitoring unit is used to calculate the transmission rate and loss rate in real time and generate a health report.
[0011] The data preprocessing module includes a missing value filling unit, a noise filtering unit, a format standardization unit, and a data fusion unit; The missing value filling unit is used to process field missing values by the mean imputation method and record the filling log; the noise filtering unit is used to identify noise data based on Z-score and generate a noise abnormal report; the format standardization unit is used to unify the date format and normalize numerical data; the data fusion unit is used to associate multi-source data through timestamps and generate an integrated data set.
[0012] The intelligent analysis module includes a feature extraction unit, an abnormal detection unit, and an index summarization unit; The feature extraction unit is used to calculate the field mean, standard deviation, and extreme values; the abnormal detection unit is used to identify three types of abnormalities through thresholds, isolation forests, and difference methods; the index summarization unit is used to build an abnormal index system and generate a comprehensive abnormal index by weighting.
[0013] The pattern recognition module includes a parameter update unit and a model construction unit; The parameter update unit is used to dynamically adjust the classification threshold according to the actual feedback; the model construction unit is used to divide the abnormal categories based on the decision tree; The result output module includes a classification determination unit, a response rule unit, and a log generation unit; The classification determination unit is used to output the abnormal classification label of the new data; the response rule unit is used to trigger the alarm, retransmission, or manual review process; the log generation unit is used to record the abnormal classification result and the system response action.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: Through the preprocessing of multi-source data, including operations such as missing value filling, noise filtering, format standardization, and normalization, the present invention can effectively improve the quality and consistency of the data, reduce errors and noise in the data, and provide a reliable data basis for subsequent data analysis and decision-making. Compared with the traditional single data preprocessing method, it can handle more complex and diverse data problems; The present invention comprehensively considers various abnormal situations and establishes a comprehensive abnormal index system, making the abnormal detection more accurate and comprehensive; It can not only detect the abnormalities of the data itself, but also discover problems in the data transmission and conversion processes. Compared with the prior art that only focuses on the abnormalities of the data itself, it can more timely and accurately discover various abnormal situations of the data in the entire process, which helps to take measures in a timely manner for correction and processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic diagram of the steps of a multi-source data analysis method based on artificial intelligence according to the present invention; Figure 2 It is a schematic diagram of the process of a multi-source data analysis system based on artificial intelligence according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0017] Embodiment: As Figure 1 - Figure 2 shown, the present invention provides a technical solution, a multi-source data analysis method based on artificial intelligence, and the method includes the following steps: Step 1: Obtain structured data and metadata from different data sources, monitor the transmission rate and loss rate, and generate a transmission health report; Step 2: Preprocess the metadata to obtain an integrated data set through conversion; Step 3: Extract statistical features, set thresholds to judge abnormal situations, summarize abnormal information to establish an index system, and generate an abnormal report; Step 4: Construct a decision tree classification model based on multi-dimensional features, classify according to the abnormal index, and incrementally update the model parameters; Step 5: Use the classification model to judge the abnormal situation of new data, set response rules according to the abnormal classification results, push feedback messages and generate logs.
[0018] In Step 1, for database data: The staff pre-configures a standard database connection protocol (such as JDBC / ODBC), and based on the standard database connection protocol, identifies the data table structure and reads the data according to the data source type; For formatted text files: Use a file parser (such as parsing rules for CSV, JSON, XML) to automatically parse the file content and identify each data field in the file; For standard API interfaces: Construct an API call module according to the RESTful protocol, automatically initiate a request and parse the returned JSON / XML data; Obtain structured data and metadata in a unified format; The metadata description includes field name, data type, collection time t, and data volume m; Among them, m is a positive integer representing the number of metadata records; The field is denoted as X j ; Among them, j = 1, 2, …, n, used to distinguish different fields; n is a positive integer representing the number of metadata description fields; The record is represented as x i ; Among them, i = 1, 2, …, m, used to distinguish different records; Monitor the transmission rate v(t) and the loss rate L: Use a sliding window to calculate the average transmission rate: v bar = 1 / w·Σ w i=1 v(t i ); Among them, w represents the window length; Calculate the packet loss rate, the number of lost packets is denoted as N loss , the total number of data packets is denoted as N total , and the loss rate calculation formula is: L = N loss / N total ; Generate a transmission health report H, including: the value range of v bar : and judge whether it reaches the predetermined threshold v min ; The loss rate L: whether it exceeds the set upper limit L max ; In Step 2, for metadata, perform missing value processing: For the field X jThe missing values are filled using the mean filling formula; the data matrix after missing value processing and the filling log are obtained; For the data matrix after missing value processing, the Z-score method is used to detect noise: for each field, the mean and standard deviation are calculated, and the Z-score value is further calculated; when the Z-score value is outside the preset threshold, the value is considered abnormal noise data; the data matrix after noise filtering and the noise anomaly report are obtained; Format standardization is performed on the data matrix after noise filtering: date and time data are converted using a fixed format (such as the ISO8601 standard); the original date is represented as d raw , and after conversion, it is denoted as d std =f date (d raw ); where f date represents the date conversion function corresponding to the fixed format; the data matrix with unified format is obtained; Based on the data matrix with unified format, Min-Max normalization is used to convert numerical data items to the [0,1] interval; the normalized data matrix is obtained; For the normalized data matrix, data matching and fusion are performed: using the common field timestamp t, an inner join or outer join is performed; an integrated dataset is generated; In step 3, statistical features are extracted: for each standardized field X j ' the mean μ j , standard deviation σ j , maximum value max(X j ') and minimum value min(X j ') are calculated; the index set I={μ j ,σ j ,max(X j '),min(X j ')} is obtained; Set a threshold to judge anomalies: when the standardized data record x ij ' is not within the range of [μ j -k·σ j ,μ j +k·σ j , it is regarded as an anomaly; where k is a positive integer representing the preset coefficient; For data transmission anomalies: using the transmission rate v(t) as the signal, mean filtering or median filtering is used to process and judge the mutation points in the time series; calculate the first-order difference Δv(t i )=v(t i )-v(t i-1), when |Δv(ti)| > γ, it is regarded as a transmission anomaly point; where γ is a set threshold; obtain the set A of transmission anomaly time points trans ; For source data anomalies: Use Isolation Forest to calculate the anomaly score s(x i ); The anomaly score is determined based on the distance of the data in the feature space. When s(x i ) > δ, it is judged as an anomaly; where δ represents a preset threshold; obtain the set A of source data anomaly records source and the corresponding anomaly levels; For data conversion anomalies: Compare the converted data X’ with the expected value X* described by the original metadata, and calculate the semantic consistency index: S = 1 - ||X’ - X*|| / ||X*||; When S < λ, it is judged as a data conversion anomaly; where λ is the set minimum alignment degree; mark and generate the set A of conversion anomaly records conv ; Summarize various anomaly scores and records, and establish an anomaly index system E = {A trans , A source , A conv}; Generate an anomaly index E(x i ) for each data record xi, and the calculation formula is: E(x i ) = w1·F(x i ∈ A trans ) + w2·F(x i ∈ A source ) + w3·F(x i ∈ A conv ); Among them, F is an index function, which takes the value of 1 when the condition in the parentheses holds, otherwise it takes the value of 0; w1, w2, w3 are the weights of each type of anomaly, 0 < w1, w2, w3 < 1 and w1 + w2 + w3 = 1; Generate an anomaly report and classification labels.
[0019] In step 4, construct a classification model based on multi-dimensional features: Use the branch structure of the decision tree to compare the anomaly index E(x i ) of the input data with the preset category thresholds τ1, τ2: When E(x i ) < τ1, the category is normal; when τ1 ≤ E(x i ) < τ2, the category is warning; when E(x i ) ≥ τ2, the category is abnormal; Among them, τ1, τ2 are preset category thresholds, and τ1 < τ2; Establish an anomaly classification model M and generate an anomaly index for each new data; Compare the predicted results of model M with the actual feedback, adjust the parameters and classification thresholds, and achieve incremental update of the model.
[0020] In step 5, for new data, after preprocessing, determine its anomaly level through model M; set response rules according to the anomaly classification results: when E(x i )≥τ2, automatically trigger an alarm and retransmit the data; when E(x i )∈[τ1,τ2), mark it for manual review; push the feedback message and generate an automatic processing log.
[0021] A multi-source data analysis system based on artificial intelligence, which includes a data access module, a data preprocessing module, an intelligent analysis module, a pattern recognition module, and a result output module; The data access module is used to obtain structured data and metadata from different data sources, monitor the transmission rate and loss rate, and generate a transmission health report; the data preprocessing module is used to preprocess the metadata and obtain an integrated data set through conversion; the intelligent analysis module is used to extract statistical features, set thresholds to judge abnormal situations, summarize abnormal information to establish an index system, and generate an abnormal report; the pattern recognition module is used to construct a decision tree classification model based on multi-dimensional features, classify according to the anomaly index, and incrementally update the model parameters; the result output module is used to judge the anomaly situation of new data according to the classification model, set response rules according to the anomaly classification results, push feedback messages and generate logs.
[0022] The data access module includes a database unit, a file parsing unit, an API call unit, and a transmission monitoring unit; The database unit is used to connect to the database through a protocol and extract structured data; the file parsing unit is used to parse files and identify fields and data content; the API call unit is used to call interfaces according to the RESTful protocol and parse JSON / XML responses; the transmission monitoring unit is used to calculate the transmission rate and loss rate in real time and generate a health report.
[0023] The data preprocessing module includes a missing value filling unit, a noise filtering unit, a format standardization unit, and a data fusion unit; The missing value filling unit is used to process field missing values by the mean imputation method and record the filling log; the noise filtering unit is used to identify noise data based on Z-score and generate a noise anomaly report; the format standardization unit is used to unify the date format and normalize numerical data; the data fusion unit is used to associate multi-source data through timestamps and generate an integrated data set.
[0024] The intelligent analysis module includes a feature extraction unit, an anomaly detection unit, and an index summarization unit; The feature extraction unit is used to calculate the field mean, standard deviation, and extreme values; the anomaly detection unit is used to identify three types of anomalies through thresholds, isolation forests, and difference methods; the metric summarization unit is used to construct an anomaly metric system and generate a comprehensive anomaly index by weighting.
[0025] The pattern recognition module includes a parameter update unit and a model construction unit; The parameter update unit is used to dynamically adjust the classification threshold according to actual feedback; the model construction unit is used to divide anomaly categories based on decision trees; The result output module includes a classification determination unit, a response rule unit, and a log generation unit; The classification determination unit is used to output the anomaly classification label of new data; the response rule unit is used to trigger alarm, retransmission, or manual review processes; the log generation unit is used to record the anomaly classification results and system response actions.
[0026] In this embodiment, an e-commerce enterprise needs to analyze multi-source data such as sales data and user evaluation data; Step 1: Data acquisition and transmission monitoring; Database data: The staff pre-configures the JDBC protocol to connect to the enterprise's sales database. Based on this protocol, the sales data table structure is identified, and sales data containing fields such as order number, product ID, sales quantity, sales amount, and order time is read; Formatted text file: The enterprise regularly receives a CSV-format logistics distribution information file provided by a cooperative logistics company. A file parser for CSV is used to automatically parse the file content and identify data fields such as package number, shipping address, receiving address, and delivery time; Standard API interface: An API call module is built according to the RESTful protocol to call the API of a third-party user evaluation platform; requests are automatically initiated and the returned JSON-format data is parsed to obtain information such as user evaluation content, score, and evaluation time for products.
[0027] Metadata generation: While obtaining structured data from the above data sources, metadata in a unified format is generated. For the metadata of sales data, the field names include "order number", "product ID", etc., the data type such as "order number" is string type, the acquisition time t records the moment of each data acquisition, and the data volume m counts the number of sales records obtained this time.
[0028] Transmission Monitoring: Calculate the average transmission rate vbar using a sliding window. The window length w is 10 minutes, and the transmission rate v(t) is recorded every 1 minute. Calculate the average value of the transmission rates at these 10 time points; Calculate the packet loss rate L by counting the number of lost packets Nloss and the total number of packets Ntotal, and obtain the loss rate. Generate a transmission health report H. The average transmission rate vbar is 10, and the predetermined threshold vmin is 8 MB / s, meeting the requirements; The loss rate L is 0.5%, and the set upper limit Lmax is 1%, not exceeding the upper limit.
[0029] Step 2: Metadata Preprocessing; Missing Value Handling: For the "Sales Amount" field Xj in the sales data, if there are missing values, calculate the mean of the existing values in this field and use this mean to fill in the missing values, and record the filling log.
[0030] Noise Detection: Use the Z-score method to detect noise in the sales data. Calculate the mean and standard deviation of the "Sales Quantity" field, and further calculate the Z-score value; When the Z-score value exceeds the preset threshold, the corresponding data is regarded as abnormal noise data, obtaining the noise-filtered data matrix and the noise anomaly report; Format Standardization: For the date data draw in the order, use the date conversion function fdate corresponding to the ISO8601 standard to convert it to the fixed format dstd, converting "2023 / 10 / 15" to "2023-10-15T00:00:00Z".
[0031] Normalization: For numerical data items such as sales amount, use Min-Max normalization to convert them to the [0,1] interval.
[0032] Data Matching and Fusion: Use the common field "Order Placement Time" t to perform an inner join on the sales data and the user evaluation data to generate an integrated dataset, associating the sales orders at the same time point with the corresponding user evaluations.
[0033] Step 3: Statistical Feature Extraction and Anomaly Judgment; Statistical Feature Extraction: For the standardized "Sales Quantity" field Xj', calculate the mean μj, standard deviation σj, maximum value max(Xj'), and minimum value min(Xj'), obtaining the index set I.
[0034] Set Threshold to Judge Anomaly: When the sales quantity xij' of a certain sales record after standardization is not within the range of [μj - k・σj, μj + k・σj], it is regarded as abnormal; k = 2, μj = 5, σj = 1, and if the sales quantity of a certain record is 8, it is regarded as abnormal.
[0035] Data transmission anomaly: Using the transmission rate v(t) as the signal, mean filtering is performed. Calculate the first-order difference Δv(ti)=v(ti)-v(ti-1). When |Δv(ti)|>γ, it is regarded as a transmission anomaly point. γ = 2. If at time v(ti)=12 and v(ti-1)=9, |Δv(ti)| = 3>2, this time point is regarded as a transmission anomaly point, and the set Atrans of transmission anomaly time points is obtained.
[0036] Source data anomaly: The isolation forest is used to calculate the anomaly score s(xi) of each record of the user evaluation data. When s(xi)>δ, it is judged as an anomaly. δ = 0.8. For a certain evaluation record, the anomaly score s(xi)=0.9, which is judged as a source data anomaly, and the set Asource of source data anomaly records and the corresponding anomaly levels are obtained.
[0037] Data conversion anomaly: Compare the converted sales data X’ with the expected value X* described by the original metadata, and calculate the semantic consistency index S. Assuming λ = 0.9, when S<λ, it is judged as a data conversion anomaly, and the conversion anomaly record set Aconv is marked and generated.
[0038] Summarize anomaly information: Summarize various anomaly scores and records, and establish an anomaly index system E. Generate an anomaly index E(xi) for each sales record xi. w1 = 0.3, w2 = 0.4, w3 = 0.3. If a certain record belongs to the source data anomaly, E(xi)=0.3×0 + 0.4×1 + 0.3×0 = 0.4.
[0039] Generate anomaly reports and classification labels: Generate anomaly reports based on the anomaly index, and label the anomaly data with classification labels.
[0040] Step 4: Construct and update the decision tree classification model; Construct the classification model: Using the branch structure of the decision tree, compare the anomaly index E(xi) of the input sales data with the preset category thresholds τ1 = 0.3 and τ2 = 0.6. When E(xi)<0.3, the category is normal; when 0.3≤E(xi)<0.6, the category is warning; when E(xi)≥0.6, the category is abnormal. The anomaly classification model M is established.
[0041] Model update: Compare the prediction results of the model M with the actual feedback. If it is found that the actual anomaly data is misjudged as normal within a certain period of time, adjust the model parameters and classification thresholds to achieve incremental update of the model.
[0042] Step 5: Process and respond to new data; Abnormal level determination: For the newly obtained sales data, after the same preprocessing as in Step 2, its abnormal level is determined by Model M. Set the response rules: When E(xi) ≥ 0.6, an alarm is automatically triggered and the data is retransmitted; when E(xi) ∈ [0.3, 0.6), it is marked for manual review. Push the feedback message and generate an automatic processing log to record the time of a certain alarm, the content of the alarm data, etc.
[0043] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced by the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
Claims
1. A multi-source data analysis method based on artificial intelligence, characterized in that: The method includes the following steps: Step 1: Obtain structured data and metadata from different data sources, monitor the transmission rate and loss rate, and generate a transmission health report; Step 2: Preprocess the metadata and obtain an integrated data set through conversion; Step 3: Extract statistical features, set thresholds to judge abnormal situations, summarize abnormal information to establish an index system, and generate an abnormal report; Step 4: Construct a decision tree classification model based on multi-dimensional features, classify according to the abnormal index, and incrementally update the model parameters; Step 5: Use the classification model to judge the abnormal situation of new data, set response rules according to the abnormal classification results, push feedback messages and generate logs.
2. The multi-source data analysis method based on artificial intelligence according to claim 1, wherein: In Step 1, for database data: The staff pre-configures the standard database connection protocol, identifies the data table structure based on the standard database connection protocol according to the data source type, and reads the data; for formatted text files: Use a file parser to automatically parse the file content and identify each data field in the file; For standard API interfaces: Construct an API call module according to the RESTful protocol, automatically initiate requests and parse the returned JSON / XML data; Obtain structured data and metadata in a unified format; The metadata description includes the field name, data type, collection time t, and data volume m; where m is a positive integer representing the number of metadata records; The field is represented as X j ; The record is represented as x i ; where j = 1, 2, …, n, which is used to distinguish different fields; n is a positive integer representing the number of metadata description fields; where i = 1, 2, …, m, which is used to distinguish different records Monitor the transmission rate v(t) and the loss rate L: Use a sliding window to calculate the average transmission rate v bar : v bar = 1 / w·Σ w i=1 v(t i ); where w represents the window length; v(t i ) represents the transmission rate corresponding to the acquisition time of the i-th record; Calculate the data packet loss rate. The number of lost packets is denoted as N loss , and the total number of data packets is denoted as N total . The formula for calculating the loss rate is: L = N loss / N total ; Generate a transmission health report H, including: v bar Value range: and determine whether the preset threshold v is reached min ; Loss rate L: whether it exceeds the preset upper limit L max .
3. The multi-source data analysis method based on artificial intelligence according to claim 2, characterized in that: In step 2, for the metadata, missing value processing is performed: for the missing values in field X j the mean filling formula is used for filling; Obtain the data matrix after missing value processing and the filling log; For the data matrix after missing value processing, use the Z-score method to detect noise: For each field, calculate the mean and standard deviation, and further calculate the Z-score value; when the Z-score value is preset threshold, consider this value as abnormal noise data; Obtain the data matrix after noise filtering and the noise abnormal report; Format standardization is performed on the data matrix after noise filtering: the date and time data are converted using a fixed format; the original date is represented as d raw , and after conversion, it is denoted as d std : d std = f date (d raw ); where f date represents the date conversion function corresponding to the fixed format; a data matrix with a unified format is obtained; Based on the data matrix with unified format, adopt Min-Max normalization to convert numerical data items to the interval [0,1]; obtain the normalized data matrix; For the normalized data matrix, perform data matching and fusion: Use the common field timestamp t to perform inner join or outer join; generate an integrated data set.
4. The multi-source data analysis method based on artificial intelligence according to claim 3, characterized in that: In step 3, extract statistical features: for each standardized field X j 'Calculate the mean μ j , standard deviation σ j , maximum value max(X j '), and minimum value min(X j '); obtain the index set I = {μ j , σ j , max(X j '), min(X j ')}; Set threshold to judge abnormality: When the standardized data record x ij ' is not within [μ j - k·σ j , μ j + k·σ j , it is regarded as abnormal; where k is a positive integer representing the preset coefficient; For data transmission anomalies: using the transmission rate v(t) as a signal, mean filtering or median filtering is employed to process and determine the mutation points in the time series; calculate the first-order difference Δv(t i ) = v(t i ) - v(t i-1 ), when |Δv(ti)| > γ, it is regarded as a transmission anomaly point; where γ is the set threshold; obtain the set A of transmission anomaly time points trans ; For source data anomalies: The isolation forest is used to calculate the anomaly score s(x i ); The anomaly score is determined based on the distance of the data in the feature space. When s(x i ) > δ, it is judged as an anomaly; where δ represents a preset threshold; The source data anomaly record set A source and the corresponding anomaly levels are obtained. For data conversion anomalies: Compare the converted data X’ with the expected value X* described by the original metadata, and calculate the semantic consistency index S: S = 1 - ||X’ - X*|| / ||X*||; when S < λ, it is judged as a data conversion anomaly; where λ is the set minimum alignment degree; mark and generate a set A of conversion anomaly records conv ; Summarize various anomaly scores and records, and establish an anomaly index system \(E = \{A trans , A source , A conv \}\); Generate an anomaly index \(E(x i )\) for each data record \(x_i\), and the calculation formula is: \(E(x i ) = w_1·F(x i \in A trans ) + w_2·F(x i \in A source ) + w_3·F(x i \in A conv )\); where \(F\) is an index function, which takes the value of 1 when the condition in the parentheses holds, and 0 otherwise; \(w_1, w_2, w_3\) are the weights of each type of anomaly, \(0 < w_1, w_2, w_3 < 1\) and \(w_1 + w_2 + w_3 = 1\); Generate an abnormal report and classification labels.
5. A multi-source data analysis method based on artificial intelligence according to claim 4, characterized in that: In step 4, a classification model based on multi-dimensional features is constructed: using the branch structure of a decision tree, the anomaly index E(x i ) of the input data is compared with the preset class thresholds τ1 and τ2: when E(x i ) < τ1, the class is normal; when τ1 ≤ E(x i ) < τ2, the class is warning; when E(x i ) ≥ τ2, the class is abnormal; where τ1 and τ2 are preset class thresholds, and τ1 < τ2; Establish an abnormal classification model M and generate an abnormal index for each new data; Compare the prediction result of model M with the actual feedback, adjust the parameters and classification thresholds, and realize the incremental update of the model; In step 5, for new data, after preprocessing, it is classified by model M; response rules are set according to the abnormal classification results: when E(x i ) ≥ τ2, an alarm is automatically triggered and the data is retransmitted; when E(x i ) ∈ [τ1, τ2), it is marked for manual review; the feedback message is pushed and an automatic processing log is generated.
6. A multi-source data analysis system based on artificial intelligence, which is applied to any one of the multi-source data analysis methods based on artificial intelligence described in claims 1-5, and is characterized in that: The system includes a data access module, a data preprocessing module, an intelligent analysis module, a pattern recognition module, and a result output module; The data access module is used to obtain structured data and metadata from different data sources, monitor the transmission rate and loss rate, and generate a transmission health report; the data preprocessing module is used to preprocess the metadata and obtain an integrated data set through conversion; the intelligent analysis module is used to extract statistical features, set thresholds to judge abnormal situations, summarize abnormal information to establish an index system, and generate an abnormal report; the pattern recognition module is used to construct a decision tree classification model based on multi-dimensional features, classify according to the abnormal index, and incrementally update the model parameters; the result output module is used to judge the abnormal situation of new data according to the classification model, set response rules according to the abnormal classification results, push feedback messages and generate logs.
7. An artificial intelligence-based multi-source data analysis system according to claim 6, characterized in that: The data access module includes a database unit, a file parsing unit, an API call unit, and a transmission monitoring unit; The database unit is used to connect to the database through a protocol and extract structured data; the file parsing unit is used to parse files and identify fields and data content; The API call unit is used to call interfaces according to the RESTful protocol and parse JSON / XML responses; the transmission monitoring unit is used to calculate the transmission rate and loss rate in real time and generate a health report.
8. An artificial intelligence-based multi-source data analysis system according to claim 7, characterized in that: The data preprocessing module includes a missing value filling unit, a noise filtering unit, a format standardization unit, and a data fusion unit; The missing value filling unit is used to process field missing values by the mean imputation method and record the filling log; The noise filtering unit is used to identify noise data based on Z-score and generate a noise anomaly report; The format standardization unit is used to unify the date format and normalize numerical data; the data fusion unit is used to associate multi-source data through timestamps and generate an integrated dataset.
9. An artificial intelligence-based multi-source data analysis system according to claim 8, characterized in that: The intelligent analysis module includes a feature extraction unit, an anomaly detection unit, and an index summarization unit; The feature extraction unit is used to calculate the field mean, standard deviation, and extreme values; the anomaly detection unit is used to identify three types of anomalies through thresholds, isolation forests, and difference methods; The index summarization unit is used to construct an anomaly index system and generate a comprehensive anomaly index by weighting.
10. A multi-source data analysis system based on artificial intelligence according to claim 9, characterized in that: The pattern recognition module includes a parameter update unit and a model construction unit; The parameter update unit is used to dynamically adjust the classification threshold according to actual feedback; the model construction unit is used to divide anomaly categories based on decision trees; The result output module includes a classification determination unit, a response rule unit, and a log generation unit; The classification determination unit is used to output the anomaly classification label of new data; The response rule unit is used to trigger alarm, retransmission, or manual review processes; the log generation unit is used to record the anomaly classification results and system response actions.
Citation Information
Patent Citations
Neural network algorithm-based network security spatial data asset threat identification method
CN118820949A
Intelligent data analysis method based on space-time big data
CN119494018A
Virtualized network fault rapid positioning method fusing multi-source information
CN119520229A
Intelligent cold chain management method and system based on Internet and multi-dimensional analysis
CN119850069A
Power communication equipment anomaly detection method based on decision tree
CN119880021A
Cited By
Drainage system abnormal working condition identification method based on data fusion model
CN121278778A
A method for identifying abnormal operating conditions in drainage systems based on data fusion models
CN121278778B