Multi-source heterogeneous data intelligent fusion and analysis system based on cloud computing

Through the intelligent fusion and analysis system of multi-source heterogeneous data based on cloud computing, the problem of low efficiency of multi-source heterogeneous data fusion analysis in the existing technology is solved, and efficient data processing and accurate analysis results are achieved.

CN120067986APending Publication Date: 2025-05-30BEIJING ZHIYUAN XUANDA TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510141333.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is inefficient in the fusion and analysis of multi-source heterogeneous data, and it is difficult to meet the requirements of efficient fusion, cleaning, analysis and real-time processing of data.

Method used

A multi-source heterogeneous data intelligent fusion and analysis system based on cloud computing is adopted, which includes data acquisition module, association module, fusion module, multiple analysis modules and comprehensive analysis modules. By preprocessing the data, extracting key features, establishing correlation maps, calculating fusion features using attention mechanisms, and conducting comprehensive analysis through multiple analytical methods to finally generate a comprehensive analysis report.

Benefits of technology

It improves the processing efficiency and analysis accuracy of multi-source heterogeneous data, enhances the consistency and analysis dimensions of data, thereby improving the quality and efficiency of data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067986A_ABST
    Figure CN120067986A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source heterogeneous data intelligent fusion and analysis system based on cloud computing, and the system comprises a data obtaining module which is used for obtaining a plurality of pieces of multi-source original data, and carrying out the preprocessing of the plurality of pieces of multi-source original data, and obtaining a plurality of pieces of multi-source heterogeneous data; the association module is used for performing feature extraction on the plurality of multi-source heterogeneous data to obtain a plurality of key features, and establishing an association map according to the plurality of key features; the fusion module is used for calculating to obtain fusion features by adopting an attention mechanism according to the association map; the multi-item analysis module is used for obtaining a plurality of analysis results by adopting a plurality of analysis methods according to the fusion features; and the comprehensive analysis module is used for inputting the plurality of analysis results into a preset analysis model to obtain a comprehensive analysis report. According to the method, the association map is established through the key features, the consistency among multi-source heterogeneous data is improved, the data processing efficiency is improved, then the association map is weighted through an attention mechanism, and the accuracy of feature fusion is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and particularly to an intelligent fusion and analysis system for multi-source heterogeneous data based on cloud computing. Background Art

[0002] With the development of information technology, data from different fields, different formats, and different structures (i.e., multi-source heterogeneous data) is increasing day by day. Such data includes structured data (such as database tables), semi-structured data (e.g., JSON / XML files), and unstructured data (e.g., text, pictures, videos, etc.). Traditional data processing and analysis methods are difficult to meet the requirements of multi-source heterogeneous data in terms of fusion, cleaning, analysis, and real-time processing.

[0003] However, the existing technologies still have deficiencies in data fusion efficiency, analysis accuracy, and system flexibility, and there is an urgent need for a new solution to solve these problems. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide an intelligent fusion and analysis system for multi-source heterogeneous data based on cloud computing to solve the technical problem of low efficiency in the fusion analysis process of existing multi-source heterogeneous data.

[0005] To achieve the above object, the present invention provides an intelligent fusion and analysis system for multi-source heterogeneous data based on cloud computing, including:

[0006] A data acquisition module, configured to acquire a plurality of multi-source raw data and preprocess the plurality of multi-source raw data to obtain a plurality of multi-source heterogeneous data;

[0007] An association module, configured to extract features from the plurality of multi-source heterogeneous data to obtain a plurality of key features, and establish an association graph according to the plurality of key features;

[0008] A fusion module, configured to calculate a fusion feature according to the association graph by using an attention mechanism;

[0009] A multi-analysis module, configured to obtain a plurality of analysis results by using a plurality of analysis methods according to the fusion feature;

[0010] A comprehensive analysis module, configured to input the plurality of analysis results into a preset analysis model to obtain a comprehensive analysis report.

[0011] The above technical solution has the following beneficial technical effects:

[0012] The present invention first preprocesses the acquired multi-source raw data to reduce redundant data, obtaining multi-source heterogeneous data. Then, it extracts key features from the multi-source heterogeneous data, establishes an association graph through the key features, represents the multi-source heterogeneous data as points in the association graph, and represents the connections between the multi-source heterogeneous data as edges, improving the consistency between the multi-source heterogeneous data, thereby enhancing the data processing efficiency. Subsequently, the association graph is weighted through an attention mechanism to improve the accuracy of the fusion features. The subsequent analysis process is comprehensively analyzed by multiple methods, and finally, a comprehensive analysis report is obtained based on the results of the multiple-method analysis, increasing the dimension of the analysis and thus improving the analysis accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention. Among them:

[0014] Figure 1 is a structural block diagram of a multi-source heterogeneous data intelligent fusion and analysis system based on cloud computing according to the present invention;

[0015] Figure 2 is a structural block diagram of the data acquisition module in an embodiment of the present invention;

[0016] Figure 3 is a structural block diagram of the association module in an embodiment of the present invention;

[0017] Figure 4 is a structural block diagram of the fusion module in an embodiment of the present invention;

[0018] Figure 5 is a structural block diagram of the multiple analysis module in an embodiment of the present invention;

[0019] Figure 6 is a structural block diagram of the comprehensive analysis module in an embodiment of the present invention;

[0020] Figure 7 is a flowchart of a multi-source heterogeneous data intelligent fusion and analysis method based on cloud computing according to the present invention;

[0021] Figure 8 is a structural schematic diagram of a computer system in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description below omits the description of well-known functions and structures.

[0023] Embodiment 1

[0024] As Figure 1 shown, an embodiment of the present invention provides an intelligent fusion and analysis system for multi-source heterogeneous data based on cloud computing, including: a data acquisition module, an association module, a fusion module, a multi-analysis module, and a comprehensive analysis module.

[0025] The data acquisition module is used to acquire a number of multi-source raw data and preprocess the number of multi-source raw data to obtain a number of multi-source heterogeneous data;

[0026] The association module is used to extract features from a number of the multi-source heterogeneous data to obtain a number of key features, and establish an association graph according to the number of key features;

[0027] The fusion module is used to calculate a fusion feature by using an attention mechanism according to the association graph;

[0028] The multi-analysis module is used to obtain a number of analysis results by using a number of analysis methods according to the fusion feature;

[0029] The comprehensive analysis module is used to input a number of the analysis results into a preset analysis model to obtain a comprehensive analysis report.

[0030] Specifically, as Figure 2 shown, the data acquisition module includes: an access recognition unit, a basic processing unit, a semantic alignment unit, and an output unit.

[0031] Specifically, the access recognition unit is used to obtain raw data from various data sources and identify the raw data by using dynamic protocol adaptation technology to obtain a number of multi-source raw data. The data sources include Internet of Things device sensors, Web API (Application Programming Interface) interfaces, databases (such as SQL or NoSQL), and file storage (such as CSV, JSON, XML, etc.). The dynamic protocol adaptation technology is used for automatic recognition of data source types and transmission protocols. For example, when obtaining sensor data from Internet of Things devices, the system automatically recognizes the Message Queuing Telemetry Transport (MQTT) protocol, parses HTTP / HTTPS data streams from Web APIs, and at the same time uses database drivers to adapt to different databases (such as MySQL, PostgreSQL, MongoDB). In terms of technical implementation, the system introduces a rule-based protocol parser and a machine learning model. The parser identifies the data type by analyzing data meta-information (such as header fields or metadata), and the model identifies the data structure through the feature patterns of data samples (such as distinguishing JSON and XML formats). In the access recognition unit, the data structure type of the raw data is first identified, and the structure types of the raw data include structured data, semi-structured data, and unstructured data. For example, in the customer behavior analysis of a bank, the system extracts information from ATM transaction logs (structured data), customer service chat records (unstructured data), and market change data (semi-structured data) obtained from third-party APIs).

[0032] Specifically, the basic processing unit is used to perform data cleaning, normalization, and data denoising on a number of the multi-source raw data to obtain a number of first data. In the basic processing unit, there are: an outlier removal subunit, a filling subunit, a standardization subunit, and a denoising subunit. The outlier removal subunit is used to calculate the interquartile range of a number of the multi-source raw data, calculate the outlier range according to the interquartile range, and screen out a number of outlier data from a number of the multi-source raw data according to the outlier range and the interquartile range, and delete a number of the outlier data to obtain a number of outlier-removed data.

[0033] The calculation formula of the interquartile range is as follows:

[0034] IQR = Q3 - Q1;

[0035] In the formula, IQR represents the interquartile range, Q1 represents the first quartile, Q3 represents the third quartile, and the first quartile and the third quartile can be directly obtained from the multi-source raw data.

[0036] The calculation formula for the abnormal range is as follows:

[0037] Abnormal range = [min, max];

[0038] Min = Q1 - 1.5 × IQR;

[0039] Max = Q3 + 1.5 × IQR.

[0040] Specifically, the filling subunit is used to generate repair values at the positions of abnormal data according to several pieces of the outlier-removed data by using the interpolation method, so as to obtain several repaired data. The interpolation method includes any one of mean filling, median filling, nearest neighbor algorithm or regression-based prediction algorithm, and can be filled by calculating the mean value on the left and right of the abnormal data or the median value of the data. If the data lacks a timestamp, linear interpolation can be performed based on the timestamps of adjacent records.

[0041] Specifically, the normalization subunit is used to perform normalization processing on several pieces of the repaired data by using a normalization algorithm or a timestamp unification method to obtain several normalized data. Taking normalization processing and standardization processing as examples, the normalization formula is as follows:

[0042] x′ = (x - xmin) / (xmax - xmin);

[0043] In the formula, x′ represents the normalized data, x represents the input data, xmin represents the minimum value of the input data, and xmax represents the maximum value of the input data;

[0044] Specifically, the standardization processing is Z-score normalization, and its formula is as follows:

[0045] x′ = (x - μ) / σ;

[0046] Among them, μ is the mean value of the input data, and σ is the standard deviation of the input data. When applied to the Internet of Things scenario, if the temperature and humidity data recorded by the sensor comes from devices of different manufacturers, the system ensures consistent temperature ranges through standardization (for example, unified conversion between Celsius and Fahrenheit).

[0047] Specifically, the denoising subunit is used to remove noise from several pieces of the normalized data by using a time series smoothing algorithm to obtain several first data. For example, in the Internet of Things environment, the temperature sensor may be affected by external interference (such as electromagnetic signals or environmental fluctuations), resulting in random noise in the data. The moving average method is preferably used in the time series smoothing algorithm. The formula of the moving average method is as follows:

[0048]

[0049] In the formula, y tIt represents the smoothed data, and n represents the size of the sliding window.

[0050] Specifically, in the denoising unit, more advanced denoising techniques such as wavelet transform (WaveletTransform) or Kalman filter can also be used to process time series data or complex signals. For example, in real-time heart rate monitoring, the Kalman filter can dynamically adjust the denoising parameters to balance signal smoothness and real-time performance.

[0051] Specifically, the semantic alignment unit is used to calculate the semantic similarity of several of the first data and align the semantics between several of the first data according to the semantic similarity to obtain several second data. For structured and semi-structured data, the system ensures unified representation of data from different sources through field mapping

[0052] and semantic alignment. For example, user data obtained from two different databases are marked as "user_id" and "uid" respectively, but semantically they represent the same content. The system calculates the semantic similarity of field names through field mapping rules (such as defining a field alignment dictionary) or deep learning-based semantic embedding models (such as Word2Vec or BERT) to achieve alignment. The output unit is used to convert several of the second data into several multi-source heterogeneous data by using data modality conversion technology. For example, converting a categorical variable (such as a customer category field) into a numerical vector representation (such as one-hot encoding or embedding encoding).

[0053] Specifically, it further includes a data storage module. The data storage module is used to store several of the multi-source heterogeneous data. The data storage module includes several distributed databases, and the several distributed databases are used to store several of the multi-source heterogeneous data. For example, the system stores real-time Internet of Things data in a distributed time series database (such as InfluxDB), stores structured customer transaction data in a distributed SQL database (such as Amazon Aurora), and uses distributed object storage (such as AWS S3) for unstructured data (such as JSON files). At the same time, the system generates data metadata, including data source, data quality score, and processing history, to ensure data traceability. Finally, the system prepares the data into a standardized input format that can be directly called by subsequent analysis modules through a data access interface (such as RESTful API or GraphQL). In summary, the data acquisition module realizes the efficient access, cleaning, format standardization, and denoising of multi-source heterogeneous data, ensuring the high quality and usability of the data in the subsequent analysis stage. The above process is applicable to a wide range of application scenarios, including sensor data analysis in the Internet of Things environment, customer behavior data processing in the financial industry, and real-time log data management in Web services.

[0054] Specifically, such as Figure 3As shown in the figure, the association module includes a feature extraction unit, a normalization unit, a relationship calculation unit, and an association graph generation unit. The feature extraction unit is used to divide a number of the multi-source heterogeneous data into several categories according to the data type, and perform feature extraction on each category of multi-source heterogeneous data respectively to obtain a number of key features. Specifically, the feature extraction unit includes a classification sub-unit, a first feature extraction sub-unit, a second feature extraction sub-unit, and a third feature extraction sub-unit. The classification sub-unit is used to classify a number of the multi-source heterogeneous data according to the data type, and the data type includes structured data, semi-structured data, and unstructured data. The first feature extraction sub-unit is used to perform feature extraction on a number of the multi-source heterogeneous data whose classification result is structured data according to statistical features to obtain a number of first features. The second feature extraction sub-unit is used to perform feature extraction on a number of the multi-source heterogeneous data whose classification result is semi-structured data by using the Scripting Object Representation (SOR) to obtain a number of second features. The third feature extraction sub-unit is used to perform feature extraction on a number of the multi-source heterogeneous data whose classification result is unstructured data by using natural language processing technology to obtain a number of third features. The number of key features includes a number of first features, a number of second features, and a number of third features. The Scripting Object Representation (SOR) is a way to describe and organize data structures, used to process semi-structured or complex data. It represents data in the form of scripted objects, making this data easier to operate, access, and process. In this method, each data item is regarded as an object, containing attributes and methods, and can flexibly map and represent data from different sources and formats, usually used to represent content without a strict table structure (such as JSON, XML, log files, etc.). Through the Scripting Object Representation, data can be extracted, processed, and transformed to adapt to a variety of different data sources and requirements. In the fund investment scenario, the system extracts fund historical performance features from structured data, such as yield, volatility, maximum drawdown rate, etc.; parses field values from semi-structured data (such as market quotation data in JSON format), such as industry growth rate or market liquidity; extracts text embedding features from unstructured data (such as fund managers' investment reports, news texts), such as investment style keywords or potential sentiment indicators.To implement this process, the system uses specific technologies to process different types of data: for example, structured data extracts tabular information through SQL (Structured Query Language) or dedicated APIs; semi-structured data uses regular expressions, XPath, or API tools to parse field values; unstructured data uses natural language processing (NLP) technologies, such as the Transformer-based BERT (Bidirectional Encoder Representations from Transformers) model, to extract the embedding features of the text and combines sentiment analysis tools to refine semantic information. In addition, for unstructured media data such as pictures and videos, the system extracts image features or key frame information in the video through a convolutional neural network (CNN) to ensure comprehensive coverage of multi-modal features.

[0055] Specifically, the standardization unit is used to standardize a number of the key features to obtain a number of standardized key features. The standardization unit is used to ensure that data from different sources is fused under a unified measurement unit and format. For example, the historical return rate of a fund is stored in percentage form, while market data is represented by absolute values. The system uses a normalization method to map these data to the same numerical range (e.g., between 0 and 1) to eliminate the differences in data dimensions. At the same time, for timestamp information, the system adopts a time formatting rule (e.g., the ISO 8601 standard) to uniformly represent time data from different sources. In addition, for categorical features (such as investment style types or fund ratings), the system uses one-hot encoding or embedded encoding to convert them into numerical representations for fusion with numerical features. During the standardization process, data anomalies are also automatically detected and processed, such as removing outliers or repairing missing values through statistical analysis and machine learning models to ensure that the data quality meets the requirements of subsequent processing.

[0056] Specifically, the relationship calculation unit is used to calculate the similarity and correlation of a number of the standardized features. First, the correlation between different standardized features is calculated through the Pearson correlation coefficient or mutual information; second, the similarity of data features is evaluated through distance-based measurement methods (such as Euclidean distance or cosine similarity). The similarity is used to reflect the similarity between standardized features, and the greater the similarity, the more similar the two are. Taking the Pearson correlation coefficient as an example, the Pearson correlation coefficient is used to measure the linear correlation between two variables. Its value ranges from -1 to 1, where 1 represents a perfect positive correlation, -1 represents a perfect negative correlation, and 0 represents no correlation. Its calculation formula is as follows:

[0057]

[0058] In the formula, ρ x,y represents the Pearson correlation coefficient, ε x represents the standard deviation of the standardized feature x, and ε y represents the standard deviation of the standardized feature y.

[0059] Specifically, the association graph generation unit is used to generate an association graph based on the similarity and correlation of a number of the standardized features. Through the association graph, a number of the standardized features from different data sources are associated into a graph structure, where nodes represent standardized features and edges represent the relationship weights between standardized features. In this analysis process, the system can also automatically generate high-weight feature combinations. For example, through machine learning algorithms, the interaction relationship between fund managers' preferences and market liquidity is discovered to form an association graph.

[0060] Specifically, the specific implementation process of generating the association graph is described in detail as follows:

[0061] In the feature similarity calculation step, the present invention first calculates the similarity of the standardized features. For different types of standardized features, different similarity measurement methods are adopted. For continuous numerical features, the Euclidean distance or cosine similarity is used to calculate the similarity between features. The specific formula is: sim(x,y) = 1 - (||x - y|| / (||x|| + ||y||)). For categorical features, the Jaccard similarity coefficient or information entropy similarity calculation method is adopted. For example, the calculation formula of the Jaccard similarity coefficient is: J(A,B) = |A ∩ B| / |A ∪ B|. Through these methods, an n×n feature similarity matrix S can be obtained, where n is the total number of standardized features.

[0062] In the correlation evaluation step, based on the similarity calculation, the correlation between standardized features is further evaluated. The Pearson correlation coefficient or mutual information method is used to quantify the association strength between standardized features. For continuous features, the formula for calculating the Pearson correlation coefficient is: ρ(X,Y) = cov(X,Y) / (σx * σy). For discrete features, mutual information is used to calculate the association strength. Through this method, a correlation matrix C can be constructed to reflect the deep association relationships between features.

[0063] In the association graph construction step, based on the similarity matrix S and the correlation matrix C, a network construction algorithm in graph theory is used to generate an association graph. The specific steps include: (1) setting the thresholds threshold_s and threshold_c for similarity and correlation; (2) for feature pairs (x,y) that satisfy sim(x,y) > threshold_s and corr(x,y) > threshold_c, establishing connection edges in the association graph; (3) the weight w(x,y) of the edge is calculated by comprehensively considering similarity and correlation, and the formula is: w(x,y) = α * sim(x,y) + β * corr(x,y), where α and β are adjustment coefficients, and the sum of α and β is 1.

[0064] Furthermore, in the node attribute and semantic annotation step, to enhance the semantic expression of the association graph, rich attribute information is added to each node. Node attributes include: the original type of the feature, data source, numerical range, coefficient of variation, etc. For features with a concept hierarchy, ontology mapping is introduced to add semantic labels to the nodes. For example, for geography-related features, the geographical concept hierarchy to which they belong can be annotated, such as "city - province - country". This semantic annotation is helpful for subsequent knowledge reasoning and semantic association analysis.

[0065] Furthermore, in the graph structure optimization strategy step, the generated association graph is structurally optimized mainly through one or more of the following methods: (1) using the maximum spanning tree algorithm to retain the most representative feature connections; (2) applying the community discovery algorithm to identify closely associated subgraphs in the graph; (3) adopting graph sparsification techniques to control the connection density of the graph by setting the maximum number of edges or the minimum weight threshold. The optimization goal is to improve the interpretability and computational efficiency of the graph while retaining the key structural information.

[0066]

[0067] ​Furthermore, in the visualization and interactive display step, the force-directed layout algorithm is used to visualize the association graph. According to the weights and connection relationships between nodes, the positions of the nodes are automatically adjusted to make the graph present intuitive structural features. An interactive exploration interface is provided, allowing users to deeply analyze the feature associations in the graph by means of zooming, clicking, etc. For large-scale graphs, dimensionality reduction techniques (such as t-SNE) can be combined for visualization processing to provide a clearer display effect while retaining the key structural information.

[0068] Specifically, the association module can also obtain the association relationships between key features by adopting multi-modal embedding technology or causal inference models, so as to perform subsequent fusion. Through the multi-modal embedding model in deep learning, the data from different data sources are uniformly mapped into a high-dimensional vector space. For example, a text embedding model (such as BERT) is used to represent the features of text data, while a dedicated network (such as MLP (Multilayer Perceptron) or CNN) is used to process structured data and image data, and then feature alignment is performed in the high-dimensional space. This method is suitable for feature representation and correlation analysis of large-scale heterogeneous data. Causal inference technology reveals the causal relationships between data by constructing a causal relationship network. For example, in the field of fund investment, a causal graph can be constructed through historical data to analyze how market fluctuations affect fund returns, thus replacing some functions of the knowledge graph. Causal inference can provide a more interpretable basis for decision-making.

[0069] Specifically, as Figure 4 shown, in the fusion module, it includes: an aggregation unit, a weight calculation unit, and a fusion unit. The aggregation unit is used to obtain a number of updated features by using a graph convolutional network according to the association graph. The nodes in the association graph represent the normalized key features, and the edges represent the weights between the normalized key features. Node aggregation by using the graph neural network algorithm means integrating the features of a node with the features of its neighbor nodes to update the feature representation of the node. For each node, the feature representations of its neighbor nodes are extracted, and the features of the node are weighted and calculated or concatenated with the features of the neighbor nodes to complete the aggregation, obtaining aggregated features. Linear transformation is performed on the aggregated features, and then the linearly transformed aggregated features are activated by a non-linear function (such as ReLU) to map the linearly transformed features into a non-linear space to obtain updated features. All nodes in the association graph are traversed to obtain a number of updated features, and the data representation ability is enhanced through the number of updated features.

[0070] Specifically, the weight calculation unit is used to calculate a number of attention weights corresponding to the number of updated features by adopting an attention mechanism. The attention weight calculation formula is as follows:

[0071] α_i = softmax(score(h_i, h_query));

[0072] In the formula, α_i represents the attention weight of the i-th updated feature; h_i represents the vector of the i-th updated feature; h_query represents the query vector (global feature or reference vector); score() represents the similarity scoring function; softmax() represents converting the score into a probability distribution.

[0073] Specifically, the fusion unit is used to perform weighted summation on a plurality of the attention weights and a plurality of the updated features to obtain a fusion feature. The calculation formula of the fusion feature is as follows:

[0074] Fusion feature = Σ(α_i * h_i).

[0075] Specifically, in the fusion module, there are also included: a dimensionality reduction unit and a verification unit. The dimensionality reduction unit is used to perform dimensionality reduction on the fusion feature by using the principal component algorithm to obtain a dimensionality-reduced fusion feature. The dimensionality reduction formula of the principal component algorithm is as follows:

[0076] Y = P T ·(X - μ);

[0077] In the formula, P is the principal component matrix, X is the original feature matrix, μ is the mean vector, and P T is the transpose of the principal component matrix.

[0078] The verification unit is used to perform comprehensive verification on the dimensionality-reduced fusion feature. If the comprehensive verification is passed, the dimensionality-reduced fusion feature is output; if the verification is not passed, the parameters of the attention mechanism are adjusted, the attention weights are updated, and a new fusion feature is generated. The fusion feature importance analysis is used to understand the usefulness or value of each fusion feature. The goal is to determine the fusion feature that has the greatest impact on the model output, thereby improving the model performance, reducing overfitting, accelerating the training and inference speed, and enhancing the interpretability of the model.

[0079] Specifically, in the verification unit, the methods of the comprehensive verification include one or a combination of permutation importance, cross-validation, hold-out method, and largest eigenvalue.

[0080] Permutation Importance, which refers to monitoring the degree of decline in model performance by randomly permuting the values of each fusion feature. The greater the decline in performance, the more important the fusion feature.

[0081] Cross-validation is a statistical method used to evaluate the performance of machine learning models. By dividing the dataset into multiple subsets and repeating the training and validation processes, it assesses the generalization ability of the model. Cross-validation includes k-fold cross-validation, which randomly divides the dataset into k parts without replacement. k - 1 parts are used to train the model, and the remaining one part is used for model performance evaluation. This process is repeated k times to obtain k models and performance evaluation results, and the average value is taken as the final performance evaluation.

[0082] The holdout method (holdout cross validation) divides the original dataset into three parts: the training set, the validation set, and the test set. The training set is used to train the model, the validation set is used for model parameter selection and configuration, and the test set is used to evaluate the generalization ability of the model. The consistency index is used to evaluate the consistency level of the decision matrix, ensuring that the judgments in the decision-making process are reasonable and reliable. The consistency index includes the consistency ratio and the consistency index. Consistency Ratio (CR): CR is an important parameter for judging whether the judgment matrix satisfies consistency, and it is calculated based on the random consistency theory. Consistency Index (CI): CI reflects the degree of difference between the judgment matrix and the most consistent judgment matrix, and it reflects the inconsistency of the judge's weighting of elements at the same level. If CI is less than a predetermined threshold (e.g., 0.1), the judgment matrix is considered to have a high level of consistency.

[0083] Largest Eigenvalue: For an orthogonal judgment matrix, its largest eigenvalue should be close to the order of the judgment matrix. If the eigenvalue is close to the order, it indicates that the judgment matrix is closer to consistency. The fused features need to be further processed to ensure their semantic consistency so that features from different sources have the same semantic basis in the analysis. For example, although the return rate in fund performance and the market growth rate have different data sources, semantically they both represent a growth trend. The system achieves feature unification through semantic consistency mapping technology. Specifically, the system uses a deep learning-based semantic embedding model (such as BERT or Word2Vec) to calculate the semantic similarity between features and aggregates features with similar semantics through Relation Inference technology. For example, the system discovers that the historical return trend of a certain fund has a high similarity with the market growth rate of a certain industry, and then unifies them as the "growth potential" feature after semantic consistency processing. Through this process, the system ensures that the fused features are not only unified numerically but also form a consistent feature space semantically. Finally, the fused multi-modal features are stored in the distributed storage system of the system for subsequent analysis modules to call. For example, in the fund investment scenario, the system stores the fused feature vectors in a distributed file system (such as HDFS) and provides fast query and retrieval services for users through a distributed indexing technology (such as Elasticsearch). During the storage process, the system also assigns a unique identifier to each feature set and supports cross-modal retrieval functions, enabling users to quickly locate relevant data based on any feature (such as historical performance or investment style). At the same time, the system supports a hierarchical storage strategy, such as storing recent data in memory to accelerate real-time analysis and archiving historical data in low-cost storage to save resources. This efficient storage design ensures that the system has good scalability and can handle massive multi-source heterogeneous data.

[0084] After completing data fusion, the system stores the data in a distributed storage system (such as HDFS) and performs efficient real-time analysis tasks through a memory computing framework (such as Apache Spark). For example, in the securities market, the system performs real-time processing on large-scale stock trading data, including calculating the real-time trading volume, price fluctuation range, and short-term trend indicators of each stock. Through a dynamic resource scheduling algorithm, the system can adjust computing resources in real time according to the complexity of tasks. For example, when the market fluctuates violently, more computing nodes are allocated for real-time analysis to ensure the rapid output of analysis results.

[0085] Specifically, as Figure 5 shown, in the multiple analysis modules, it includes any two or more of: a statistical feature calculation unit, a real-time monitoring and analysis unit, a machine learning unit, a deep analysis unit, or a predictive analysis unit.

[0086] Specifically, the statistical feature calculation unit is used to calculate several statistical features of the fusion feature, and the several statistical features are expressed as statistical analysis results. The statistical features include mean, variance, median, or distribution, etc. This process is usually used to quickly understand the overall characteristics of the data and provide a preliminary basis for subsequent in-depth analysis. For example, in the field of fund investment, the system calculates the average return rate, standard deviation of returns, and volatility of each fund. In the implementation process, the system uses statistical formulas for processing. In addition, the system supports visual distribution analysis. For example, it displays the return distribution of different funds through a histogram to help users intuitively understand the data trend and outliers.

[0087] Specifically, the real-time monitoring and analysis unit is used to obtain the real-time data of the object to be analyzed, and adopt the sliding window technology and a preset anomaly detection model for the real-time data to obtain the real-time analysis result. While processing static data, the system supports the processing of real-time data streams. For example, in the field of bank risk control, the system detects abnormal transaction behaviors from real-time transaction data. The data stream processing is based on a distributed stream processing framework (such as Apache Kafka and Apache Flink). The system batches the real-time data stream through the sliding window technology, and the window size and sliding step can be configured according to the scenario. The anomaly detection model can be based on a clustering algorithm (such as DBSCAN) to mark in real time the transactions that deviate significantly from normal transaction behaviors. For example, when an account completes multiple large transfers in a short period of time, the system can immediately mark it as a high-risk transaction and issue an alarm.

[0088] Specifically, the machine learning unit is used to obtain real-time data of the object to be analyzed, input the real-time data into a pre-trained machine learning model to obtain a machine learning analysis result, and the pre-trained machine learning model is trained based on the fusion features. The system performs machine learning modeling on the fusion features, trains a prediction or classification model suitable for the scenario, and performs real-time or offline inference. For example, in the field of fund investment, the system uses historical data to train a return prediction model, and predicts the possible return rate of the fund in the next period through the model. In the implementation process, the system uses a regression algorithm (such as Gradient Boosting Decision Tree GBDT or Support Vector Regression SVR) to predict numerical targets, and can also use a classification model (such as Random Forest or XGBoost) to classify the risk level of the fund. The specific implementation includes the following steps: In the data splitting step, the system divides the data into a training set, a validation set, and a test set (for example, in a ratio of 7:2:1). In the feature selection step, the most useful features for the prediction target are screened through feature importance analysis (such as feature contributions based on SHAP values or LIME explanations). In the model training step, the selected algorithm and optimizer (such as Adam optimizer) are used to train the model to ensure model convergence. In the model evaluation step, the performance of the model is evaluated through metrics such as Mean Squared Error (MSE) or classification accuracy. The calculation formula of the mean squared error is as follows:

[0089]

[0090] In the formula, represents the predicted value, and y i represents the actual value, and N is the total number of samples.

[0091] The deep analysis unit is used to generate an explanatory analysis result for the fusion features using a deep learning algorithm. For complex data patterns and unstructured data, the system introduces deep learning technology for advanced analysis. For example, in the field of digital humans, the system analyzes the interaction records between users and virtual customer service through Natural Language Processing (NLP) technology to improve the user experience. The system uses the Transformer architecture (such as BERT or GPT) to extract features and perform semantic analysis on text data. For example, a text classification model is used to classify user feedback into categories such as "satisfied" and "needs improvement". For image data, the system uses a Convolutional Neural Network (CNN) to extract visual features. For example, in a recruitment scenario, the image resume or presentation video of a job applicant is analyzed to extract key content to assist in decision-making. The implementation steps of the deep learning model include data preprocessing (such as image scaling and text tokenization), model construction (such as multi-layer Transformer or ResNet network), and model inference and output interpretation.

[0092] Specifically, the prediction analysis unit is configured to obtain a prediction analysis result according to the fusion features by using a preset time series model. The prediction analysis unit includes: a time series generation sub-unit, a prediction sub-unit, and a prediction analysis sub-unit. The time series generation sub-unit is configured to generate time series data according to the fusion features. The prediction sub-unit is configured to input the time series data into a preset long short-term memory network model to obtain a prediction result. The prediction analysis sub-unit generates a prediction analysis result according to the prediction result and a preset risk threshold. Based on machine learning and deep learning analysis, the system further performs prediction analysis to provide users with clear decision-making suggestions. For example, in a risk control scenario, the system predicts the future transaction risk distribution of an account based on transaction history and recommends whether to increase the monitoring intensity. Prediction analysis is usually based on time series models (such as ARIMA, LSTM) or reinforcement learning algorithms (such as DQN (Deep Q-Network)), where the former is used for univariate trend prediction and the latter is used for complex policy generation.

[0093] For example, the LSTM model predicts the value of the future time point y_{T + 1} by inputting the time series data X = {x_1, x_2,..., x_T}. The hidden state update formula of the long short-term memory network model is as follows:

[0094] h t = σ(W x x t + W h h t-1 + b);

[0095] In the formula, σ represents an activation function (such as sigmoid or tanh), W x represents the input weight matrix, W h represents the hidden layer weight matrix, b represents the bias term, h t is the hidden state at the current time t, and h t-1 is the hidden state at the previous time. The prediction value calculation formula is as follows:

[0096] y T+1 = W y h T + b y ;

[0097] In the formula, y T+1 is the predicted output of the model at time T + 1, and this output is the future value predicted according to the current hidden state h T . Wy is the weight matrix of the output layer. It controls how the current hidden state h T is mapped to the predicted output y T+1 h Tis the hidden state at time T, and by represents the output layer bias term.

[0098] In addition, the predictive analysis unit has a built-in optimization engine that can generate execution strategies based on the analysis results. For example, in the investment field, the system can recommend the optimal asset allocation ratio to maximize returns and suggest that users reduce risky investments under high-risk market conditions.

[0099] Specifically, it also includes a visualization unit. After the analysis is completed, the system uses the result interpretation module and visualization tools to present the analysis results to the user in an intuitive way. For example, in the fund investment scenario, the system generates a line chart of return predictions and a heat map of risk distributions to help users quickly understand the model output. At the same time, the system combines interpretive AI (such as SHAP or LIME) to provide users with the contribution degree of each feature to the analysis results, helping users understand the logic behind the prediction. For example, the system can explain that the high-risk rating of a certain fund is mainly due to high historical volatility and low industry growth rate. In addition, the system supports users to customize the display method of the analysis results through an interactive dashboard, such as filtering data according to the time range or specific funds.

[0100] Specifically, each analysis unit in the multiple analysis module updates its parameters through an adaptive mechanism. For example, in the financial management scenario, when the cash flow of an enterprise suddenly fluctuates, the system automatically updates the data and adjusts the analysis model to provide more accurate predictions. The adaptive analysis module uses reinforcement learning algorithms to dynamically optimize the analysis process of the system, such as real-time adjusting the parameter configuration through the Policy Gradient method.

[0101] The formula for the optimization process is as follows:

[0102]

[0103] where θ is the model parameter, α is the learning rate, and R(τ) is the cumulative reward value of the policy.

[0104] Specifically, as Figure 6 shown, the comprehensive analysis module specifically includes: an index generation unit, a weighting unit, and a comprehensive analysis report unit. The index generation unit is used to generate a number of indexes based on a number of the analysis results and assign values to the number of the indexes according to the number of the analysis results, and the number of the indexes is the same as the number of the analysis results. The weighting unit is used to weight a number of the indexes by using a valuation algorithm. The comprehensive analysis report unit is used to calculate a comprehensive score by weighted calculation according to the weights of a number of the indexes and the values of a number of the indexes, and generate a comprehensive analysis report according to the comprehensive score and the values of each of the indexes.

[0105] Embodiment 2

[0106] As shown Figure 7 in the figure, an embodiment of the present invention further provides an intelligent fusion and analysis method for multi-source heterogeneous data based on cloud computing, including the following steps:

[0107] S10: Obtain a number of multi-source raw data, and preprocess the number of multi-source raw data to obtain a number of multi-source heterogeneous data;

[0108] S20: Extract features from the number of multi-source heterogeneous data to obtain a number of key features, and establish an association graph according to the number of key features;

[0109] S30: Calculate the fusion features according to the association graph by using the attention mechanism;

[0110] S40: Obtain a number of analysis results by using a number of analysis methods according to the fusion features;

[0111] S50: Input the number of analysis results into a preset analysis model to obtain a comprehensive analysis report.

[0112] Specifically, in step S10, it may specifically include:

[0113] S101: Obtain raw data from each data source, and identify the raw data by using dynamic protocol adaptation technology to obtain a number of multi-source raw data;

[0114] S102: Perform data cleaning, normalization, and data denoising on the number of multi-source raw data to obtain a number of first data;

[0115] S103: Calculate the semantic similarity of the number of first data according to the calculation, and align the semantics between the number of first data according to the semantic similarity to obtain a number of second data;

[0116] S104: Use data modality conversion technology to convert the number of second data into a number of multi-source heterogeneous data.

[0117] Specifically, in step S20, it specifically includes:

[0118] S201: Divide the number of multi-source heterogeneous data into several categories according to the data type, and extract features from each category of multi-source heterogeneous data respectively to obtain a number of key features;

[0119] S202: Perform standardization processing on the number of key features to obtain a number of standardized key features;

[0120] S203: Calculate the similarity and correlation of the number of standardized features according to the number of standardized features;

[0121] S204: Generate an association graph based on the similarities and correlations of a number of the standardized features.

[0122] Specifically, step S30 includes:

[0123] S301: Obtain a number of updated features by using a graph convolutional network according to the association graph;

[0124] S301: Calculate a number of attention weights corresponding to the number of the updated features by using an attention mechanism;

[0125] S301: Perform weighted summation calculation on the number of the attention weights and the number of the updated features to obtain a fused feature.

[0126] Specifically, step S40 includes:

[0127] S401: Calculate a number of statistical features of the fused feature, and the number of the statistical features is represented as a statistical analysis result;

[0128] S402: Obtain real-time data of the object to be analyzed, and use a sliding window technique and a preset anomaly detection model for the real-time data to obtain a real-time analysis result;

[0129] S403: Obtain real-time data of the object to be analyzed, input the real-time data into a pre-trained machine learning model to obtain a machine learning analysis result, and the pre-trained machine learning model is trained based on the fused feature;

[0130] S404: Generate an interpretation analysis result for the fused feature by using a deep learning algorithm;

[0131] S405: Obtain a prediction analysis result according to the fused feature by using a preset time series model.

[0132] Specifically, step S50 includes:

[0133] S501: Generate a number of metrics according to the number of the analysis results, and assign values to the number of the metrics according to the number of the analysis results, and the number of the metrics is the same as the number of the analysis results;

[0134] S502: Assign weights to the number of the metrics by using an assignment algorithm;

[0135] S503: Perform weighted calculation according to the weights of the number of the metrics and the values of the number of the metrics to obtain a comprehensive score, and generate a comprehensive analysis report according to the comprehensive score and the values of each of the metrics.

[0136] Specifically, throughout the data processing and analysis process, the system ensures data security and compliance through the security and privacy protection module. For example, in cross-organizational financial risk control cooperation, using federated learning technology, banks can collaborate to train anti-fraud models without directly sharing raw data. At the same time, through differential privacy technology, the system can desensitize sensitive information when publishing customer analysis results to ensure that personal privacy is not leaked. For example, in a financial risk control scenario, the system can share the results of the risk scoring model without exposing the specific transaction details of customers.

[0137] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.

[0138] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements any one of the above methods.

[0139] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. Of course, there are also other ways of readable storage media, such as quantum memory, graphene memory, and so on. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0140] The present invention also provides an electronic device. The electronic device according to an embodiment of the present invention includes: one or more processors; a storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement a multi-source heterogeneous data intelligent fusion and analysis method based on cloud computing provided by the present invention.

[0141] Reference is made below to Figure 8 , which shows a schematic structural diagram of a computer system 800 of an electronic device suitable for implementing an embodiment of the present invention. Figure 8 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0142] As Figure 8As shown, computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the computer system 800 are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0143] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read therefrom is installed into the storage section 808 as needed.

[0144] Specifically, according to an embodiment disclosed in the present invention, the process described in the main step diagram above can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the method shown in the main step diagram. In the above embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the central processing unit 801, the above functions defined in the system of the present invention are executed.

[0145] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0146] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0147] The units involved in the embodiments of the present invention can be implemented in software or in hardware. The described units can also be provided in a processor, and the names of these units do not, in some cases, constitute a limitation to the units themselves.

[0148] The above specific embodiments do not limit the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention should be included within the protection scope of the present invention.

Claims

1. A multi-source heterogeneous data intelligent fusion and analysis system based on cloud computing, characterized in that: include: A data acquisition module, used for acquiring a plurality of multi-source original data, and preprocessing the plurality of multi-source original data to obtain a plurality of multi-source heterogeneous data; An association module, used for extracting features from the multi-source heterogeneous data to obtain a number of key features, and establishing an association map based on the key features; A fusion module, used for calculating fusion features using an attention mechanism according to the association graph; A multiple analysis module, used to obtain a number of analysis results by using a number of analysis methods according to the fusion features; The comprehensive analysis module is used to input a number of the analysis results into a preset analysis model to obtain a comprehensive analysis report.

2. According to the cloud computing-based multi-source heterogeneous data intelligent fusion and analysis system of claim 1, it is characterized by: The data acquisition module specifically includes: An access identification unit is used to obtain original data from various data sources and identify the original data using a dynamic protocol adaptation method to obtain a plurality of multi-source original data; A basic processing unit, configured to perform data cleaning, normalization and data denoising on the plurality of multi-source original data to obtain a plurality of first data; a semantic alignment unit, configured to calculate semantic similarity of a plurality of the first data, and align semantics of the plurality of the first data according to the semantic similarity to obtain a plurality of the second data; The output unit is used to convert the plurality of the second data into a plurality of multi-source heterogeneous data by adopting a data modality conversion method.

3. According to the cloud computing-based multi-source heterogeneous data intelligent fusion and analysis system of claim 1, it is characterized by: The association module specifically includes: A feature extraction unit, used to classify the plurality of multi-source heterogeneous data into a plurality of categories according to data types, and extract features from each category of the multi-source heterogeneous data to obtain a plurality of key features; A standardization unit, used for performing standardization processing on the key features to obtain a number of standardized key features; A relationship calculation unit, used for calculating similarities and correlations of the plurality of standardized features according to the plurality of standardized features; The association map generating unit is used to generate an association map according to the similarity and correlation of a plurality of the standardized features.

4. The multi-source heterogeneous data intelligent fusion and analysis system based on cloud computing according to claim 1 is characterized in that: The fusion module specifically includes: An aggregation unit, configured to obtain a plurality of updated features using a graph convolutional network according to the association graph; A weight calculation unit, used to calculate a number of attention weights corresponding to the updated features using an attention mechanism; The fusion unit is used to perform weighted summation on the attention weights and the update features to obtain a fusion feature.

5. The multi-source heterogeneous data intelligent fusion and analysis system based on cloud computing according to claim 1 is characterized in that: The multiple analysis modules specifically include: any two or more of a statistical feature calculation unit, a real-time monitoring analysis unit, a machine learning unit, a deep analysis unit or a prediction analysis unit; A statistical feature calculation unit, used for calculating a plurality of statistical features of the fusion feature, wherein the plurality of statistical features are represented as statistical analysis results; A real-time monitoring and analysis unit, used to obtain real-time data of the object to be analyzed, and to obtain real-time analysis results by using a sliding window method and a preset anomaly detection model on the real-time data; A machine learning unit, used for acquiring real-time data of an object to be analyzed, and inputting the real-time data into a pre-trained machine learning model to obtain a machine learning analysis result, wherein the pre-trained machine learning model is trained based on the fusion feature; A deep analysis unit, used to generate an interpretation and analysis result using a deep learning algorithm for the fusion feature; The prediction and analysis unit is used to obtain prediction and analysis results based on the fusion features using a preset time series model.

6. The multi-source heterogeneous data intelligent fusion and analysis system based on cloud computing according to claim 1 is characterized in that: The comprehensive analysis module specifically includes: An indicator generating unit, used for generating a plurality of indicators according to the plurality of analysis results, and assigning values ​​to the plurality of indicators according to the plurality of analysis results, wherein the number of the indicators is the same as the number of the analysis results; A weighting unit, used for weighting the plurality of indicators using a value assignment algorithm; The comprehensive analysis report unit is used to obtain a comprehensive score by weighted calculation based on the weights of several of the indicators and the values ​​of several of the indicators, and to generate a comprehensive analysis report based on the comprehensive score and the values ​​of each of the indicators.

7. The multi-source heterogeneous data intelligent fusion and analysis system based on cloud computing according to claim 2 is characterized in that: The basic processing unit specifically includes: An outlier elimination subunit is used to calculate the interquartile range of the plurality of multi-source original data, calculate the abnormal range according to the interquartile range, filter out a plurality of abnormal data from the plurality of multi-source original data according to the abnormal range and the interquartile range, delete the plurality of abnormal data, and obtain a plurality of eliminated data; A filling subunit is used to generate a repair value at the position of the abnormal data by using an interpolation method according to the plurality of the eliminated data to obtain a plurality of repaired data; A normalization subunit, used for normalizing the repaired data by using a normalization algorithm or a timestamp unification method to obtain a number of normalized data; The denoising subunit is used to remove noise from the normalized data by using a time series smoothing algorithm to obtain a plurality of first data.

8. The multi-source heterogeneous data intelligent fusion and analysis system based on cloud computing according to claim 3 is characterized in that: The feature extraction unit specifically includes: A classification subunit, used for classifying the plurality of multi-source heterogeneous data according to data types, wherein the data types include structured data, semi-structured data and unstructured data; A first feature extraction subunit is used to extract features from the plurality of multi-source heterogeneous data whose classification results are structured data according to statistical features to obtain a plurality of first features; A second feature extraction subunit is used to extract features of the multi-source heterogeneous data whose classification results are semi-structured data by using a script object representation method to obtain a plurality of second features; A third feature extraction subunit is used to extract features of the plurality of multi-source heterogeneous data whose classification results are unstructured data by using a natural language processing method to obtain a plurality of third features; The several key features include several first features, several second features and several third features.

9. The multi-source heterogeneous data intelligent fusion and analysis system based on cloud computing according to claim 1 is characterized in that: The fusion module also includes: A dimension reduction unit, used for reducing the dimension of the fused features by using a principal component algorithm to obtain a reduced-dimensional fused feature; The verification unit is used to perform a comprehensive verification on the dimensionality reduction fusion feature. If the comprehensive verification is passed, the dimensionality reduction fusion feature is output; if the verification is not passed, the parameters of the attention mechanism are adjusted, the attention weight is updated, and a new fusion feature is generated.

10. The multi-source heterogeneous data intelligent fusion and analysis system based on cloud computing according to claim 5, characterized in that: The prediction analysis unit specifically includes: A time series generating subunit, used for generating time series data according to the fusion features; A prediction subunit, used for inputting the time series data into a preset long short-term memory network model to obtain a prediction result; The prediction analysis subunit generates a prediction analysis result based on the prediction result and a preset risk threshold.

Citation Information

Cited By

  • Efficient two-stage reverse osmosis petroleum coke water treatment system

    CN120757261A

  • NLP verification method for commercial password security assessment report based on agent

    CN121262111A

  • NLP checking method for agent-based commercial cryptographic security evaluation report

    CN121262111B

  • Data intelligent evaluation system based on multi-dimensional feature fusion

    CN122796454A