Fusion intelligent lake and warehouse system and method based on trusted data

By using classifiers and blockchain technology to classify and store data in the intelligent lake warehouse system, and building a data fusion model with deep learning algorithms, the balance between data governance, architecture flexibility and real-time nature of intelligent analysis in the existing technology is solved, and efficient data processing and intelligent analysis are achieved.

CN120105342AActive Publication Date: 2025-06-06JIANGSU ZHENYUN TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510262583.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-06
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The existing technology is difficult to balance data governance, architectural flexibility and intelligent analytics real-time, resulting in challenges for enterprises to handle massive heterogeneous data, AI convergence and cross-domain collaboration.

Method used

By using a classifier to classify mixed data in the intelligent lake warehouse system, and using blockchain technology to classify and store content, and combining deep learning algorithms to build a data fusion model to realize the fusion of data in different formats and abnormal risk judgments.

Benefits of technology

It realizes efficient classification and storage of massive heterogeneous data, improves data quality and security, enhances the real-time and accuracy of intelligent analysis, and allows more detailed implementation of security policies and abnormal risk warnings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105342A_ABST
    Figure CN120105342A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent lake and warehouse fusion system and method based on credible data, and relates to the technical field of data fusion. Three types of credible data in history are used for training to obtain a classifier, and format classification is carried out on mixed credible data collected in real time in the data; main information of each piece of credible data is extracted, and content classification is carried out on all the collected real-time credible data according to the main information of each piece of credible data; storing the credible data after content classification by using a block chain technology; constructing a data fusion model, and fusing all the credible data features in different formats to generate a comprehensive feature vector; respectively extracting features of three types of credible data, inputting the features into the data fusion model, and calculating an abnormal threshold value; and extracting the stored credible data of each content classification, obtaining a comprehensive feature vector through a data fusion model, judging the real-time comprehensive feature vector by using an abnormal threshold value, and sending out an early warning when judging that the credible data has an abnormal risk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data fusion, and in particular to a fusion intelligent lake warehouse system and method based on trusted data. Background Art

[0002] Traditional lake warehouses are computing-centric and have limited support for AI-driven real-time analysis, incremental computing and other scenarios, making it difficult to meet the data processing needs of the big model era. With the deepening of digital transformation and the tightening of data security and privacy protection, enterprises need to be compatible with historical systems (such as relational databases) and emerging architectures (such as lake warehouse integration), and achieve seamless management in hybrid environments through logical unification and virtualization technology. Cloud computing provides on-demand computing and storage resources, reducing the hardware construction costs of enterprises; distributed file systems and distributed database technologies are also becoming more and more perfect, which can support distributed storage and parallel processing of massive data, laying a solid underlying architecture foundation for building large-scale trusted data lake warehouse systems; federal architectures (such as the Digital Object Identifier Resolution Protocol) realize data sharing and trusted management among multiple subjects, breaking organizational boundaries. However, existing technologies have failed to effectively balance the strictness of data governance, the flexibility of architecture and the real-time nature of intelligent analysis, resulting in challenges for enterprises in dealing with massive heterogeneous data, AI integration and cross-domain collaboration. Summary of the invention

[0003] The purpose of the present invention is to provide a fusion intelligent lake warehouse system and method based on trusted data to solve the problems raised in the prior art.

[0004] To achieve the above object, the present invention provides the following technical solutions:

[0005] A fusion intelligent lake warehouse method based on trusted data, the method comprising the following steps:

[0006] S100, collect all the trusted data collected by the sensor, determine the format of the collected trusted data, and classify them into numerical trusted data, text trusted data, and video trusted data; use the three types of historical trusted data to train a classifier to classify the format of the mixed trusted data collected in real time;

[0007] Furthermore, the specific steps for format classification of the mixed trusted data collected in real time are as follows:

[0008] S101, collect the trusted data collected by the sensor, judge the format of the collected data, and divide it into numerical trusted data, text data, and video data; extract the format status of the three types of data as the format features of the three types of data respectively, use the format features of the three types of data as training sets, and calculate the information entropy of each type of data. The formula is:

[0009] ;

[0010] In the formula, H(X) represents the information entropy of each type of data calculated, p(x i ) represents the xth i The probability of type data, n represents the number of data in the training set for each type of data;

[0011] S102. After calculating the information entropy of each type of data, the information gain of the three data format features is calculated. The formula is:

[0012] ;

[0013] In the formula, Gain (X, S) represents the information gain of each of the three data format features calculated, S represents the format feature, and the information gain of the three data format features is calculated; the format feature with the largest information gain is selected as the classification standard;

[0014] S103, using the selected classification standard as a classification node, inputting the classification node into a decision tree to construct a classifier; after real-time collection of all data in the trusted space to obtain mixed data, inputting the mixed data into the classifier to automatically classify the collected mixed data.

[0015] Faced with massive and complex trusted data, different formats of data storage and access requirements vary. After using the classifier to classify the mixed data according to the format, the encryption product authorization key table and interface call key table can be quickly stored in the relational database. Relying on its mature index system, the target data can be accurately located within milliseconds; the text product usage log is placed in the document database, and when querying the log of a specific product usage period, it can be quickly screened based on the document identifier; the video has an exclusive storage area for easy management. In the classification process, it is easier to find data with confusing formats and non-compliant data, and clean or correct them in time to maintain data quality.

[0016] S200, after classifying the mixed trusted data collected in real time, analyzing the three types of trusted data after classification, extracting the main information of each trusted data respectively, and classifying the content of all the collected real-time trusted data according to the main information of each trusted data;

[0017] Furthermore, the specific steps of content classification of all collected real-time credible data according to the main information of each data are as follows:

[0018] S201. After classifying the mixed data collected in real time, extract all types of information values ​​in the numerical data, calculate the average value of all types of information values ​​as Zp, and then calculate the importance of each type of information value. The formula is:

[0019] ;

[0020] In the formula, Zd represents the importance of the calculated information value, Z represents each information value, and Zp represents the average value of all calculated information values. After calculating the importance of each information value, the information value with the greatest importance is selected as the main information of the corresponding numerical data;

[0021] S202: extract all the words contained in the text data, and calculate the importance of each word. The formula is:

[0022] ;

[0023] In the formula, Wd represents the importance of each word, m j represents the number of the jth word in the document, M represents the total number of words in the document; N represents the total number of documents, L j Indicates the number of documents containing the jth word in the document; repeatedly calculate the importance of all kinds of words in the text data, and select the word with the greatest importance as the main information of the text data;

[0024] S203, in the video data, the video data is processed by frame division, all key frames in the video data are extracted by using the CNN model, objects in all key frame images are extracted, and the total number of extracted object types is set to p, and the importance of each object in the video data is calculated, and the formula is:

[0025] ;

[0026] In the formula, Td represents the importance of each object in the calculated video data, k represents the number of occurrences of each object in all key frames, and kz represents the total number of occurrences of all objects in all key frames; r min and r max Represents the minimum and maximum grayscale values ​​of each object in all key frames, R min and R max Indicates the minimum and maximum grayscale values ​​of each key frame in all key frames; after calculating the importance of each object, the object with the largest importance value is selected as the main information of the video data;

[0027] S204. After the main information of numerical data, text data and video data is calculated respectively, the information content corresponding to the main information is used for mapping to obtain the information content of the main information of each data, and the information content of the main information is used for division among the three types of data, and the data with the same information content of the main information is divided into the same content data.

[0028] After segmenting by content, managers can quickly lock in key business segments, accurately grasp the operating details of specific business processes, and understand the data transaction process within the trusted space; through content classification, they can also discover hidden connections between different businesses; data content reflects differences in sensitivity, and content classification facilitates the implementation of more sophisticated security strategies.

[0029] S300, after the collected real-time trusted data is classified, the classified trusted data is stored using blockchain technology;

[0030] Furthermore, the specific steps of using blockchain technology to store data after content classification are as follows:

[0031] S301. After classifying the data in the trusted space, the blockchain technology is used to store the data with the same content in the same block. When storing the data, the hash value of each data is calculated using the hash algorithm. The formula is: HA=SHA256(D). In the formula, HA represents the calculated hash value of each data, and D represents the original data of each data. When storing, the hash value of each data is stored accordingly.

[0032] S400, after storing the trusted data after content classification, extract the features of all the trusted data of different formats in the same content, map the extracted features into the same space, perform dimension reduction on the features in the same space, build a data fusion model, and fuse the features of all the trusted data of different formats to generate a comprehensive feature vector;

[0033] Furthermore, the specific steps of fusing the features of all data in different formats to generate a comprehensive feature vector are as follows:

[0034] S401. After classifying and storing the data collected in real time, use the deep learning algorithm to analyze the stored data, extract the features of all data in the same content category, set the number of features extracted from each data as g, standardize each extracted feature to form a feature matrix, and calculate the covariance matrix of each feature matrix after standardization. The formula is:

[0035] ;

[0036] In the formula, F represents the covariance matrix of each feature matrix, e represents the number of each feature, V represents the feature matrix, and V T Represent the transposed matrix of the feature matrix; construct the characteristic equation, calculate the eigenvalue and eigenvector of the characteristic equation according to the covariance matrix, and select the feature with the largest eigenvalue as the main feature of the corresponding data;

[0037] S402: extract the main features of the three types of data in the same content category after step S401 as T 1 、T2 and T 3 ; Map the extracted features in the same space and reduce the dimension of the features in the same space; collect the three main features of the abnormal records in the historical credible space, calculate the weights of the three main features, and the data of the same content category contains three types of data. The formula is:

[0038] ;

[0039] In the formula, β represents the weight of each main feature calculated, Y i It indicates the number of records whose main characteristics change when an anomaly occurs, and Yz indicates the total number of abnormal records;

[0040] S403, constructing a data fusion model using the weights of the main features of the three types of data, the model is:

[0041] ;

[0042] In the formula, Zt represents the integrated feature vector after fusion, β 1 , β 2 and β 3 Represents the weights of the main features of numerical data, text data, and video data, Tz 1 , Tz 2 and Tz 3 Represents the main characteristics of numerical data, text data and video data.

[0043] S500, collecting abnormal records that occurred in the trusted space in history, extracting three types of trusted data from the abnormal records, extracting features of the three types of trusted data respectively and inputting them into the data fusion model, and calculating the abnormal threshold;

[0044] Furthermore, the specific steps for calculating the abnormal threshold are:

[0045] S501. Collect the abnormal records that occurred in the trusted space in history, extract three types of data from the abnormal records, use the features of the three types of data in the deep learning model to input into the data fusion model, obtain the comprehensive feature vector in each record, and calculate the abnormal threshold. The formula is:

[0046] ;

[0047] In the formula, In represents the calculated anomaly threshold, Zsp represents the average value of the comprehensive feature vector in history, and Zst represents the standard deviation of the comprehensive feature vector in history.

[0048] S600. Extract the trustworthy data for each content classification stored. After obtaining the comprehensive feature vector through the data fusion model, use the anomaly threshold to judge the real-time comprehensive feature vector. When it is judged that there is an anomaly risk in the trustworthy data, give an early warning.

[0049] Furthermore, the specific steps for giving an early warning when it is judged that there is an anomaly risk in the trustworthy data are as follows:

[0050] S601. Extract the data for each content classification stored. After obtaining the real-time comprehensive feature vector Zt through the data fusion model, use the anomaly threshold to judge the real-time comprehensive feature vector. When Zt≥In, it is judged that there is an anomaly risk in the trustworthy data and an early warning is given; when Zt<In, it is judged that there is no anomaly risk in the trustworthy data.

[0051] After fusing the three types of data using the data fusion model, judge whether there is an anomaly risk in the trustworthy data, examine the data transaction operation from multiple dimensions, no longer limited to the one-sided perspective of a single data type, and comprehensively capture anomaly signs. And isolated data types are difficult to show deep connections. After fusion, it may be possible to discover the association between the launch of data products and the generation of relevant data elements and the data loss problem caused by frequent data transactions, creating conditions for accurately positioning the root cause of anomalies; different types of data have their own advantages and disadvantages. After fusion, they confirm and complement each other, reduce misjudgments, and significantly improve the accuracy of anomaly judgment.

[0052] A fusion intelligent lakehouse system based on trustworthy data. The fusion intelligent lakehouse system includes a data collection module, a classification module, a storage module, a data fusion module, a threshold calculation module, and an anomaly judgment module;

[0053] The data collection module is used to collect all the data when an anomaly occurs in the trustworthy space in history;

[0054] The classification module is used to generate a classifier by judging the formats of all the data in the trustworthy space, classify the format of the mixed data; and extract the main information of each type of data for content classification;

[0055] The storage module is used to store the trustworthy data after content classification using blockchain technology;

[0056] The data fusion module is used to extract the features of all different format data in the same content, map the extracted features to the same space, reduce the dimension of the features in the same space, and construct a data fusion model;

[0057] The threshold calculation module is used to extract three types of data from the anomaly records, respectively extract the features of the three types of data and input them into the data fusion model, and calculate the anomaly threshold;

[0058] The abnormality judgment module is used to judge the real-time comprehensive feature vector using the abnormality threshold, and issue an early warning when it is judged that the trusted data has an abnormal risk.

[0059] The classification module includes a format classification unit and a content classification unit;

[0060] The format classification unit is used to judge the format of all data in the trusted space and to classify the format of mixed data;

[0061] The content classification unit is used to extract the main information of each data respectively, and classify the content of all the collected real-time data according to the main information of each data.

[0062] The data fusion module includes the main feature extraction unit and the model building unit;

[0063] The main feature extraction unit is used to extract the features of all data in the same content, calculate the feature value of each feature, and select the largest feature value as the main feature;

[0064] The model building unit is used to calculate the weight of each main feature and use the weight to build a data fusion model of the three types of data.

[0065] Compared with the prior art, the present invention has the following beneficial effects:

[0066] 1. The present invention utilizes a classifier to classify mixed data according to format. During the classification process, it is easier to find data with confusing formats and not meeting the specifications, and clean or correct them in a timely manner to maintain data quality.

[0067] 2. The present invention can also discover hidden connections between different businesses through content classification; data content reflects differences in sensitivity, and more sophisticated security strategies can be implemented based on content classification.

[0068] 3. The present invention uses a data fusion model to fuse the three types of data to determine whether the trusted data has abnormal risks, and examines the operation of products in the trusted space from multiple dimensions. It is no longer limited to the one-sided perspective of a single data type and can capture abnormal signs in all directions. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 This is a module distribution diagram of a fusion intelligent lake warehouse system based on trusted data of the present invention;

[0070] Figure 2 This is a schematic diagram of the steps of a fusion intelligent lake warehouse method based on trusted data in the present invention. DETAILED DESCRIPTION

[0071] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0072] Example: Figure 1-Figure 2 As shown, the present invention provides a technical solution.

[0073] A fusion intelligent lake warehouse method based on trusted data, the method comprising the following steps:

[0074] S100, collect the trusted data collected by all sensors in the trusted space, determine the format of the collected trusted data, and classify them into numerical trusted data, text trusted data, and video trusted data; use the three types of historical trusted data to train a classifier to classify the format of the mixed trusted data collected in real time;

[0075] The specific steps for format classification of mixed trusted data collected in real time are as follows:

[0076] S101. Collect data collected by all sensors in the trusted space, determine the format of the collected data, and classify the data into numerical data, text data, and video data; extract the format status of the three types of data as format features of the three types of data respectively, use the format features of the three types of data as training sets, and calculate the information entropy of each type of data. The formula is:

[0077] ;

[0078] In the formula, H(X) represents the information entropy of each type of data calculated, p(x i ) represents the xth i The probability of type data, n represents the number of data in the training set for each type of data;

[0079] S102. After calculating the information entropy of each type of data, the information gain of the three data format features is calculated. The formula is:

[0080] ;

[0081] In the formula, Gain (X, S) represents the information gain of each of the three data format features calculated, S represents the format feature, and the information gain of the three data format features is calculated; the format feature with the largest information gain is selected as the classification standard;

[0082] S103, using the selected classification standard as a classification node, inputting the classification node into a decision tree to construct a classifier; after real-time collection of all data in the trusted space to obtain mixed data, inputting the mixed data into the classifier to automatically classify the collected mixed data.

[0083] Faced with massive and complex trusted data, different formats of data storage and access requirements vary. After using the classifier to classify the mixed data according to the format, the encryption product authorization key table and interface call key table are quickly stored in the relational database. Relying on its mature index system, the target data can be accurately located within milliseconds; the text product usage log is placed in the document database, and when querying the log of a specific product usage period, it can be quickly filtered based on the document identifier; the video has an exclusive storage area for easy management. In the classification process, it is easier to find data with confusing formats and non-compliant data, and clean or correct them in time to maintain data quality.

[0084] S200, after classifying the mixed trusted data collected in real time, analyzing the three types of trusted data after classification, extracting the main information of each trusted data respectively, and classifying the content of all the collected real-time trusted data according to the main information of each trusted data;

[0085] The specific steps for content classification of all collected real-time trusted data according to the main information of each trusted data are:

[0086] S201. After classifying the mixed data collected in real time, extract all types of information values ​​in the numerical data, calculate the average value of all types of information values ​​as Zp, and then calculate the importance of each type of information value. The formula is:

[0087] ;

[0088] In the formula, Zd represents the importance of the calculated information value, Z represents each information value, and Zp represents the average value of all calculated information values. After calculating the importance of each information value, the information value with the greatest importance is selected as the main information of the corresponding numerical data;

[0089] S202: extract all the words contained in the text data, and calculate the importance of each word. The formula is:

[0090] ;

[0091] In the formula, Wd represents the importance of each word, m j represents the number of the jth word in the document, M represents the total number of words in the document; N represents the total number of documents, L jIndicates the number of documents containing the jth word in the document; repeatedly calculate the importance of all kinds of words in the text data, and select the word with the greatest importance as the main information of the text data;

[0092] S203, in the video data, the video data is processed by frame division, all key frames in the video data are extracted by using the CNN model, objects in all key frame images are extracted, and the total number of extracted object types is set to p, and the importance of each object in the video data is calculated, and the formula is:

[0093] ;

[0094] In the formula, Td represents the importance of each object in the calculated video data, k represents the number of occurrences of each object in all key frames, and kz represents the total number of occurrences of all objects in all key frames; r min and r max Represents the minimum and maximum grayscale values ​​of each object in all key frames, R min and R max Indicates the minimum and maximum grayscale values ​​of each key frame in all key frames; after calculating the importance of each object, the object with the largest importance value is selected as the main information of the video data;

[0095] S204. After the main information of numerical data, text data and video data is calculated respectively, the information content corresponding to the main information is used for mapping to obtain the information content of the main information of each data, and the information content of the main information is used for division among the three types of data, and the data with the same information content of the main information is divided into the same content data.

[0096] After segmenting by content, managers can quickly lock in key business segments, accurately grasp the operating details of specific business processes, and understand the data transaction process within the trusted space; through content classification, they can also discover hidden connections between different businesses; data content reflects differences in sensitivity, and content classification facilitates the implementation of more sophisticated security strategies.

[0097] S300, after the collected real-time trusted data is classified, the classified trusted data is stored using blockchain technology;

[0098] The specific steps of using blockchain technology to store trusted data after content classification are as follows:

[0099] S301. After classifying the data in the trusted space, the blockchain technology is used to store the data with the same content in the same block. When storing the data, the hash value of each data is calculated using the hash algorithm. The formula is: HA=SHA256(D). In the formula, HA represents the calculated hash value of each data, and D represents the original data of each data. When storing, the hash value of each data is stored accordingly.

[0100] S400, after storing the trusted data after content classification, extract the features of all the trusted data of different formats in the same content, map the extracted features into the same space, perform dimension reduction on the features in the same space, build a data fusion model, and fuse the features of all the trusted data of different formats to generate a comprehensive feature vector;

[0101] The specific steps of fusing the features of all credible data in different formats to generate a comprehensive feature vector are as follows:

[0102] S401. After classifying and storing the data collected in real time, use the deep learning algorithm to analyze the stored data, extract the features of all data in the same content category, set the number of features extracted from each data as g, standardize each extracted feature to form a feature matrix, and calculate the covariance matrix of each feature matrix after standardization. The formula is:

[0103] ;

[0104] In the formula, F represents the covariance matrix of each feature matrix, e represents the number of each feature, V represents the feature matrix, and V T Represent the transposed matrix of the feature matrix; construct the characteristic equation, calculate the eigenvalue and eigenvector of the characteristic equation according to the covariance matrix, and select the feature with the largest eigenvalue as the main feature of the corresponding data;

[0105] S402: extract the main features of the three types of data in the same content category after step S401 as T 1 、T 2 and T 3 ; Map the extracted features in the same space and reduce the dimension of the features in the same space; collect the three main features of the abnormal records of the historical credible data, calculate the weights of the three main features, and the data of the same content category contains three types of data. The formula is:

[0106] ;

[0107] In the formula, β represents the weight of each main feature calculated, Y i It indicates the number of records whose main characteristics change when an anomaly occurs, and Yz indicates the total number of abnormal records;

[0108] S403. Construct a data fusion model using the weights of the main features of three types of data. The model is as follows:

[0109] ;

[0110] In the formula, Zt represents the fused comprehensive feature vector, and β 1 , β 2 and β 3 represent the weights of the main features of numerical data, text data, and video data respectively. Tz 1 , Tz 2 and Tz 3 represent the main features of numerical data, text data, and video data respectively.

[0111] S500. Collect the abnormal records that occurred in the trusted space in history, extract the three types of trusted data in the abnormal records, respectively extract the features of the three types of trusted data and input them into the data fusion model, and calculate the abnormal threshold;

[0112] The specific steps for calculating the abnormal threshold are as follows:

[0113] S501. Collect the abnormal records that occurred in the trusted space in history, extract the three types of data in the abnormal records, input the features of the three types of data into the data fusion model using the deep learning model, obtain the comprehensive feature vector in each record, and calculate the abnormal threshold. The formula is as follows:

[0114] ;

[0115] In the formula, In represents the calculated abnormal threshold, Zsp represents the average value of the comprehensive feature vectors in history, and Zst represents the standard deviation of the comprehensive feature vectors in history.

[0116] S600. Extract the trusted data of each content classification stored. After obtaining the comprehensive feature vector through the data fusion model, use the abnormal threshold to judge the real-time comprehensive feature vector. When it is judged that there is an abnormal risk in the trusted data, an early warning is issued.

[0117] The specific steps for issuing an early warning when it is judged that there is an abnormal risk in the trusted data are as follows:

[0118] S601. Extract the data of each content classification stored. After obtaining the real-time comprehensive feature vector Zt through the data fusion model, use the abnormal threshold to judge the real-time comprehensive feature vector. When Zt ≥ In, it is judged that there is an abnormal risk in the trusted data and an early warning is issued; when Zt < In, it is judged that there is no abnormal risk in the trusted data.

[0119] After the three types of data are fused using the data fusion model, it is possible to determine whether the trusted data has abnormal risks, examine the data transaction operation from multiple dimensions, no longer limited to the one-sided perspective of a single data type, and capture abnormal signs in all directions. In addition, isolated data types are difficult to show deep connections. After fusion, it may be possible to find the connection between the launch of data products and the generation of related data elements and the problem of data missing caused by frequent data transactions, creating conditions for accurately locating the root cause of abnormalities; different types of data have their own advantages and disadvantages. After fusion, they can verify and supplement each other, reduce misjudgments, and significantly improve the accuracy of abnormal judgments.

[0120] A fusion intelligent lake warehouse system based on trusted data, the fusion intelligent lake warehouse system includes a data collection module, a classification module, a storage module, a data fusion module, a threshold calculation module and an abnormality judgment module;

[0121] The data collection module is used to collect all data when anomalies occur in the trusted space in history;

[0122] The classification module is used to generate a classifier by judging the format of all data in the trusted space, classify the format of the mixed data, and extract the main information of each type of data for content classification;

[0123] The storage module is used to store the trusted data after content classification using blockchain technology;

[0124] The data fusion module is used to extract features of all data of different formats in the same content, map the extracted features into the same space, reduce the dimension of the features in the same space, and build a data fusion model;

[0125] The threshold calculation module is used to extract three types of data from the abnormal records, respectively extract the features of the three types of data and input them into the data fusion model, and calculate the abnormal threshold;

[0126] The abnormality judgment module is used to judge the real-time comprehensive feature vector using the abnormality threshold, and issue an early warning when it is judged that the trusted data has an abnormal risk.

[0127] The classification module includes a format classification unit and a content classification unit;

[0128] The format classification unit is used to judge the format of all data in the trusted space and to classify the format of mixed data;

[0129] The content classification unit is used to extract the main information of each data respectively, and classify the content of all the collected real-time data according to the main information of each data.

[0130] The data fusion module includes the main feature extraction unit and the model building unit;

[0131] The main feature extraction unit is used to extract the features of all data in the same content, calculate the feature value of each feature, and select the largest feature value as the main feature;

[0132] The model building unit is used to calculate the weight of each main feature and use the weight to build a data fusion model of the three types of data.

[0133] Embodiment: A trusted space is now monitored for anomalies in real time, and numerical, textual, and video data in the trusted space are extracted respectively; all data are classified into product, transaction, and personnel categories according to the content of the data; each content category contains numerical, textual, and video data;

[0134] When collecting trusted data in real time, assume that a product class of data is obtained, including product record table, transaction log and operation record; extract the main features of the three types of data as transaction time, transaction record and operation video; map the three main features in the same space as transaction time 8:00, operation time 50min, number of data products 80; convert the three types of data into feature vectors;

[0135] The abnormal threshold is calculated as 80 based on the abnormal records in history;

[0136] The data fusion model Zt=β is obtained by calculating the weights through the abnormal records in history. 1 ×Tz 1 +β 2 ×Tz 1 +β 3 ×Tz 1 ; Calculate the real-time comprehensive feature vector 95; Determine whether the trusted data has abnormal risks.

[0137] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

Claims

1. A fusion intelligent lake warehouse method based on trusted data, characterized by: The method comprises the following steps: S100, collect the trusted data collected by all sensors, determine the format of the collected trusted data, and classify them into numerical trusted data, text trusted data, and video trusted data; use the three types of historical trusted data to train a classifier to classify the format of the mixed trusted data collected in real time; S200, after classifying the mixed trusted data collected in real time, analyzing the three types of trusted data after classification, extracting the main information of each trusted data respectively, and classifying the content of all the collected real-time trusted data according to the main information of each trusted data; S300, after the collected real-time trusted data is classified, the classified trusted data is stored using blockchain technology; S400, after storing the trusted data after content classification, extract the features of all data of different formats in the same content, map the extracted features in the same space, reduce the dimension of the features in the same space, build a data fusion model, and fuse the features of all trusted data of different formats to generate a comprehensive feature vector; S500, collecting abnormal records of trusted data in history, extracting three types of trusted data from the abnormal records, respectively extracting the three types of trusted data features and inputting them into the data fusion model, and calculating the abnormal threshold; S600: extract the stored trusted data of each content classification, obtain the comprehensive feature vector through the data fusion model, use the abnormal threshold to judge the real-time comprehensive feature vector, and issue an early warning when it is judged that there is an abnormal risk in the trusted data.

2. According to the method of claim 1, the fusion intelligent lake warehouse method based on trusted data is characterized by: The specific steps of format classification of the hybrid trusted data collected in real time in S100 are: S101, collect the trusted data collected by all sensors, determine the format of the collected trusted data, and classify them into numerical trusted data, text trusted data, and video trusted data; The format states of the three types of trusted data are extracted as three types of data format features respectively, and the format features of the three types of trusted data are used as training sets to calculate the information entropy of each type of trusted data. The formula is: ; In the formula, H(X) represents the information entropy of each type of trusted data calculated, and p(x i ) represents the xth i The probability of credible data of each type, n represents the number of data in the credible data training set of each type; S102. After calculating the information entropy of each type of trusted data, calculate the information gain of three types of trusted data format features, the formula is: ; In the formula, Gain (X, S) represents the information gain of each of the three types of credible data format features, S represents the format feature, and the information gain of the three types of credible data format features is calculated; the format feature with the largest information gain is selected as the classification standard; S103, using the selected classification standard as a classification node, and inputting the classification node into a decision tree to construct a classifier; After the mixed trusted data is obtained according to all the trusted data collected in real time, the mixed trusted data is input into the classifier to automatically classify the collected mixed trusted data.

3. The method for integrating intelligent lake storage based on trusted data according to claim 1 is characterized in that: The specific steps of content classification of all collected real-time trusted data according to the main information of each trusted data in S200 are: S201, after classifying the mixed trusted data collected in real time, extract all types of information values ​​in the numerical trusted data, calculate the average value of all types of information values ​​as Zp, and then calculate the importance of each information value, the formula is: ; In the formula, Zd represents the importance of the calculated information value, Z represents each information value, and Zp represents the average value of all calculated information values. After calculating the importance of each information value, the information value with the greatest importance is selected as the main information of the corresponding numerical credible data; S202: extract all types of words contained in the text-based credible data, and calculate the importance of each word. The formula is: ; In the formula, Wd represents the importance of each word, m j represents the number of the jth word in the document, M represents the total number of words in the document; N represents the total number of documents, L j represents the number of documents containing the jth word in the document; repeatedly calculate the importance of all kinds of words in the text-based reliable data, and select the word with the greatest importance as the main information of the text-based reliable data; S203, in the video type trusted data, the video type trusted data is processed by frame division, all key frames in the video type trusted data are extracted by using the CNN model, objects in all key frame images are extracted, and the total number of extracted object types is set to p, and the importance of each object in the video type trusted data is calculated, and the formula is: ; In the formula, Td represents the importance of each object in the calculated video-based credible data, k represents the number of occurrences of each object in all key frames, and kz represents the total number of occurrences of all objects in all key frames; r min and r max Represents the minimum and maximum grayscale values ​​of each object in all key frames, R min and R max Indicates the minimum and maximum grayscale values ​​of each key frame in all key frames; after calculating the importance of each object, the object with the largest importance value is selected as the main information of the video-type trusted data; S204. After the main information of numerical trusted data, text-based trusted data and video-based trusted data is calculated and obtained respectively, the information content corresponding to the main information is used for mapping to obtain the information content of the main information of each trusted data. The information content of the main information is used for division among the three types of trusted data, and the trusted data with the same information content of the main information are divided into trusted data of the same content.

4. The method for integrating intelligent lake storage based on trusted data according to claim 3 is characterized in that: The specific steps of using blockchain technology to store the credible data after content classification in S300 are: S301. After the trusted data is classified by content, the trusted data with the same content is stored in the same block using blockchain technology. When storing the trusted data, the hash value of each data is calculated using a hash algorithm. The formula is: HA=SHA256(D). In the formula, HA represents the calculated hash value of each trusted data, and D represents the original data of each trusted data. When storing, the hash value of each trusted data is stored accordingly.

5. The method for integrating intelligent lake storage based on trusted data according to claim 4 is characterized in that: The specific steps of fusing the features of all the trusted data in different formats to generate a comprehensive feature vector in S400 are: S401. After the trusted data collected in real time is classified and stored, the stored trusted data is analyzed using a deep learning algorithm to extract features of all trusted data in the same content category. The number of features extracted from each trusted data is set to g. Each extracted feature is standardized to form a feature matrix. The covariance matrix of each feature matrix after standardization is calculated. The formula is: ; In the formula, F represents the covariance matrix of each feature matrix, e represents the number of each feature, V represents the feature matrix, and V T represents the transposed matrix of the feature matrix; Construct the characteristic equation, calculate the eigenvalue and eigenvector of the characteristic equation according to the covariance matrix, and select the feature with the largest eigenvalue as the main feature corresponding to the reliable data; S402, extract the main features of three types of trusted data in the same content category through step S401 as T1, T2 and T3; map the extracted features in the same space, and reduce the dimension of the features in the same space; collect the three main features in the abnormal records of historical trusted data, and calculate the weights of the three main features. The data of the same content category contains three types of trusted data, and the formula is: ; In the formula, β represents the weight of each main feature calculated, Y i It indicates the number of records whose main characteristics change when an anomaly occurs, and Yz indicates the total number of abnormal records; S403. Use the weights of the main features of the three types of credible data to build a data fusion model. The model is: ; In the formula, Zt represents the fused comprehensive feature vector, β1, β2, and β3 represent the weights of the main features of numerical credible data, text credible data, and video credible data, and Tz1, Tz2, and Tz3 represent the main features of numerical credible data, text credible data, and video credible data.

6. The method for integrating intelligent lake storage based on trusted data according to claim 5 is characterized in that: The specific steps for calculating the anomaly threshold in S500 are as follows: S501. Collect the anomaly records of historical credible data, extract the three types of credible data in the anomaly records, input the features of the three types of credible data into the data fusion model using a deep learning model, obtain the comprehensive feature vector for each record, and calculate the anomaly threshold. The formula is: ; In the formula, In represents the calculated anomaly threshold, Zsp represents the average value of the comprehensive feature vectors in history, and Zst represents the standard deviation of the comprehensive feature vectors in history.

7. The method for integrating intelligent lake storage based on trusted data according to claim 6 is characterized in that: The specific steps for issuing a warning when it is determined that there is an anomaly risk in the credible data in S600 are as follows: S601. Extract the credible data of each content classification stored, obtain the real-time comprehensive feature vector Zt through the data fusion model, and use the anomaly threshold to judge the real-time comprehensive feature vector. When Zt≥In, it is determined that there is an anomaly risk in the credible data and a warning is issued; when Zt<In, it is determined that there is no anomaly risk in the credible data.

8. A fusion intelligent lake warehouse system based on trusted data, characterized by: The fusion intelligent lakehouse system includes a data collection module, a classification module, a storage module, a data fusion module, a threshold calculation module, and an anomaly judgment module; The data collection module is used to collect all credible data when an anomaly occurs historically; The classification module is used to generate a classifier by judging the formats of all credible data, classify the formats of the mixed credible data; and extract the main information of each type of credible data for content classification; The storage module is used to store the content-classified credible data using blockchain technology; The data fusion module is used to extract the features of all credible data in the same content with different formats, map the extracted features to the same space, reduce the dimensions of the features in the same space, and construct a data fusion model; The threshold calculation module is used to extract the three types of credible data in the anomaly records, respectively extract the features of the three types of credible data and input them into the data fusion model, and calculate the anomaly threshold; The anomaly judgment module is used to judge the real-time comprehensive feature vector using the anomaly threshold and issue a warning when it is determined that there is an anomaly risk in the credible data.

9. The fusion intelligent lake warehouse system based on trusted data according to claim 8 is characterized by: The classification module includes a format classification unit and a content classification unit; The format classification unit is used to judge the data formats of all credible data and classify the formats of the mixed credible data; The content classification unit is used to extract the main information of each credible data respectively, and classify the content of all collected real-time data according to the main information of each credible data.

10. The fusion intelligent lake warehouse system based on trusted data according to claim 8 is characterized by: The data fusion module includes a main feature extraction unit and a model construction unit; The main feature extraction unit is used to extract the features of all credible data in the same content, calculate the feature values of each feature, and select the one with the largest feature value as the main feature; The model building unit is used to calculate the weight of each main feature, and use the weight to build a data fusion model of three types of credible data.

Citation Information

Patent Citations

  • Method and system for realizing automatic data governance based on data standard

    CN114880318A

  • Visual image abnormal behavior detection method

    CN117152673A

  • Large model data management system and method based on deep learning

    CN117556065A

  • Scene data fusion analysis management system and method based on multi-source heterogeneity

    CN118656596A

  • Intelligent analysis system and method based on artificial intelligence

    CN118709045A