Artificial intelligence cloud data analysis method and system based on knowledge graph
Through the AI cloud data analysis method based on knowledge graph, the problem that traditional data analysis methods are difficult to deal with big data is solved, efficient data analysis and visualization is realized, and trend prediction support is provided in multiple fields.
Patent Information
- Application Number
- CN202510432385.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-08
AI Technical Summary
Traditional data analysis methods are difficult to cope with massive, high-speed and diversified data challenges, and cannot efficiently process and understand valuable information in big data.
Using an artificial intelligence cloud data analysis method based on knowledge graph, intelligent analysis and prediction of data is achieved through data collection, classification regularization, knowledge graph construction, multimodal feature extraction and trend prediction, combined with random forest model and visualization technology.
It realizes efficient processing and understanding of big data, improves the efficiency and accuracy of data analysis, and provides visual support in multiple fields.
Smart Images

Figure CN120277251A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and particularly to an artificial intelligence cloud data analysis method and system based on a knowledge graph. Background Art
[0002] With the advent of the big data era, the amount of data faced by enterprises and organizations has increased explosively. Traditional data analysis methods are difficult to cope with the challenges of massive, high-speed, and diverse data. At the same time, the development of cloud computing technology provides powerful computing resources and storage capabilities, while the progress of artificial intelligence technology makes it possible to extract valuable information from complex data. Against this background, artificial intelligence cloud data visualization analysis has emerged, which integrates artificial intelligence algorithms into the cloud platform, enabling users to process, analyze, and understand data more efficiently. Based on this, the present invention proposes an artificial intelligence cloud data analysis method and system based on a knowledge graph. Summary of the Invention
[0003] The present invention provides an artificial intelligence cloud data analysis method based on a knowledge graph, including:
[0004] S10. Determine cloud data sources for data collection, and perform preliminary regularization according to data classification to obtain a cloud data set.
[0005] S20. Construct a knowledge graph according to the cloud data set classification, and calculate the shortest path in the graph to simplify the graph.
[0006] S30. Extract multi-modal data features in the data set according to the cloud data set classification, and form a classification multi-modal feature set.
[0007] S40. Perform intelligent feature mining according to different classification multi-modal feature sets, and perform trend prediction tasks for each classification.
[0008] S50. Map the trend prediction situations of each classification to the corresponding knowledge graph, update the graph, and visualize the graph.
[0009] For the artificial intelligence cloud data analysis method based on a knowledge graph as described above, when selecting cloud data sources, one or more data sources should be selected, such as storage buckets of Amazon Web Services, object storage of Alibaba Cloud, data warehouses of Google Cloud Platform, etc. For different data sources, corresponding application programming interfaces are used for data access to ensure the comparability of the data collected from different data sources.
[0010] For the artificial intelligence cloud data analysis method based on a knowledge graph as described above, important entities in the knowledge graph can be identified using a named entity recognition tool, or entities can be directly determined according to data patterns and field definitions.
[0011] An artificial intelligence cloud data analysis method based on a knowledge graph as described above, in which for different types of data, corresponding feature extraction algorithms are adopted. The principal component analysis method is used to extract the numerical principal components as features; the word frequency and inverse document frequency of words in the document are calculated to extract text features; a neural network is used to extract image features; and the Mel spectrum coefficients are calculated to extract audio features.
[0012] An artificial intelligence cloud data analysis method based on a knowledge graph as described above, in which according to different classification multi-modal feature sets, a random forest model is used for feature mining. The random forest consists of multiple decision trees, and the importance of features is evaluated by analyzing the features in each decision tree, and the most predictive data features are identified and selected.
[0013] An artificial intelligence cloud data analysis method based on a knowledge graph as described above, in which the importance of features is obtained by summarizing their performance in all trees. If a feature is frequently used for effective splitting in multiple trees, that is, the Gini index is significantly reduced, then this feature is considered very important.
[0014] An artificial intelligence cloud data analysis method based on a knowledge graph as described above, in which the data features and prediction results are mapped to the corresponding nodes and edges in the knowledge graph, and additional attributes need to be added to the nodes and edges. And according to the data type and analysis purpose of the knowledge graph, mapping rules from data to visualization elements need to be formulated.
[0015] The present invention also provides an artificial intelligence cloud data analysis system based on a knowledge graph, including:
[0016] Collection module: used to determine the cloud data source for data collection and perform preliminary regularization according to data classification to obtain a cloud data set.
[0017] Atlas module: construct a knowledge graph according to the cloud data set classification and calculate the shortest path in the graph to simplify the graph.
[0018] Feature extraction module: extract multi-modal data features in the data set according to the cloud data set classification to form a classification multi-modal feature set.
[0019] Prediction module: perform intelligent feature mining according to different classification multi-modal feature sets and perform trend prediction tasks for each classification.
[0020] Visualization module: used to map the trend prediction situation of each classification to the corresponding knowledge graph, update the graph and visualize the graph.
[0021] The beneficial effects achieved by the present invention are as follows: The present invention visualizes data based on artificial intelligence cloud data, and can process, analyze and understand data more efficiently. Brief Description of the Drawings
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0023] Figure 1 It is a flowchart of an artificial intelligence cloud data analysis method based on a knowledge graph provided in Embodiment 1 of this application.
[0024] Figure 2 It is a schematic diagram of an artificial intelligence cloud data analysis system based on a knowledge graph provided in Embodiment 2 of this application. Detailed implementation manners
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0026] Embodiment 1
[0027] As Figure 1 shown, Embodiment 1 of this application provides an artificial intelligence cloud data analysis method based on a knowledge graph, including:
[0028] S10. Determine the cloud data source for data collection, and perform preliminary regularization according to data classification to obtain a cloud data set.
[0029] Sort out various cloud data sources, select one or more data sources, such as the storage buckets of Amazon Web Services, the object storage of Alibaba Cloud, the data warehouses of Google Cloud Platform, etc. For different data sources, use the corresponding application programming interfaces for data access to ensure the comparability of the data collected from different data sources. Collect structured data by directly reading data from database tables; for semi-structured data, convert it into a processable format through a parser, and for unstructured data, use techniques such as text extraction and image recognition for preliminary processing to remove obsolete, skewed distribution, repetitive, and poor-quality data.
[0030] After obtaining the data, perform preliminary regularization according to classification, including cleaning, denoising, handling missing values, and standardization processing, unify the data format, and convert all data into a general format that is convenient for subsequent processing. Classification means that data belonging to the same field is classified as one category, and different classifications include medical, financial, etc. For numerical data, perform standardization processing to make it have the same magnitude and distribution range.
[0031] S20. Construct a knowledge graph based on the classification of the cloud data set, and calculate the shortest paths in the graph to simplify the graph.
[0032] For the classification data in the cloud data set, extract the entities and relationships in each category of data respectively to construct their own knowledge graphs. For text data, use a named entity recognition tool to identify important entities. For structured data, directly determine the entities according to the data schema and field definitions. Identify the relationships between entities in the text and extract relevant attributes for each entity through custom syntax rules and semantic patterns. Integrate the extracted entities, relationships, and attributes to construct a knowledge graph. Use a graph database for storage. The graph database represents entities as nodes, relationships as edges, and attributes as the attribute values of nodes or edges. Use the existing knowledge graph for reasoning to discover new knowledge or fill in missing information.
[0033] Use the formula to quickly locate the most widely connected nodes in the knowledge graph, so as to determine the central node of the knowledge graph and the nodes close to the central point, where σ st is the number of shortest paths from node s to node t, and σ st (v) is the number of shortest paths from node s to node t passing through node v.
[0034] In the huge relationship network of the knowledge graph, by calculating the shortest paths, indirect connections between entities can be discovered, which helps to simplify the graph. Let the knowledge graph be graph G=(V, E), where V is the set of nodes, E is the set of edges, and each edge (u, v)∈E has a weight ω(u, v). u and v are the two nodes of an edge. For a given source node s and target node t, create a distance array d, where d[v] represents the shortest distance from the source node s to node v. Initially, d[s]=0, and for other nodes v≠s, d[v]=∞. Create a set S to store the nodes for which the shortest paths have been determined. Initially, S is an empty set. Among the nodes in the set V - S for which the shortest paths have not been determined, select the node u with the smallest d value and add it to the set S. For all adjacent nodes v of node u, if d[u]+ω(u, v)<d[v], then update d[v]=d[u]+ω(u, v). Repeat the above steps until the target node t is added to the set S. At this time, d[t] is the length of the shortest path from the source node s to the target node t. Connect the nodes according to the shortest path length to simplify the graph.
[0035] Adopt a graph embedding algorithm to map the nodes in the knowledge graph to a low-dimensional vector space so that the semantic relationships between nodes can be reflected in the vector space.
[0036] S30. Classify and extract the multi-modal data features in the dataset according to the cloud dataset to form a classified multi-modal feature set.
[0037] For different types of data, corresponding feature extraction algorithms are adopted.
[0038] S31. Extract numerical features.
[0039] For numerical data, the principal component analysis method is used to extract the main components as features. Calculate the covariance matrix based on the standardized numerical data. For a dataset X with n samples and p variables for each sample, the covariance matrix calculation formula is X T represents the transpose of the dataset. Perform eigenvalue decomposition on the covariance matrix A to find its eigenvalues λ1, λ2,..., λ p and the corresponding eigenvectors v1, v2,..., v p . The eigenvalues represent the variance sizes explained by each principal component, and the eigenvectors define the directions of the new coordinate system, that is, the directions of the principal components. Sort the eigenvectors according to the magnitudes of the eigenvalues. The larger the eigenvalue, the greater the variation of the data in the direction represented by the corresponding eigenvector. Therefore, select the first k largest eigenvalues and their corresponding eigenvectors as the principal components, and project the original data into a new subspace using the selected k principal components. The specific formula is B = X × W k , where B is the transformed data matrix and W k is the matrix composed of k principal components.
[0040] Determine how many principal components need to be retained according to the cumulative contribution rate. The cumulative contribution rate is the proportion of the total variance explained by the first k principal components. The specific calculation formula is The cumulative contribution rate can effectively extract the most important features from the numerical dataset while reducing the data dimension. The numerical feature set is N = {n1, n2,..., n a}
[0041] S32. Extract text features.
[0042] Perform necessary preprocessing operations on the original text, including removing punctuation marks, converting to lowercase, removing stop words, etc., to reduce noise and improve the extraction effect. Based on the preprocessed text data, construct a vocabulary, which involves collecting all the words that appear in all documents and assigning a unique index number to each word. For each word in each document, calculate its term frequency TF(t, d), that is, the number of times a certain word appears in the document divided by the total number of words in the document. The inverse document frequency IDF(t) is used to measure the general importance of a word. According to the TF-IDF values of each word in each document, generate a feature vector for each document. The specific formula is nt,d denotes the number of times the term t appears in the document d, denotes the sum of the occurrences of all terms in the document d, t' is a variable used to iterate over each possible term in the document d, N represents the total number of documents in the document set, |{d ∈ D: t ∈ d}| represents the number of documents containing the term t, and D represents the set containing all documents. The length of the vector is equal to the size of the vocabulary, and the value at each position corresponds to the TF-IDF weight of the corresponding vocabulary. If a term does not appear in the document, the value at that position is 0. The set of text features is T = {t1, t2,..., t b}.
[0043] S33. Extract image features.
[0044] The original image I is input into the input layer. The original image is represented as a three-dimensional tensor I ∈ R H×W×C , where R represents the set of real numbers, H represents the height of the image, W represents the width of the image, and C represents the number of channels of the image. The convolutional layer detects local features in the image by applying multiple filters. For each filter the convolution operation is defined as For each element K(m, n, c) in the convolutional kernel, multiply it by the value I(i + m, j + n, c) at the corresponding position in the input image, and sum all these product results to obtain the convolution result at that position, the feature map. i, j are the position coordinates on the output feature map, k h , k w are the height and width of the filter respectively.
[0045] Apply the activation function O(i, j) = max(0, (I * K)(i, j) + b) to introduce non-linearity. O(i, j) represents the output value after being processed by the activation function, max(0, x) represents the activation function, and b represents the bias term.
[0046] Use the pooling layer to reduce the spatial size of the feature map, reduce the computational complexity, and at the same time improve the invariance of the features. The max-pooling formula is p(i, j) = max{O(i × s h : i × s h + k h , j × s w : j × s w + k w )}, where s h , s w represent the strides in the height and width directions respectively, which determine the distance the pooling window moves. k h , k w represent the height and width of the pooling window respectively, O represents the feature map input to the pooling layer, and max{O(i × s h : i × sh +k h ,j×s w :j×s w +k w )} represents taking the region determined by i and j from the feature map O, and taking the maximum value within this region as the output after pooling.
[0047] The fully connected layer converts the feature map after multiple convolutions and poolings into a prediction result. Assuming that the output of the previous layer is flattened into a vector a, the output of the fully connected layer can be expressed as y = Wa + b, where W is the weight matrix and b is the bias vector. The output layer outputs image features according to the task requirements, and the parameter adjustment depends on the backpropagation algorithm and the gradient descent method during the feature extraction process. The image feature set is I = {i1, i2,..., i c}.
[0048] S34. Extract audio features.
[0049] Perform pre-emphasis processing on the audio signal to enhance the high-frequency components, divide the continuous audio signal into several short-time frames, and window each frame signal to reduce spectral leakage. Perform Fourier transform on each frame signal to obtain the frequency-domain representation where X(k,t) is the k-th frequency component of the t-th frame, x(n,t) is the n-th sampling point of the t-th frame, j represents the imaginary unit, and N is the total number of samples in the frame. Calculate the squared magnitude of each Fourier transform result to obtain the power spectrum, which is used to quantify the energy of each frequency component. The specific formula is P(k) = |X(k,t)| 2 . Apply a set of triangular filters to the power spectrum to convert the linearly spaced frequencies to frequencies on the non-linear Mel scale perceived by the human auditory system. Each filter covers a specific range of frequencies, and its output is the sum of the energies of all frequency components within the corresponding range.
[0050] The specific formula is where S(m) is the logarithm of the output of the m-th Mel filter, H m (k) is the transfer function of the m-th Mel filter, and P(k) represents the power spectrum. The logarithmically compressed Mel spectrum coefficients are further processed by discrete cosine transform to produce the final audio features. The specific formula is c l is the l-th Mel spectrum coefficient, that is, the audio feature, M represents the number of Mel filters, L is the number of coefficients, and π represents the circumference ratio. The audio feature set is Q = {q1, q2,..., q d}.
[0051] S40. Perform intelligent feature mining based on different classification multi-modal feature sets and perform trend prediction tasks for each classification.
[0052] Feature mining is performed using a random forest, which consists of multiple decision trees. By analyzing the features in each decision tree, the importance of the features is evaluated to identify and select the most predictive data features.
[0053] Training dataset where α i is a feature vector containing various features such as numerical values, text, images, audio, etc., and β i is the corresponding label, represents the number of features. The random forest contains T decision trees. For each tree, samples are drawn with replacement from the dataset φ to construct a subset, and then a decision tree is trained on this subset. During the construction of the decision tree, each feature helps make a splitting decision by reducing the impurity. Therefore, the importance of a feature can be measured by its contribution to the reduction of the overall impurity. The Gini index is used to measure the impurity of a node, and the Gini index formula is where f i is the proportion of samples of the i-th class in the node, and C is the number of classes. For each decision tree T t , when splitting a node, the contribution of each feature to the impurity is calculated. Suppose at a certain node, splitting is based on feature j, and the contribution I j,t of the feature to the impurity at this node is where Gini before represents the Gini index of the node before splitting, Gini left and Gini right are the Gini indices of the two child nodes after splitting respectively, n left and n right are the number of samples of the two child nodes respectively, and the total number of samples is n = n left + n right .
[0054] The importance of a feature is obtained by summarizing its performance in all trees. If feature j is frequently used for effective splitting in multiple trees, that is, it significantly reduces the Gini index, then this feature is considered very important. The total amount of impurity reduction of feature j during the splitting process is averaged to obtain the importance score According to the importance score, features that meet the score threshold setting are selected from numerical, text, image, and audio features. If there are no features in a certain category of numerical, text, image, or audio that meet the threshold setting, the threshold can be relaxed.
[0055] After determining the most predictive data features, a formula is used for trend prediction. The specific formula is where the numerical feature set is N = {n1, n2,..., n a} and its weight set is W N = {wn1 , w n2 ,..., w na} and the text feature set is T = {t1, t2,..., t b}, and the weight set is W T = {w t1 , w t2 ,..., w tb}, and the image feature set is I = {i1, i2,..., i c}, and the weight set is W I = {w i1 , w i2 ,..., w ic}, and the audio feature set is Q = {q1, q2,..., q d}, and the weight set is W A = {w a1 , w a2 ,..., w ad}. Considering the time factor, the time period T cycle , and the external environmental factor set E = {e1, e2,..., e f}, and the external environmental factors include economic factors, policy factors, sociocultural factors, and technological factors, and its weight set is W E = {w e1 , w e2 ,..., w ef}. The entire formula comprehensively considers the influence of numerical features, text features, image features, audio features, and external environmental factors on trend prediction by multiplying two parts.
[0056] S50. Map the trend prediction situations of each classification to the corresponding knowledge graph, update the graph, and visualize the graph.
[0057] Map the data features and prediction results to the corresponding nodes and edges in the knowledge graph, which involves adding additional attributes to the nodes and edges. For each entity v j ∈ V and its corresponding prediction result y j ∈ Y, and for each edge e k ∈ E and its corresponding prediction result y k ∈ Y, define a function f v to update the attributes of the entity and the edge. The specific formula is L = f v (v j , y j | e k , y k ), where L represents the updated knowledge graph.
[0058] According to the data types and analysis purposes of the knowledge graph, mapping rules from data to visualization elements are formulated. For numerical data, map the numerical magnitude to the vertical coordinate of the line chart; for categorical data, map different categories to different colors and shapes, and adjust the node size and color based on feature importance to reflect their importance. Mapping rules are formulated for different types of data. After formulating the mapping rules, select a suitable visualization tool to visualize the graph. Each classification graph is visualized according to the mapping rules, providing visualization information support for multiple fields.
[0059] Design user interaction methods to improve the efficiency of visualization, such as hovering to display detailed information, clicking to zoom in on a specific area, and dragging to re-layout, etc. Optimize chart parameters to enhance the efficiency of information transmission, and select an appropriate layout algorithm to clearly display the graph structure. And adjust parameters such as the compactness, balance, color saturation, and graphic ratio of the visualization layout according to user feedback.
[0060] Embodiment 2
[0061] As Figure 2 shown, Embodiment 2 of the present application provides an artificial intelligence cloud data analysis system based on a knowledge graph, including:
[0062] Collection module: used to determine cloud data sources for data collection, and perform preliminary regularization according to data classification to obtain a cloud data set.
[0063] Sort out various cloud data sources, select one or more data sources, such as the S3 bucket of Amazon Web Services, the object storage of Alibaba Cloud, the data warehouse of Google Cloud Platform, etc. For different data sources, use the corresponding application programming interfaces for data access to ensure the comparability of the data collected from different data sources. Collect structured data by directly reading data from database tables; for semi-structured data, convert it into a processable format through a parser, and for unstructured data, use techniques such as text extraction and image recognition for preliminary processing to remove outdated, skewed distribution, repetitive, and poor-quality data.
[0064] After obtaining the data, perform preliminary regularization according to classification, including cleaning, denoising, handling missing values, and standardization processing, unify the data format, and convert all data into a common format that is convenient for subsequent processing. Classification means that data belonging to the same field is grouped into one category, and different classifications include medical, financial, etc. For numerical data, perform standardization processing to make it have the same magnitude and distribution range.
[0065] Graph module: includes a construction sub-module and a path sub-module.
[0066] Construction sub-module: construct a knowledge graph according to the cloud data set classification.
[0067] For the categorical data in the cloud dataset, extract the entities and relationships in each category of data respectively to construct their respective knowledge graphs. For text data, use a named entity recognition tool to identify important entities. For structured data, directly determine the entities according to the data schema and field definitions. Identify the relationships between entities in the text and extract relevant attributes for each entity through custom syntax rules and semantic patterns. Integrate the extracted entities, relationships, and attributes to construct a knowledge graph. Use a graph database for storage. The graph database represents entities as nodes, relationships as edges, and attributes as the attribute values of nodes or edges. Use the existing knowledge graph for reasoning to discover new knowledge or fill in missing information.
[0068] Path sub-module: Used to calculate the shortest paths in the graph to simplify the graph.
[0069] Use the formula as Quickly locate the most widely connected nodes in the knowledge graph to determine the central node of the knowledge graph and the nodes close to the central point, where σ st is the number of shortest paths from node s to node t, and σ st (v) is the number of shortest paths from node s to node t passing through node v.
[0070] In the huge relationship network of the knowledge graph, by calculating the shortest paths, the indirect connections between entities can be discovered, which helps to simplify the graph. Let the knowledge graph be a graph G=(V, E), where V is the set of nodes, E is the set of edges, and each edge (u, v)∈E has a weight ω(u, v). u and v are the two nodes of an edge. For the given source node s and target node t, create a distance array d, where d[v] represents the shortest distance from the source node s to node v. Initially, d[s]=0, and for other nodes v≠s, d[v]=∞. Create a set S to store the nodes for which the shortest paths have been determined. Initially, S is an empty set. Among the nodes in the set V - S for which the shortest paths have not been determined, select the node u with the smallest d value and add it to the set S. For all adjacent nodes v of node u, if d[u]+ω(u, v)<d[v], then update d[v]=d[u]+ω(u, v). Repeat the above steps until the target node t is added to the set S. At this time, d[t] is the length of the shortest path from the source node s to the target node t. Connecting the nodes according to the shortest path length can simplify the graph.
[0071] Adopt a graph embedding algorithm to map the nodes in the knowledge graph to a low-dimensional vector space so that the semantic relationships between nodes can be reflected in the vector space.
[0072] Feature extraction module: Includes numerical feature sub-module, text feature sub-module, image feature sub-module, audio feature sub-module.
[0073] Numerical feature sub-module: used to extract numerical features.
[0074] For numerical data, the principal component analysis method is used to extract the main components as features. Calculate the covariance matrix based on the standardized numerical data. For a dataset X with n samples and p variables for each sample, the covariance matrix calculation formula is X T represents the transpose of the dataset. Perform eigenvalue decomposition on the covariance matrix A to find its eigenvalues λ1, λ2,..., λ p and the corresponding eigenvectors v1, v2,..., v p . The eigenvalues represent the variance sizes explained by each principal component, and the eigenvectors define the directions of the new coordinate system, that is, the directions of the principal components. Sort the corresponding eigenvectors according to the sizes of the eigenvalues. The larger the eigenvalue, the greater the change in the data in the direction represented by the corresponding eigenvector. Therefore, select the first k largest eigenvalues and their corresponding eigenvectors as the principal components, and project the original data into the new subspace using the selected k principal components. The specific formula is B = X × W k , where B is the transformed data matrix, and W k is the matrix composed of k principal components.
[0075] Determine how many principal components need to be retained according to the cumulative contribution rate. The cumulative contribution rate is the proportion of the total variance explained by the first k principal components. The specific calculation formula is The cumulative contribution rate can effectively extract the most important features from the numerical dataset and reduce the data dimension at the same time. The numerical feature set is N = {n1, n2,..., n a}
[0076] Text feature sub-module: used to extract text features.
[0077] Perform necessary preprocessing operations on the original text, including removing punctuation marks, converting to lowercase, removing stop words, etc., to reduce noise and improve the extraction effect. Based on the preprocessed text data, construct a vocabulary, which involves collecting all the words appearing in all documents and assigning a unique index number to each word. For each word in each document, calculate its term frequency TF(t, d), that is, the number of times a certain word appears in the document divided by the total number of words in the document. The inverse document frequency IDF(t) is used to measure the general importance of a word. Generate a feature vector for each document according to the TF-IDF value of each word in each document. The specific formula is n t,d represents the number of times the word t appears in the document d, Denote the sum of the occurrences of all words in document d. Let t' be a variable used to iterate through each possible word in document d. Let N be the total number of documents in the document set, |{d ∈ D: t ∈ d}| denote the number of documents containing word t, and D denote the set containing all documents. The length of the vector is equal to the size of the vocabulary, and the value at each position corresponds to the TF-IDF weight of the corresponding vocabulary. If a certain word does not appear in the document, the value at that position is 0. The set of text features is T = {t1, t2,..., t b}.
[0078] Image feature sub-module: used to extract image features.
[0079] The original image I is input into the input layer. The original image is represented as a three-dimensional tensor I ∈ R H×W×C , where R represents the set of real numbers, H represents the height of the image, W represents the width of the image, and C represents the number of channels of the image. The convolutional layer detects local features in the image by applying multiple filters. For each filter The convolution operation is defined as For each element K(m, n, c) in the convolutional kernel, multiply it by the value I(i + m, j + n, c) at the corresponding position in the input image, and sum all these product results to obtain the convolution result at that position, the feature map. i, j are the position coordinates on the output feature map, k h , k w are the height and width of the filter respectively.
[0080] Apply the activation function O(i, j) = max(0, (I * K)(i, j) + b) to introduce non-linearity. O(i, j) represents the output value after being processed by the activation function, max(0, x) represents the activation function, and b represents the bias term.
[0081] Use the pooling layer to reduce the spatial size of the feature map, reduce the computational complexity, and at the same time improve the invariance of the features. The formula for max pooling is p(i, j) = max{O(i × s h : i × s h + k h , j × s w : j × s w + k w )}, where s h , s w represent the strides in the height and width directions respectively, which determine the distance the pooling window moves. k h , k w represent the height and width of the pooling window respectively, O represents the feature map input to the pooling layer, max{O(i × s h : i × s h + k h , j × sw : j × s w + k w )} represents taking the region determined by i and j from the feature map O, and taking the maximum value within this region as the output after pooling.
[0082] The fully connected layer converts the feature map after multiple convolutions and poolings into a prediction result. Assuming that the output of the previous layer is flattened into a vector a, the output of the fully connected layer can be expressed as y = Wa + b, where W is the weight matrix and b is the bias vector. The output layer outputs image features according to the task requirements, and the parameter adjustment depends on the backpropagation algorithm and the gradient descent method during the feature extraction process. The image feature set is I = {i1, i2,..., i c}.
[0083] Audio feature sub-module: Used to extract audio features.
[0084] Perform pre-emphasis processing on the audio signal to enhance the high-frequency components, divide the continuous audio signal into several short-time frames, and window each frame signal to reduce spectral leakage. Perform Fourier transform on each frame signal to obtain the frequency-domain representation where X(k, t) is the k-th frequency component of the t-th frame, x(n, t) is the n-th sampling point of the t-th frame, j represents the imaginary unit, and N is the total number of samples in the frame. Calculate the squared magnitude of each Fourier transform result to obtain the power spectrum, which is used to quantify the energy of each frequency component. The specific formula is P(k) = |X(k, t)| 2 . Apply a set of triangular filters to the power spectrum to convert the linearly spaced frequencies to frequencies on the non-linear Mel scale perceived by the human auditory system. Each filter covers a specific range of frequencies, and its output is the sum of the energies of all frequency components within the corresponding range.
[0085] The specific formula is where S(m) is the logarithm of the output of the m-th Mel filter, H m (k) is the transfer function of the m-th Mel filter, and P(k) represents the power spectrum. The logarithmically compressed Mel spectrum coefficients are further processed by the discrete cosine transform to generate the final audio features. The specific formula is c l is the l-th Mel spectrum coefficient, that is, the audio feature, M represents the number of Mel filters, L is the number of coefficients, and π represents the circumference ratio. The audio feature set is Q = {q1, q2,..., q d}.
[0086] Prediction module: Includes a feature mining sub-module and a trend prediction sub-module.
[0087] Feature mining sub-module: Uses a random forest model to mine the most predictive data features.
[0088] Feature mining is performed using a random forest, which consists of multiple decision trees. By analyzing the features in each decision tree, the importance of the features is evaluated to identify and select the most predictive data features.
[0089] Training data set where α i is a feature vector containing various features such as numerical values, text, images, and audio, and β i is the corresponding label. represents the number of features. The random forest contains T decision trees. For each tree, samples are drawn with replacement from the data set φ to construct a subset, and then a decision tree is trained on this subset. During the construction of the decision tree, each feature helps to make a splitting decision by reducing the impurity. Therefore, the importance of a feature can be measured by its contribution to the reduction of the overall impurity. The Gini index is used to measure the impurity of a node, and the Gini index formula is where f i is the proportion of samples of the i-th class in the node, and C is the number of classes. For each decision tree T t , when splitting a node, the contribution of each feature to the impurity is calculated. Suppose at a certain node, splitting is performed based on feature j, and the contribution I j,t of the feature to the impurity at this node is where Gini before represents the Gini index of the node before splitting, Gini left and Gini right are the Gini indices of the two child nodes after splitting respectively, n left and n right are the number of samples of the two child nodes respectively, and the total number of samples is n = n left + n right .
[0090] The importance of a feature is obtained by summarizing its performance in all trees. If feature j is frequently used for effective splitting in multiple trees, that is, the Gini index is significantly reduced, then this feature is considered very important. The total amount of impurity reduction of feature j during the splitting process is averaged to obtain the importance score According to the importance score, features that meet the score threshold setting are selected from numerical, text, image, and audio features. If there are no features in a certain category of numerical, text, image, and audio that meet the threshold setting, the threshold can be relaxed.
[0091] Trend prediction sub-module: used for trend prediction.
[0092] After determining the most predictive data features, a formula is used for trend prediction. The specific formula is Among them, the numerical feature set is N = {n1, n2,..., n a}, and its weight set is W N = {w n1 , w n2 ,..., w na}; the text feature set is T = {t1, t2,..., t b}, and its weight set is W T = {w t1 , w t2 ,..., w tb}; the image feature set is I = {i1, i2,..., i c}, and its weight set is W I = {w i1 , w i2 ,..., w ic}; the audio feature set is Q = {q1, q2,..., q d}, and its weight set is W A = {w a1 , w a2 ,..., w ad}; considering the time factor, the time period is T cycle , and the external environmental factor set is E = {e1, e2,..., e f}, where the external environmental factors include economic factors, policy factors, sociocultural factors, and technological factors, and its weight set is W E = {w e1 , w e2 ,..., w ef}. The entire formula comprehensively considers the impacts of numerical features, text features, image features, audio features, and external environmental factors on trend prediction by multiplying two parts.
[0093] Visualization module: includes a mapping sub-module and an optimization sub-module.
[0094] Mapping sub-module: Maps data features and prediction results to corresponding nodes and edges in the knowledge graph.
[0095] Mapping data features and prediction results to corresponding nodes and edges in the knowledge graph involves adding additional attributes to the nodes and edges. For each entity v j ∈ V and its corresponding prediction result y j ∈ Y, and for each edge e k ∈ E and its corresponding prediction result y k ∈ Y, define a function f v to update the attributes of the entity and the edge. The specific formula is L = f v (v j , yj |e k ,y k ), where L represents the updated knowledge graph.
[0096] According to the data type and analysis purpose of the knowledge graph, formulate the mapping rules from data to visualization elements. For numerical data, map the numerical size to the vertical coordinate of the line chart; for categorical data, map different categories to different colors and shapes, and adjust the node size and color based on feature importance to reflect their importance. Formulate the mapping rules for different types of data. After formulating the mapping rules, select a suitable visualization tool to visualize the graph. Each categorical graph is visualized according to the mapping rules respectively, providing visualization information support for multiple fields.
[0097] Optimization sub-module: used to optimize the visualization layout.
[0098] Design the user interaction method to improve the efficiency of visualization, such as hovering to display detailed information, clicking to zoom in on a specific area, dragging to re-layout, etc. Optimize the chart parameters to enhance the information transmission efficiency, and select an appropriate layout algorithm to clearly display the graph structure. And adjust the parameters such as the compactness, balance, color saturation, and graphic ratio of the visualization layout according to user feedback.
[0099] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention shall be included in the protection scope of the present invention.
Claims
1. An artificial intelligence cloud data analysis method based on a knowledge graph, characterized in that, Including: S10. Determine the cloud data source for data collection, and conduct preliminary regularization according to data classification to obtain a cloud data set. S20. Construct a knowledge graph according to the cloud data set classification, and calculate the shortest path in the graph to simplify the graph. S30. Extract multi-modal data features from the data set according to the cloud data set classification, and form a classified multi-modal feature set. S40. Conduct intelligent feature mining based on different classified multi-modal feature sets, and perform trend prediction tasks for each classification. S50. Map the trend prediction situations of each classification to the corresponding knowledge graph, update the graph and visualize the graph.
2. The artificial intelligence cloud data analysis method based on a knowledge graph according to claim 1, characterized in that When selecting the cloud data source, one or more data sources should be selected, such as the storage bucket of Amazon Web Services, the object storage of Alibaba Cloud, the data warehouse of Google Cloud Platform, etc. For different data sources, use the corresponding application programming interfaces for data access to ensure the comparability of the data collected from different data sources.
3. The artificial intelligence cloud data analysis method based on a knowledge graph according to claim 1, characterized in that The entities in the knowledge graph can use a named entity recognition tool to identify important entities, or directly determine the entities according to the data pattern and field definition.
4. The artificial intelligence cloud data analysis method based on a knowledge graph according to claim 1, characterized in that For different types of data, corresponding feature extraction algorithms are adopted. Use the principal component analysis method to extract the main components of the numerical values as features; calculate the term frequency-inverse document frequency of the words in the document to extract text features; use a neural network to extract image features; calculate the Mel frequency cepstral coefficients to extract audio features.
5. The artificial intelligence cloud data analysis method based on a knowledge graph according to claim 1, characterized in that According to different classified multi-modal feature sets, use a random forest model for feature mining. The random forest consists of multiple decision trees, and the importance of features is evaluated by analyzing the features in each decision tree to identify and select the most predictive data features.
6. The artificial intelligence cloud data analysis method based on a knowledge graph according to claim 5, characterized in that The importance of a feature is obtained by summarizing its performance in all the trees. If a feature is frequently used for effective splitting in multiple trees, that is, it significantly reduces the Gini index, then this feature is considered very important.
7. The artificial intelligence cloud data analysis method based on a knowledge graph according to claim 1, characterized in that When mapping the data features and prediction results to the corresponding nodes and edges in the knowledge graph, additional attributes need to be added to the nodes and edges. And according to the data type and analysis purpose of the knowledge graph, mapping rules from data to visualization elements need to be formulated.
8. An artificial intelligence cloud data analysis system based on a knowledge graph, characterized in that, Including: Collection module: Used to determine the cloud data source for data collection, and conduct preliminary regularization according to data classification to obtain a cloud data set. Graph module: Construct a knowledge graph according to the cloud data set classification, and calculate the shortest path in the graph to simplify the graph. Feature extraction module: Classify and extract the multi-modal data features in the cloud data set according to the classification, and form a classified multi-modal feature set. Prediction module: Perform intelligent feature mining based on different classified multi-modal feature sets, and carry out the trend prediction tasks for each classification. Visualization module: Used to map the trend prediction situations of each classification to the corresponding knowledge graph, update the graph and visualize the graph.
Citation Information
Patent Citations
Method and system for quickly constructing industry question and answer knowledge base
CN117290489A
Power grid regulation and control knowledge graph construction system, method and program product
CN118211647A
Cloud service data processing method and device, equipment, medium and program product
CN118822584A