Statistical data association mining platform and method based on knowledge graph

Through a statistical data correlation mining platform based on knowledge graph, the challenges in multi-source heterogeneous data processing are solved, efficient data integration, real-time updates and transparent pattern interpretation are achieved, and powerful decision-making support capabilities are provided.

CN120144749AInactive Publication Date: 2025-06-13SHENZHEN ZHIXIN DATA TECHNOLOGY SERVICE CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510282607.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art faces the problems of manual intervention needs, inaccurate entity links, lack of dynamic update mechanisms, low efficiency of traditional algorithms, and opaque interpretation of data mining results when processing multi-source heterogeneous data.

Method used

Provide a statistical data correlation mining platform based on knowledge graphs, including data collection and processing module, knowledge graph construction and update module, feature extraction module, potential pattern mining and report generation module. The platform performs text preprocessing through the BERT model, updates knowledge graphs in real time, extracts path features and subgraph features, and uses the Apriori algorithm to perform pattern mining and association rules generation.

Benefits of technology

It realizes efficient integration and analysis of multi-source heterogeneous data, supports real-time updates and dynamic correlation discovery, improves data mining efficiency and transparency and interpretability of results, and provides valuable decision support information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144749A_ABST
    Figure CN120144749A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, discloses a statistical data association mining platform and method based on a knowledge graph, and aims to provide a comprehensive platform for extracting, integrating, analyzing and explaining complex economic data from various heterogeneous data sources so as to support more effective decision making. Collecting original data from various heterogeneous data sources, and preprocessing the original data to generate a comprehensive data set; secondly, mapping the comprehensive data set into a predefined knowledge graph to form a network structure of nodes and edges, and updating the knowledge graph in real time; then, feature extraction is carried out on the updated knowledge graph, and a comprehensive feature vector is generated; then, the comprehensive feature vector is used as input, a potential mode in the data is obtained, and statistical association is mined through the potential mode; and finally, generating an interpretation report, displaying the potential mode in the knowledge graph and the mined statistical association, and increasing the transparency of mode interpretation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis and mining. More specifically, the present invention relates to a statistical data association mining platform and method based on a knowledge graph. Background Art

[0002] With the development of information technology, enterprises and organizations are faced with a vast amount of information from various data sources. These data sources include a variety of heterogeneous data sources, covering structured, semi-structured, and unstructured data types. In order to make full use of these data resources, a method and technology that can integrate multi-source heterogeneous data and mine valuable information from it are needed. In recent years, the knowledge graph, as an effective data representation and management tool, has been widely used in multiple fields to enhance the connection and understanding between data.

[0003] Existing technologies face multiple challenges in processing multi-source heterogeneous data, including the need for a large amount of manual intervention to adapt to different data formats, which increases the difficulty of preprocessing; inaccurate entity linking affects the quality of the knowledge graph; the lack of a dynamic update mechanism makes it impossible to reflect data changes in real time; traditional pattern discovery and association rule mining algorithms are inefficient on large-scale data sets and are difficult to quickly identify valuable information; the interpretation of data mining results is opaque, which limits the practical application value. Summary of the Invention

[0004] To overcome the above-mentioned defects of the prior art and to achieve the above object, the present invention provides the following technical solutions: A statistical data association mining platform based on a knowledge graph, comprising: A data collection and processing module: collecting raw data from a variety of heterogeneous data sources, performing preprocessing, and performing data integration processing on the preprocessed raw data to generate a comprehensive data set; A knowledge graph construction and update module: mapping the comprehensive data set into a predefined knowledge graph to form a network structure of nodes and edges, and performing real-time updates on the knowledge graph; A feature extraction module: extracting features from the updated knowledge graph, extracting path features and subgraph features to generate a comprehensive feature vector; Potential pattern mining: using the comprehensive feature vector as input, using the Apriori algorithm to obtain potential patterns in the data, and mining statistical associations through potential patterns; A report generation module: generating an explanatory report, displaying the potential patterns in the knowledge graph and the mined statistical associations, and increasing the transparency of pattern interpretation.

[0005] Furthermore, the manner of generating the comprehensive data set includes: Multiple heterogeneous data sources include relational databases, file systems, and web service APIs, and the raw data includes structured data, semi-structured data, and unstructured data; Perform data preprocessing on all the structured, semi-structured, and unstructured raw data collected during the tp period, including data cleaning and text parsing and information extraction. Data cleaning includes removing duplicates, filling in missing values, and standardizing the data format; The methods of text parsing and information extraction include: Extract text data from the raw data after data cleaning, integrate it into a raw text dataset, and perform processing to remove HTML tags and special characters on the raw text dataset; Furthermore, use the Tokenizer tokenization algorithm of the BERT model to split the text sentences in the raw text dataset into words or sub-word units, and generate embedding vectors for each word or sub-word unit; According to the embedding vectors of each word or sub-word unit, use the BERT model for named entity recognition to obtain the entity type labels corresponding to each word or sub-word, and based on the entity type labels, identify and label all the named entities in the raw text dataset; Among them, a named entity refers to a noun phrase with a specific meaning or referring to a specific thing in the text; Among all the identified and labeled named entities, combine all the named entities in pairs to generate entity pair combinations of all the named entities; For each entity pair, construct a context window that includes the pair of entities and their surrounding context information; Use the embedding representation of the context window as the input, and use a linear classifier to classify the relationship of each pair of entities, and output the relationship type corresponding to the entity pair of the context window; if there is no relationship type between the entity pairs output by the linear classifier, it is marked as having no relationship; Perform data integration processing on the raw data of multiple heterogeneous data sources after preprocessing, and merge the raw data after data integration processing to form a comprehensive dataset.

[0006] Furthermore, the data integration processing includes entity alignment and time series alignment; The methods of entity alignment include: According to the obtained entity pairs and their corresponding context windows and relationship types, use NER technology to extract the attribute information of each entity using the context window and relationship type of the entity, where the attribute information of the entity represents specific features or details used to describe the entity; According to the attributes of each entity, use similarity calculation techniques to calculate the similarity between the same type of attributes of every two entities. Based on all the similarity scores between two entities, use weighted summation to obtain the total similarity score between the entities. If the total similarity score is less than or equal to the preset similarity score threshold, it is determined that the two entities are different. If the total similarity score is greater than the preset similarity score threshold, it is determined that the two entities are the same, and all different data determined to be the same entity are merged; The ways of time series alignment include: Based on the preprocessed original data, organize and integrate the heterogeneous data from multiple heterogeneous data sources in the original data in chronological order to form a unified time series format, including: Step 31: Convert the time information in different data sources into a unified timestamp format. According to the timestamp of each data item, allocate it to the corresponding time point or time period, and reorganize it in chronological order to construct a continuous time series; where a data item refers to a data with a timestamp, and a data item contains a timestamp and a corresponding heterogeneous data; Step 32: According to different types of heterogeneous data, adopt feature engineering techniques to extract common feature dimensions, and then merge them into the same time series structure; Step 33: For time series with different sampling frequencies, calculate the distance matrix LJD between every two time series. Each position in the distance matrix represents the Euclidean distance between the x-th time point in time series Xl 1 and the l-th time point in time series Xl 2 ; Step 34: Create a cumulative distance matrix LJC with the same size as the distance matrix. Each position in the cumulative distance matrix is denoted as the cumulative distance LJC(x, l) between x and l; Step 35: Initialize the cumulative distance matrix. For the first row and the first column of the cumulative distance matrix: LJC(1, 1) = LJD(1, 1); for all positions where x > 1: LJC(x, 1) = LJC(x - 1, 1) + LJD(x, 1); for all positions where l > 1: LJC(1, l) = LJC(1, l - 1) + LJD(1, l); Step 36: For each x and l greater than 1, each position LJC(x, l) in the cumulative distance matrix is defined as the distance at the current time point plus the minimum cumulative distance at the previous time point: LJC(x, l) = LJD(x, l) + min{LJC(x - 1, l), LJC(x, l - 1), LJC(x - 1, l - 1)}; Step 37: Starting from the lower right corner of the cumulative distance matrix, move in the direction with the minimum cumulative distance, trace back to the upper left corner, and find the path that minimizes the total distance, which is denoted as the optimal path; Step 38: According to the optimal path, for each pair of indices (x k , l k ) on the optimal path, align the corresponding elements in the time series Xl 1 and Xl 2 , and then construct the aligned time series DQ(Xl 1 ) and DQ(Xl 2 ); where x k and l k are a pair of indices of the k-th point on the optimal path, respectively pointing to the time points in the time series Xl 1 and Xl 2 , x k is an index in the time series Xl 1 , and l k is an index in the time series Xl 2 .

[0007] Furthermore, the method of mapping the comprehensive data set into a predefined knowledge graph to form a network structure of nodes and edges and updating the knowledge graph in real time includes: According to the comprehensive data set, take each entity in the knowledge graph as a node, add all the attributes corresponding to the entity into the node, and use the relationship classification between entities as the edge of each node in the knowledge graph; Define the event source, including: using a trigger to monitor the changes in the data in the database, using a file listener to detect the creation or modification of files in the file system, checking the data updates in the Web service through a callback mechanism, and defining an event when new data is added, existing data is modified, or deleted in each heterogeneous data source; When an event occurs, encapsulate the relevant information into an event object, where the event object includes the event type, timestamp, and relevant entity identifiers; For each event object, preprocess the data therein and perform data integration processing, and then add it to the comprehensive data set; According to the added comprehensive data set, update the knowledge graph, including creating new nodes, adding edges, and updating the attributes of existing nodes.

[0008] Furthermore, the method of extracting features from the updated knowledge graph to generate a comprehensive feature vector includes: Define the extracted feature types as path features and subgraph features; Extract features from the knowledge graph to obtain path features and subgraph features, and perform standardization processing on the extracted path features and subgraph features; Concatenate the standardized path features and subgraph features to form a preliminary comprehensive feature vector; Apply the chi-square test technique to perform feature screening on the preliminary comprehensive feature vector to obtain a comprehensive feature vector.

[0009] Furthermore, the method for extracting the path features includes: Based on the knowledge graph, calculate the centrality score of each node in the knowledge graph ; Among them, represents the centrality score of node jv, jdn represents the number of nodes in the knowledge graph, JD represents the set of all nodes in the knowledge graph, refers to the set operation, indicating all the remaining nodes after removing node jv from the set JD, ju represents a node in the set and represents the shortest path length from node jv to node jv, and the shortest path length represents the minimum number of edges that need to be passed between two nodes; Based on the centrality scores of each node, sort all nodes in descending order according to their scores to form a node sequence; Select the first two nodes from the node sequence in turn to form the first pair of starting point and ending point, and the next two nodes form the second pair of starting point and ending point. This process continues until all nodes in the node sequence are paired. If jdn is odd, it is defined that the last node in the node sequence does not participate in the pairing; For each pair of starting point and ending point, use the graph traversal algorithm to find all paths, and for each path, arrange the nodes and edges on the path in order into a string form to generate a unique code; Regard all the path codes between each pair of starting point and ending point as a document, and regard each path code as a word. Furthermore, use the pre-trained Word2Vec model to map each path code to a vector with a fixed dimension to obtain the path features of each path code.

[0010] Furthermore, the method for extracting the subgraph features includes: Step 71: Based on the knowledge graph, assign a unique label bq(jv) to each node jv in the knowledge graph, and initialize each node as an independent community; Step 72: Initialize the label bq(jv) as the unique identifier of node jv; Traverse all nodes in the knowledge graph. For each node, select the most frequently occurring label among its neighbors as the new bq(jv). If multiple labels have the same maximum frequency among the neighbors, randomly select one of them; Repeat until the labels of all nodes no longer change or reach the preset maximum number of iterations to obtain the initialized community partition; Step 73: Calculate the modularity MK under the current community partition according to the communities after the initialization partition; Step 74: For each node jv in the knowledge graph, take the communities where all the nodes directly connected to node jv are located as the target communities, and calculate the new modularity MK after moving node jv to each target community 1 ; If MK 1 > MK, then move node jv to the community with the maximum new modularity. If MK 1 ≤ MK, then keep the position of node jv unchanged; Evaluate and move all nodes in the knowledge graph in sequence until no node can further increase the modularity by moving or reach the predetermined maximum number of iterations; Step 75: For the community partition after moving nodes, regard each community as a community node, construct a new knowledge graph, and repeat Step S72 to Step S74 on the new knowledge graph; Step 76: Continuously iterate until the modularity cannot be improved to obtain the optimal community partition result, including the community to which each node belongs and the modularity value of the entire knowledge graph; According to the optimal community partition result, based on each community, arrange the nodes and edges within the community in sequence into a string form to generate a unique community code, and then use the pre-trained Word2Vec model to convert each community code into a feature vector to obtain the subgraph features of each community code.

[0011] Furthermore, the way of using the comprehensive feature vector as input and using the Apriori algorithm to obtain the potential patterns in the data includes: Step 81: Collect the comprehensive feature vectors corresponding to ctp time periods. According to the comprehensive feature vectors, use the cp-th percentile as the threshold (such as the 90th percentile), mark the first cp% of the features in the comprehensive feature vector as 1, and mark the features after cp% as 0 to obtain the binary feature vector of the comprehensive feature vector; Step 82: Take the comprehensive feature vector of each time period as a sample. For the binary feature vector of each sample, create an item set, and the item set contains the features marked as 1 in the sample; Combine the item sets of all samples into a transaction database, where each transaction in the transaction database corresponds to one sample, and the transaction is the item set of the sample; Take the ratio of the number of occurrences of each item set in the transaction database to the total number of transactions as the support degree of each item set; Set the minimum support degree, and define the item sets with support degrees greater than the minimum support degree as frequent item sets; Step 83: Traverse the transaction database, and take all item sets with support degrees greater than the minimum support degree as the initial frequent 1-item sets; Step 84: Based on all the frequent (pk - 1)-item sets, combine all the frequent (pk - 1)-item sets pairwise to generate the initial candidate pk-item sets; where pk represents the order of the current iteration; Step 85: For each newly generated candidate pk-item set, check whether all (pk - 1)-item subsets in it are in the frequent (pk - 1)-item sets. If not, remove the candidate pk-item set; if so, retain the candidate pk-item set; Step 86: Traverse the transaction database, calculate the support degree of each retained candidate pk-item set, and retain the candidate pk-item sets with support degrees greater than or equal to the minimum support degree threshold as the frequent pk-item sets; Step 87: Repeat Step S84 to Step S86 until no new frequent item sets can be generated, and take all the obtained frequent item sets as the potential patterns of the knowledge graph.

[0012] Step 88: For each frequent item set corresponding to a potential pattern , generate rules , where means that if , then , means is a non-empty subset of and , means the remaining part after removing from ; For each rule , calculate its confidence degree. The calculation method of the confidence degree is to calculate the ratio of the number of transactions of and to the total number of transactions, and calculate the ratio with the support degree of to obtain the confidence degree of the rule ; Set the minimum confidence degree threshold, and retain all rules greater than or equal to the minimum confidence degree threshold as the statistical associations of the knowledge graph.

[0013] Further, the ways to generate an explanation report, display potential patterns in the knowledge graph and the mined statistical associations, and increase the transparency of pattern explanations include: For each retained rule, calculate its lift, where the calculation method of lift is to divide the proportion of the number of transactions of and in the total number of transactions by the product of the support degrees of and to obtain the lift of rule ; According to the lift of each retained rule, if the lift is greater than 1, it is determined that there is a positive correlation between and within the rule, and if the lift is less than 1, it is determined that there is a negative correlation between and within the rule; Test the independence between and in each retained rule through the chi-square test method to obtain rules with statistical significance; Draw a heat map and use color gradients to represent the support degrees between different item sets; Generate an explanation report and elaborate on the specific meanings of all information in the explanation report through explainable AI technology. The explanation report includes all rules verified by calculating the lift and the chi-square test method and all frequent item sets they contain; Further, a statistical data association mining method based on a knowledge graph is characterized by including: S1: Collect raw data from multiple heterogeneous data sources, perform preprocessing, and perform data integration processing on the preprocessed raw data to generate a comprehensive data set; S2: Map the comprehensive data set to a predefined knowledge graph to form a network structure of nodes and edges, and update the knowledge graph in real time; S3: Extract features from the updated knowledge graph, extract path features and subgraph features to generate a comprehensive feature vector; S4: Use the comprehensive feature vector as input, utilize the Apriori algorithm to obtain potential patterns in the data, and mine statistical associations through the potential patterns; S5: Generate an explanation report, display potential patterns in the knowledge graph and the mined statistical associations, and increase the transparency of pattern explanations.

[0014] The technical effects and advantages of the statistical data association mining platform and method based on the knowledge graph of the present invention: The present invention aims to provide a comprehensive platform for extracting, integrating, analyzing, and interpreting complex economic data from multiple heterogeneous data sources to support more effective decision-making. First, by collecting and cleaning data from multiple heterogeneous data sources, using the BERT model to tokenize text data and generate embedding vectors, and extracting key information, high-quality data input is ensured. Secondly, map the comprehensive data set to a predefined knowledge graph to form a network structure of nodes and edges, and implement a real-time update mechanism to ensure the currency of information; discover potential associations between different entities through entity alignment and time series alignment, support the dynamic update function, and reflect the latest changes in the market or enterprise situation. Then, by extracting features from the updated knowledge graph, generating path features and subgraph features, and combining the chi-square test to screen important features, a comprehensive feature vector is formed, providing a multi-level data representation method, reducing noise, improving model performance and transparency. Next, use the Apriori algorithm to mine frequent item sets, identify potential patterns and generate association rules, evaluate their confidence and support, effectively find patterns and regularities hidden in a large amount of data, and provide valuable decision support information for enterprises. Finally, generate a detailed explanation report to display potential patterns and statistical associations in the knowledge graph, enhancing the transparency and understandability of the report. The present invention realizes the full-process management from data collection to the generation of the final explanation report, ensures data comprehensiveness, structured representation, efficient feature selection, and discovery of implicit patterns, enhances model transparency and practicality, improves data analysis efficiency and accuracy, provides strong support for actual business decisions, and enables enterprises to make more informed choices in a complex and changing market environment. Brief Description of the Drawings

[0015] Figure 1 It is a schematic diagram of the statistical data association mining platform based on the knowledge graph of the present invention; Figure 2 It is a schematic diagram of the statistical data association mining method based on the knowledge graph of the present invention; Figure 3 It is a schematic diagram of the subgraph feature extraction method of the statistical data association mining platform based on the knowledge graph of the present invention; Detailed Embodiment

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0017] Example 1 Please refer to Figure 1 and Figure 3 As shown, the statistical data association mining platform based on a knowledge graph described in this embodiment includes: Data collection and processing module: Collects raw data from multiple heterogeneous data sources, performs preprocessing, integrates the preprocessed raw data, and then generates a comprehensive data set; Knowledge graph construction and update module: Maps the comprehensive data set into a predefined knowledge graph to form a network structure of nodes and edges, and updates the knowledge graph in real time; Feature extraction module: Extracts features from the updated knowledge graph, extracts path features and subgraph features, and generates a comprehensive feature vector; Latent pattern mining: Uses the comprehensive feature vector as input, uses the Apriori algorithm to obtain latent patterns in the data, and statistically associates through latent pattern mining; Report generation module: Generates an explanatory report, displays the latent patterns in the knowledge graph and the mined statistical associations, and increases the transparency of pattern explanations.

[0018] The method for generating the comprehensive data set includes: The multiple heterogeneous data sources include relational databases (such as MySQL, PostgreSQL, SQL Server), file systems, and Web service APIs, and the raw data includes structured data, semi-structured data, and unstructured data; Extract structured data from a relational database through SQL queries (the data stored in a relational database is usually highly structured, organized in the form of tables, each table having fixed columns and rows, and such structured data can be directly obtained through SQL queries), read the file content data in the file system using a file parser, including structured, semi-structured, and unstructured data (CSV, Excel: these formats are usually used to store structured data, similar to the tables in a relational database; JSON, XML: these formats belong to semi-structured data, which have a certain structure (such as a hierarchical structure), but are not as strictly defined as in a relational database, and JSON and XML are widely used for data exchange between web services; text files (such as.txt,.pdf,.docx, etc.): such files usually contain unstructured data, such as free-form text content, and may require additional processing to extract useful information), obtain structured and unstructured data by making HTTP requests to call RESTful APIs or SOAP services (RESTful APIs: in most cases, return data in JSON or XML formats, which are considered semi-structured data. Although there is a certain structure (key-value pairs or tags), it is not as fixed as in a relational database; SOAP services: usually communicate using the XML format, which is also a semi-structured data format), or utilize intelligent data access tools or platforms to automatically identify and adapt to different data source types, simplifying the data access process (automated access tools can identify and adapt to different data source types, whether structured (such as relational databases), semi-structured (such as JSON, XML), or unstructured data (such as text in PDF documents). These tools may integrate various parsers and technologies to process different types of data and convert them into a format suitable for further analysis); For all the structured, semi-structured, and unstructured raw data collected during the tp period, perform data preprocessing, including data cleaning and text parsing and information extraction. Data cleaning includes removing duplicates (detecting and identifying whether there are exactly the same or almost the same records in the data and deleting these duplicates to reduce redundant information), filling in missing values (for data fields with missing values, the mean / median filling method, mode filling method, or prediction model filling method can be used to fill them), and standardizing the data format (converting all data into a unified standard format, for example, the date format is unified to YYYY-MM-DD, and the currency unit is unified to a specific international standard, etc., to ensure the consistency and comparability of the data); The methods of text parsing and information extraction include (specifically, in terms of text parsing and information extraction, use the BERT model to perform word segmentation and generate embedding vectors for text data): Extract text data from the original data after data cleaning, integrate it into the original text dataset, and perform processing to remove HTML tags and special characters on the original text dataset; Furthermore, use the Tokenizer word segmentation algorithm of the BERT model to split the text sentences in the original text dataset into word or sub-word units, and generate the embedding vector QR = [qr 1 , qr 2 ,..., qr i ,..., qr n , where qr i is the embedding vector of the i-th word or sub-word, QR represents the embedding vector of the text sentence, and n represents the number of words or sub-words in the text sentence; Example, a description extracted from an enterprise annual report: "Company A's operating income in 2023 was 5 million yuan, a 10% increase from the previous year, and the R & D expenditure was 800,000 yuan." Perform processing: After removing HTML tags and special characters, the pure text is obtained: "Company A's operating income in 2023 was 5 million yuan, a 10% increase from the previous year, and the R & D expenditure was 800,000 yuan." Use the Tokenizer word segmentation algorithm of the BERT model to split it into word or sub-word units: "Company A / 2023 year / of / operating / income / was / 500 / ten thousand / yuan / than / previous / year / increased / by / 10% / R & D / expenditure / was / 80 / ten thousand / yuan"; For the word "Company A", its embedding vector can be expressed as qri = BERT(Company A), where BERT(.) represents the encoding process of the BERT model for the input word, and the output is the embedding vector of the word; According to the embedding vector of each word or sub-word unit, use the BERT model for named entity recognition to obtain the entity type label corresponding to each word or sub-word. According to the entity type label, identify and label all named entities in the original text dataset (for example, in the text sentence "Zhang San is Li Si's friend, and they often work together.", two named entities can be identified: "Zhang San" and "Li Si"); Among them, a named entity refers to a noun phrase with a specific meaning or referring to a specific thing in the text (they usually represent objects in the real world, including but not limited to people, places, organizations, times, quantities); Among all the identified and labeled named entities, pair up all the named entities to generate entity pairs of all named entities. (For the example "Zhang San is Li Si's friend and they often work together.", there is only one entity pair: "Zhang San" and "Li Si". If there are more entities in the text, all possible pairwise combinations need to be generated. For example, "Zhang San works in Company A and he is friends with Li Si", there are three entities: "Zhang San", "Li Si", and "Company A", and the possible entity pairs include: (Zhang San, Li Si), (Zhang San, Company A), and (Li Si, Company A)); For each entity pair, construct a context window that includes the pair of entities and their surrounding context information. (The context window helps capture the association clues between entities. For example, in "Zhang San works in Company A and he is friends with Li Si", for the entity pair (Zhang San, Li Si), the context window may be the entire sentence "Zhang San works in Company A and he is friends with Li Si"); Use the embedding representation of the context window as input, and use a linear classifier to classify the relationship of each pair of entities, and output the relationship type corresponding to the entity pair in the context window (such as "friend", "colleague"); if there is no relationship type between the entity pairs output by the linear classifier, it is marked as having no relationship; Integrate the raw data of multiple heterogeneous data sources that have been preprocessed, and merge the raw data after data integration processing to form a comprehensive data set; The data integration processing includes entity alignment and time series alignment; Entity alignment and time series alignment are key steps in data integration. During entity alignment, calculate the similarity score based on entity attributes to determine whether two entities are the same. During time series alignment, organize and integrate heterogeneous data in chronological order to ensure the time consistency of the data; The methods of entity alignment include: According to the obtained entity pairs and their corresponding context windows and relationship types, use NER technology to extract the attribute information of each entity using the context window and relationship type of the entity. Among them, the attribute information of the entity represents the specific characteristics or details used to describe the entity, including but not limited to basic personal information (such as name, gender, date of birth, and age), occupation and educational background (such as position, work unit, industry, and educational background), contact information (such as contact number, address), and corresponding time information (time of joining the work unit, time corresponding to the educational background); Example: Suppose we have two different data sources that respectively mention "Company A" and "Company B", and in one data source, it is mentioned that the headquarters of "Company A" is located in Chaoyang District, Beijing, while in another data source, it is also mentioned that the headquarters of "Company B" is located in Chaoyang District, Beijing. By comparing the attributes (headquarters location) of these two entities, we can determine that they may be different representations of the same entity and perform a merging process on them; According to the attributes of each entity, use similarity calculation techniques (for string similarity, use Levenshtein distance, Jaccard similarity coefficient, Cosine similarity; for numerical similarity, directly compare the difference in numerical attributes; for embedding similarity, use the BERT model to generate embedding vectors and calculate the cosine similarity between the embedding vectors) to calculate the similarity between the same type of attributes of every two entities. Based on all the similarity scores between two entities, use weighted summation to obtain the total similarity score between the entities. If the total similarity score is less than or equal to the preset similarity score threshold, then determine that the two entities are different; if the total similarity score is greater than the preset similarity score threshold, then determine that the two entities are the same, and merge all the different data determined to be the same entity; The ways of time series alignment include: Based on the preprocessed original data, organize and integrate various types of data from multiple heterogeneous data sources in chronological order to form a unified time series format, including: Convert the time information in different data sources into a unified timestamp format. According to the timestamp of each data item, allocate it to the corresponding time point or time period, and reorganize it in chronological order to construct a continuous time series; where a data item refers to a data with a specific timestamp, and a data item contains a timestamp and a corresponding data (or a set of observations. The observations can be numerical data (such as thermometer readings (23.5°), stock prices (150.75 yuan)), categorical data (such as weather conditions (sunny, cloudy, rainy)), text data (such as user comments on social media) or composite data (data containing multiple attributes, for example, a record in a health monitoring application may include multiple values such as heart rate (75 bpm), steps (3000 steps))); According to different types of heterogeneous data (such as numerical, categorical, and text), adopt feature engineering techniques to extract common feature dimensions and then merge them into the same time series structure; For time series with different sampling frequencies, calculate the distance matrix LJD between every two time series. Each position in the distance matrix represents the Euclidean distance between the x-th time point in time series Xl 1 and the l-th time point in time series Xl 2 ; Create a cumulative distance matrix LJC that is the same size as the distance matrix. Each position in the cumulative distance matrix is denoted as the cumulative distance LJC(x, l) between x and l. Initialize the cumulative distance matrix. For the first row and the first column of the cumulative distance matrix: LJC(1, 1) = LJD(1, 1); for all x > 1: LJC(x, 1) = LJC(x - 1, 1) + LJD(x, 1); for all l > 1: LJC(1, l) = LJC(1, l - 1) + LJD(1, l). For each x and l greater than 1, each position LJC(x, l) in the cumulative distance matrix is defined as the distance at the current time point plus the minimum cumulative distance at the previous time point: LJC(x, l) = LJD(x, l) + min{LJC(x - 1, l), LJC(x, l - 1), LJC(x - 1, l - 1)}. Starting from the lower right corner of the cumulative distance matrix, move in the direction with the minimum cumulative distance, backtrack to the upper left corner, and find the path that minimizes the total distance, which is denoted as the optimal path. According to the optimal path, for each pair of indices (x k , l k ) on the optimal path, align the corresponding elements in the time series Xl 1 and Xl 2 , and then construct the aligned time series DQ(Xl 1 ) and DQ(Xl 2 ); where x k and l k are the pair of indices of the kth point on the optimal path, respectively pointing to the time points in the time series Xl 1 and Xl 2 , x k is an index in the time series Xl 1 , and l k is an index in the time series Xl 2 . It should be noted that the optimal path is a sequence composed of multiple points (x k , l k ), and each point represents the corresponding position in the time series Xl 1 and Xl 2 . By backtracking the cumulative distance matrix, this path can be found, and each point on the path represents a "match", that is, it is considered that the x 1 th point in Xl k corresponds to the l 2 th point in Xl k (for example, assume there are two time series Xl 1 and Xl 2 : Xl1 =[xl 1,1 ,xl 1,2 ,xl 1,3 ,xl 1,4 , Xl 2 =[xl 2,1 ,xl 2,2 ,xl 2,3 , after calculation, the optimal path obtained may be ljp = [(1,2), (2,2), (3,2), (4,3)], which means: Xl 1 The first point xl in 1,1 corresponds to the first point xl in Xl 2 , Xl 2,1 , the second point xl in 1 corresponds to the first point xl in Xl 1,2 , Xl 2 the third point xl in 2,2 , Xl 1 also corresponds to the second point xl in Xl 1,3 , Xl 2 the fourth point xl in 2,1 , Xl 1 corresponds to the third point xl in Xl 1,4 , according to the optimal path, construct the aligned time series: DQ(Xl 2 ) = [xl 2,3 , xl 1 , DQ(Xl 1,1 ), xl 1,2 , xl 1,3 , xl 1,4 , DQ(Xl 2 ) = [xl 2,1 , xl 2,2 , xl 2,2 , xl 2,3 ); The method of mapping the comprehensive data set into a predefined knowledge graph to form a network structure of nodes and edges and updating the knowledge graph in real time includes: According to the comprehensive data set, in the knowledge graph, each entity is used as a node, and all attributes corresponding to the entity are added to the node, and the relationship classification between entities is used as the edge of each node in the knowledge graph; Define the event source, including: using a trigger to monitor changes in data in the database, using a file listener to detect the creation or modification of files in the file system, checking for data updates in the Web service through a callback mechanism, and defining an event when new data is added, existing data is modified, or deleted in each heterogeneous data source; When an event occurs, relevant information is encapsulated into an event object, where the event object contains the event type, timestamp, and relevant entity identifiers; For each event object, the data therein is preprocessed and integrated, and then added to the comprehensive dataset; Based on the added comprehensive dataset, the knowledge graph is updated, including creating new nodes (entities), adding edges (relationships), and updating the attributes of existing nodes; Example: Suppose we already have a node representing "Company A" in the knowledge graph and know that the company's operating income in 2023 was 5 million yuan. When new data arrives showing that the company's operating income in 2024 was 6 million yuan, we add this information to the knowledge graph, update the attributes of the "Company A" node, and create a new edge representing the change relationship of income between years; The methods for extracting features from the updated knowledge graph to generate comprehensive feature vectors include: Define the extracted feature types as path features (considering all possible paths between two entities in the knowledge graph as features) and subgraph features (identifying important subgraph structures (such as communities, loops) in the knowledge graph as features); Extract path features and subgraph features from the knowledge graph, and perform standardization processing on the extracted path features and subgraph features; Concatenate the standardized path features and subgraph features to form a preliminary comprehensive feature vector; Apply the chi-square test technique to perform feature screening on the preliminary comprehensive feature vector to obtain the comprehensive feature vector; The extraction method of the path features includes: Based on the knowledge graph, calculate the centrality score of each node in the knowledge graph , where represents the centrality score of node jv, jdn represents the number of nodes in the knowledge graph, JD represents the set of all nodes in the knowledge graph, refers to the set operation, indicating the set of all remaining nodes after removing node jv from the set JD, that is, it is the set composed of all nodes except node jv, ju represents a node in the set In the summation process, ju will traverse all nodes different from jv, represents the shortest path length from node jv to node jv, and the shortest path length represents the minimum number of edges that need to be passed between two nodes; Based on the centrality scores of each node, arrange all nodes in descending order according to their scores to form a node sequence; Select the first two nodes from the node sequence in turn to form the first pair of start and end points, and the next two nodes form the second pair of start and end points. This process continues until all nodes in the node sequence are paired. If the number of nodes jdn is odd, it is defined that the last node in the node sequence does not participate in pairing; For each pair of start and end points, use a graph traversal algorithm (such as BFS or DFS) to find all paths. For each path, arrange the nodes and edges on the path in order to form a unique code in string form; Regard all the path codes between each pair of start and end points as a document, and regard each path code as a word. Then use the pre-trained Word2Vec model to map each path code to a vector of a fixed dimension to obtain the path features of each path code; (Example: Suppose there are multiple nodes in our knowledge graph representing different companies, such as "Company A", "Company B", etc. If we want to analyze the cooperation relationship between these companies in a certain field (such as jointly participating in R & D projects), we can measure the cooperation tightness between them by calculating the shortest path from "Company A" to "Company B". For example, "Company A" has established a cooperation relationship with "Company B" through an intermediate company "Company C", then the path is "Company A -> Company C -> Company B". All the nodes and edges on this path are arranged in order to form a unique code in string form, and further converted into a path feature vector); The extraction method of the subgraph features includes: Step 71: Based on the knowledge graph, assign a unique label bq(jv) to each node jv in the knowledge graph, and initialize each node as an independent community (here, the label bq(jv) can be the unique identifier of the node (such as the node ID) to distinguish different nodes); Step 72: Initialize the label bq(jv) as the unique identifier of the node jv (this step ensures that in the initial state, each node has a unique label and is regarded as an independent community); Traverse all nodes in the knowledge graph. For each node, select the label that appears most frequently among its neighbors as the new bq(jv). If multiple labels in the neighbors are tied for the most, randomly select one of them; Repeat until the labels of all nodes no longer change or reach the preset maximum number of iterations to obtain the initialized community partition (this process gradually clusters nodes with similar neighbors into the same community by continuously adjusting the labels of the nodes); Step 73: Calculate the modularity under the current community partition according to the initialized partitioned community ; Among them, represents the modularity under the current community partition, bm represents the number of edges in the knowledge graph, and respectively represent the number of all edges connected to node jv and node ju, and respectively represent the communities to which node jv and node ju belong, represents the Kronecker function. If node jv and node ju belong to the same community, then = 1. If node jv and node ju do not belong to the same community, then = 0 (Modularity is an important indicator to measure the quality of the community structure. It evaluates the degree to which nodes in the knowledge graph are partitioned into communities. The value of modularity ranges from -1 to 1. Generally, a positive value indicates the existence of a significant community structure, while approaching 0 means that the community partition is similar to a random network); Step 74: For each node jv in the knowledge graph, take the communities where all the nodes directly connected to node jv are located as target communities, and calculate the new modularity MK' after moving node jv to each target community respectively; If MK' > MK, then move node jv to the community with the largest new modularity. If MK' ≤ MK, then keep the position of node jv unchanged; Evaluate and move all the nodes in the knowledge graph in sequence until all nodes can no longer increase the modularity by moving or reach the predetermined maximum number of iterations; Step 75: For the community partition after moving the nodes, regard each community as a community node, construct a new knowledge graph, and repeat Step S72 to Step S74 on the new knowledge graph; Step 76: Keep iterating until the modularity cannot be improved, and obtain the optimal community partition result, including the community to which each node belongs and the modularity value of the entire knowledge graph; According to the optimal community partition result, based on each community, arrange the nodes and edges within the community in string form to generate a unique community code, and then use the pre-trained Word2Vec model to convert each community code into a feature vector to obtain the subgraph feature of each community code; (Example: If we want to analyze the enterprise clusters in a specific area and their mutual influences, we can do it according to the community partition of each node in the knowledge graph. For example, identify that "Company A", "Company B" and "Company C" all belong to the same community because they are in the same area and have many cooperation projects. Then, regard this community as a whole and extract the relationship pattern of the internal nodes and edges as the subgraph feature); The ways of using the comprehensive feature vector as input and using the Apriori algorithm to obtain the potential patterns in the data include: In this process, the comprehensive feature vectors of each time period are used as samples to create a transaction database, and the Apriori algorithm is applied to find frequent item sets. Subsequently, association rules are generated based on the frequent item sets, and their confidence and support are evaluated to discover statistically significant associations; Step 81: Collect the comprehensive feature vectors corresponding to ctp time periods. According to the comprehensive feature vectors, use the cp-th percentile as the threshold (such as the 90th percentile), mark the first cp% of the features in the comprehensive feature vector as 1, and mark the features after cp% as 0 to obtain the binary feature vector of the comprehensive feature vector; Step 82: Take the comprehensive feature vector of each time period as a sample. For the binary feature vector of each sample, create an item set that contains the features marked as 1 in the sample; Combine the item sets of all samples into a transaction database. The transaction database contains one transaction corresponding to each sample, and the transaction is the item set of the sample; Take the ratio of the number of times each item set appears in the transaction database to the total number of transactions as the support of each item set; Set the minimum support (set by industry insiders based on experience), and define the item sets with support greater than the minimum support as frequent item sets; Step 83: Traverse the transaction database and use all item sets with support greater than the minimum support as the initial frequent 1-item sets; Step 84: Based on all the frequent (pk - 1)-item sets, combine all the frequent (pk - 1)-item sets pairwise to generate the initial candidate pk-item sets; where pk represents the order of the current iteration; Step 85: For each newly generated candidate pk-item set, check whether all (pk - 1)-item subsets are in the frequent (pk - 1)-item sets. If not, remove the candidate pk-item set; if so, retain the candidate pk-item set; Step 86: Traverse the transaction database, calculate the support of each retained candidate pk-item set, and retain the candidate pk-item sets with support greater than or equal to the minimum support threshold as the frequent pk-item sets; Step 87: Repeat Step S84 to Step S86 until no new frequent item sets can be generated. Take all the obtained frequent item sets as the potential patterns of the knowledge graph; Note that all frequent item sets mined by the Apriori algorithm refer to all items (or combinations of items) in a given transaction database that appear with a frequency exceeding a set minimum support threshold. These frequent item sets are the basis for understanding the potential patterns and statistical associations in the data. A frequent item set refers to a set of items that appear more than a certain preset threshold (i.e., the minimum support) in the transaction database. For example, in a database containing purchase records, if the item set {milk, bread} appears in at least 20% of the transaction records and 20% is the minimum support we set, then {milk, bread} is a frequent item set; Frequent item sets reveal common patterns in the data. For example, they can help identify which products are often purchased together, or in web log analysis, which page access sequences are the most common, etc.; The ways to generate statistical associations through potential patterns include: For each frequent item set xj corresponding to a potential pattern, generate a rule , where yj represents a non-empty subset of xj, yj ⊆ xj and xj - yj ≠ 0, and xj - yj represents the remaining part after removing yj from xj (for example, for the frequent item set {A, B, C}, the rules that can be generated include , , ); For each rule , calculate its confidence. The confidence is calculated by computing the ratio of the number of transactions of xj and yj to the total number of transactions and taking the ratio with the support of xj to obtain the confidence of the rule ; Set a minimum confidence threshold (set by industry insiders based on experience), and retain all rules greater than or equal to the minimum confidence threshold as the statistical associations of the knowledge graph; (Example: Suppose we discover a frequent item set {Company A, R & D investment, high-tech}, which means that Company A has a high R & D investment in the high-tech field. We can generate the rule {Company A, R & D investment} {high-tech}, and calculate its confidence and support. If both the confidence and support are higher than the set threshold, it is considered a meaningful statistical association, indicating that Company A's high R & D investment in the high-tech field is closely related to its business activities); The ways to generate an explanation report to display the potential patterns in the knowledge graph and the mined statistical associations and increase the transparency of pattern explanations include: For each retained rule, calculate its lift. The lift is calculated by taking the ratio of the proportion of the number of transactions of xj and yj to the total number of transactions to the product of the supports of xj and yj to obtain the lift of the rule ; According to the lift of each retention rule, if the lift is greater than 1, it is determined that there is a positive correlation between xj and yj within the rule; if the lift is less than 1, it is determined that there is a negative correlation between xj and yj within the rule. Use the chi-square test method to test the independence between xj and yj in each retention rule, and obtain rules with statistical significance. It should be noted that the chi-square test method is a statistical method used to determine whether there is a significant correlation or independence between two categorical variables. In this method, the chi-square test can help determine whether the discovered rules are statistically significant, that is, these rules are not caused by chance. The test steps of the chi-square test are as follows: Set the null hypothesis H0 (the two variables are independent, that is, there is no association) and the alternative hypothesis H1 (the two variables are not independent and there is some association). Construct a contingency table (a contingency table is a table used to display the cross-classification frequencies of two categorical variables. For example, in market basket analysis, a contingency table can be created to show the situation of purchasing product A and product B). Under the assumption that the two variables are independent, calculate the expected frequency of the retention rule, and then measure the difference between the actual observed frequency and the expected frequency by comparing them, that is, calculate the contribution value of each cell in the contingency table and sum up the contribution values of all cells to obtain the total chi-square statistic. Set the significance level (usually set to 0.05 or 0.01), which is used to determine the standard for rejecting the null hypothesis, and then calculate the degrees of freedom according to the size of the contingency table (generally (number of rows - 1) × (number of columns - 1)). According to the calculated chi-square statistic and degrees of freedom, look up the chi-square distribution table to obtain the corresponding p-value. If the p-value is less than the set significance level, reject the null hypothesis and consider that there is a significant association between the two variables; otherwise, accept the null hypothesis and consider that the two variables are independent. Draw a heat map, using a color gradient to represent the support between different item sets (using support as the association strength between item sets to show the complex relationships between different large data points). Generate a visual interpretation report, and the interpretation report includes all the rules verified by calculating the lift and the chi-square test method and all the frequent item sets they contain. Example: For each retained rule, we calculate its lift. If the lift is greater than 1, it is determined that there is a positive correlation between xj and yj within the rule. For example, for the rule {Company A, R & D investment} {High-tech}. If its lift is 1.5, it indicates a significant positive correlation between Company A's high R & D investment and the development of the high-tech field. In addition, by drawing a heat map and using color gradients to represent the support between different item sets, the visualization effect of the results is further enhanced.

[0019] In this embodiment, by collecting and cleaning data from multiple heterogeneous data sources, using the BERT model to tokenize text data and generate embedding vectors, key information is extracted to ensure high-quality data input; the comprehensive data set is mapped into a pre-defined knowledge graph to form a network structure of nodes and edges, and a real-time update mechanism is implemented to ensure the currency of information. Through entity alignment and time series alignment, potential associations between different entities can be discovered, which is not only convenient for subsequent analysis and understanding, but also supports the dynamic update function to reflect the latest changes in the market or enterprise situation; by extracting features from the updated knowledge graph, path features and sub-graph features are generated and standardized, and important features are selected by combining the chi-square test technology to form a comprehensive feature vector, providing a multi-level data representation method, reducing noise and improving the model performance. At the same time, the transparency and interpretability of the model are improved through the feature selection process, which helps to better understand and apply the analysis results; the Apriori algorithm is used to mine frequent item sets based on the comprehensive feature vector, identify potential patterns, and generate association rules based on the frequent item sets, and their confidence and support are evaluated to effectively find the patterns and rules hidden in a large amount of data, providing valuable decision-making support information for enterprises; a detailed explanation report is generated to display the potential patterns in the knowledge graph and the mined statistical associations, enhancing the transparency and understandability of the report; the whole process management from data collection to the generation of the final explanation report is realized, which not only ensures the comprehensiveness, structured representation and efficient feature selection of data, but also can discover implicit patterns and enhance the transparency and practicality of the model. It not only improves the efficiency and accuracy of data analysis, but also provides strong support for actual business decisions, enabling enterprises to make more informed choices in the complex and ever-changing market environment.

[0020] Embodiment 2 Please refer to Figure 2 As shown, the parts not described in detail in this embodiment can be seen in the description of Embodiment 1. A statistical data association mining method based on a knowledge graph is provided, including: S1: Collect raw data from multiple heterogeneous data sources, perform preprocessing, and perform data integration processing on the preprocessed raw data to generate a comprehensive data set; S2: Map the comprehensive data set into a pre-defined knowledge graph to form a network structure of nodes and edges, and perform real-time updates on the knowledge graph; S3: Extract features from the updated knowledge graph, obtaining path features and subgraph features, and generating a comprehensive feature vector; S4: Use the comprehensive feature vector as input, utilize the Apriori algorithm to obtain potential patterns in the data, and mine statistical associations through the potential patterns; S5: Generate an explanation report, display the potential patterns in the knowledge graph and the mined statistical associations, and increase the transparency of pattern explanations.

[0021] Embodiment III This embodiment publicly provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the running mode of the above-provided statistical data association mining method based on a knowledge graph.

[0022] Since the electronic device introduced in this embodiment is the electronic device used to implement the statistical data association mining method based on a knowledge graph in the embodiments of the present application, based on the statistical data association mining method based on a knowledge graph introduced in the embodiments of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of the present application will not be described in detail here. As long as those skilled in the art implement the electronic device used for the statistical data association mining method based on a knowledge graph in the embodiments of the present application, it falls within the scope of protection of the present application.

[0023] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain a formula that is closest to the actual situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.

[0024] The above are only the preferred embodiments of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for ordinary technical users in the technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. A statistical data association mining platform based on knowledge graph, characterized by: include: Data collection and processing module: collects raw data from various heterogeneous data sources, performs preprocessing, and integrates the preprocessed raw data to generate a comprehensive data set; Knowledge graph construction and update module: maps the comprehensive data set into the predefined knowledge graph to form a network structure of nodes and edges, and updates the knowledge graph in real time; Feature extraction module: extracts features from the updated knowledge graph, extracts path features and subgraph features, and generates a comprehensive feature vector; Among them, based on the knowledge graph, the centrality score of each node is calculated, and then all nodes are arranged into a node sequence in descending order according to their scores. All nodes in the node sequence are paired. For each pair of nodes, a graph traversal algorithm is used to find all paths, and each path is converted into a unique code to obtain the path feature; Based on the knowledge graph, each node is initialized as an independent community and the community division is initialized. According to the community after the initial division, the modularity under the current community division is calculated. The community where all adjacent nodes directly connected to the node are located is taken as the target community. The new modularity after moving the node to each target community is calculated respectively. If the new modularity is greater than the original modularity, the node is moved to the community with the largest new modularity. The iteration is continued to obtain the optimal community division result, and then a unique code is generated for each community. Each community code is converted into a feature vector to obtain the subgraph feature. Latent pattern mining: Using the comprehensive feature vector as input, the Apriori algorithm is used to obtain the latent pattern in the data, and statistical associations are mined through latent patterns; Report generation module: Generates explanation reports to display potential patterns in the knowledge graph and the mined statistical associations, and increases the transparency of pattern explanations.

2. The statistical data association mining platform based on knowledge graph according to claim 1 is characterized in that: The method of generating a comprehensive data set includes: Various heterogeneous data sources include relational databases, file systems, and Web service APIs, and raw data includes structured data, semi-structured data, and unstructured data; All structured, semi-structured and unstructured raw data collected during the tp period are preprocessed, including data cleaning, text parsing and information extraction. Data cleaning includes removing duplicates, filling missing values ​​and standardizing data formats. Text parsing and information extraction methods include: Extract text data from the raw data after data cleaning, integrate it into the raw text data set, and remove HTML tags and special characters from the raw text data set; Then, the Tokenizer segmentation algorithm of the BERT model is used to segment the text sentences in the original text dataset into words or sub-word units, and an embedding vector for each word or sub-word unit is generated; Based on the embedding vector of each word or subword unit, the BERT model is used to perform named entity recognition to obtain the entity type label corresponding to each word or subword. Based on the entity type label, all named entities in the original text dataset are identified and labeled. Among them, named entities represent noun phrases in the text that have specific meanings or refer to specific things; Among all the named entities identified and marked, all the named entities are combined in pairs to generate entity pair combinations of all the named entities; For each entity pair, a context window is constructed that contains the pair of entities and their surrounding context information; Take the embedding representation of the context window as input, use a linear classifier to classify the relationship between each pair of entities, and output the relationship type of the entity pair corresponding to the context window; if there is no relationship type between the entity pairs output by the linear classifier, it is marked as no relationship; The pre-processed raw data from various heterogeneous data sources are subjected to data integration processing, and the raw data after data integration processing are merged to form a comprehensive data set.

3. The statistical data association mining platform based on knowledge graph according to claim 2 is characterized in that: The data integration process includes entity alignment and time series alignment; Entity alignment methods include: According to the obtained entity pairs and their corresponding context windows and relationship types, the attribute information of each entity is extracted using the context windows and relationship types of the entities using the NER technology, wherein the attribute information of the entity represents the specific features or details used to describe the entity; Based on the attributes of each entity, similarity calculation technology is used to calculate the similarity between the same type of attributes of each two entities. Based on all similarity scores between the two entities, a total similarity score between the entities is obtained by weighted summation. If the total similarity score is less than or equal to a preset similarity score threshold, the two entities are determined to be different. If the total similarity score is greater than the preset similarity score threshold, the two entities are determined to be the same, and all different data determined to be the same entity are merged; Ways to align time series include: Based on the preprocessed raw data, the heterogeneous data from various heterogeneous data sources in the raw data are sorted and integrated in chronological order to form a unified time series format, including: Step 31: Convert the time information in different data sources into a unified timestamp format, assign each data item to a corresponding time point or time period according to its timestamp, and reorganize it in chronological order to construct a continuous time series; wherein a data item refers to a data with a timestamp, and a data item includes a timestamp and a corresponding heterogeneous data; Step 32: Based on different types of heterogeneous data, feature engineering techniques are used to extract common feature dimensions, and then merged into the same time series structure; Step 33: For time series with different sampling frequencies, calculate the distance matrix LJD between every two time series. Each position in the distance matrix represents the time series Xl 1 The xth time point in the time series Xl 2 The Euclidean distance between the lth time points in ; Step 34: Create a cumulative distance matrix LJC of the same size as the distance matrix, and each position in the cumulative distance matrix is ​​recorded as the cumulative distance LJC(x,l) between x and l; Step 35: Initialize the cumulative distance matrix. For the first row and the first column of the cumulative distance matrix: LJC(1,1)=LJD(1,1); for all positions where x>1: LJC(x,1)=LJC(x-1,1)+LJD(x,1); for all positions where l>1: LJC(1,l)=LJC(1,l-1)+LJD(1,l); Step 36: For each x and l greater than 1, each position LJC(x,l) in the cumulative distance matrix is ​​defined as the distance of the current time point plus the minimum cumulative distance of the previous time point: LJC(x,l)=LJD(x,l)+min{LJC(x-1,l),LJC(x,l-1),LJC(x-1,l-1)}; Step 37: Starting from the lower right corner of the cumulative distance matrix, move in the direction with the smallest cumulative distance, trace back to the upper left corner, and find the path with the smallest total distance, which is recorded as the optimal path; Step 38: According to the optimal path, for each pair of indexes (x k ,l k ), the time series Xl 1 and Xl 2 The corresponding elements in are aligned, and then the aligned time series DQ (Xl 1 ) and DQ(Xl 2 ); Among them, x k and l k is a pair of indexes of the kth point in the optimal path, pointing to the time series Xl 1 and Xl 2 The time point in x k is the time series Xl 1 An index in k is the time series Xl 2 An index in .

4. The statistical data association mining platform based on knowledge graph according to claim 3 is characterized in that: The method of mapping the comprehensive data set into the predefined knowledge graph to form a network structure of nodes and edges and updating the knowledge graph in real time includes: Based on the comprehensive data set, each entity is a node in the knowledge graph, all the attributes corresponding to the entity are added to the node, and the relationship classification between entities is used as the edge of each node in the knowledge graph; Define event sources, including: using triggers to monitor changes in data in the database, using file listeners to detect the creation or modification of files in the file system, and checking data updates in Web services through callback mechanisms. When new data is added, existing data is modified or deleted in each heterogeneous data source, it is defined as an event. When an event occurs, the relevant information is encapsulated into an event object, where the event object contains the event type, timestamp, and related entity identifiers; For each event object, the data is preprocessed and integrated, and then added to the comprehensive data set; Based on the added comprehensive dataset, the knowledge graph is updated, including creating new nodes, adding edges, and updating the properties of existing nodes.

5. The statistical data association mining platform based on knowledge graph according to claim 4 is characterized in that: The method of extracting features from the updated knowledge graph and generating a comprehensive feature vector includes: Define the extracted feature types as path features and subgraph features; Perform feature extraction on the knowledge graph to obtain path features and subgraph features, and perform standardization on the extracted path features and subgraph features; The standardized path features and subgraph features are concatenated to form a preliminary comprehensive feature vector; The chi-square test technique is applied to perform feature screening on the preliminary comprehensive feature vector to obtain the comprehensive feature vector.

6. The statistical data association mining platform based on knowledge graph according to claim 5 is characterized in that: The method for extracting the path feature includes: Based on the knowledge graph, calculate the centrality score of each node in the knowledge graph ; in, represents the centrality score of node jv, jdn represents the number of nodes in the knowledge graph, JD represents the set of all nodes in the knowledge graph, Refers to the set operation, which means all the nodes remaining after removing the node jv from the set JD. Ju represents the set A node in It represents the shortest path length from node jv to node jv. The shortest path length represents the minimum number of edges that need to be passed between two nodes. Based on the centrality score of each node, all nodes are arranged in descending order according to their scores to form a node sequence; The first two nodes are selected from the node sequence to form the first pair of starting point and end point, and the next two nodes form the second pair of starting point and end point. This process continues until all nodes in the node sequence are paired. If jdn is an odd number, the last node in the node sequence is defined not to participate in the pairing. For each pair of starting point and end point, use the graph traversal algorithm to find all paths. For each path, arrange the nodes and edges on the path in order into a string to generate a unique code. All path encodings between each pair of starting points and end points are regarded as a document, and each path encoding is regarded as a word. Then, the pre-trained Word2Vec model is used to map each path encoding to a vector of fixed dimension to obtain the path features of each path encoding.

7. The statistical data association mining platform based on knowledge graph according to claim 6 is characterized in that: The sub-graph feature extraction method includes: Step 71: Based on the knowledge graph, assign a unique label bq(jv) to each node jv in the knowledge graph, and initialize each node as an independent community; Step 72: Initialize label bq(jv) as the unique identifier of node jv; Traverse all nodes in the knowledge graph. For each node, select the label that appears most frequently among its neighbors as the new bq(jv). If multiple labels in the neighbors are tied for the most, randomly select one of them. Repeat until the labels of all nodes no longer change or the preset maximum number of iterations is reached, and the initialized community division is obtained; Step 73: Calculate the modularity MK under the current community division according to the community after initial division; Step 74: For each node jv in the knowledge graph, the community where all nodes directly connected to the node jv are located is taken as the target community, and the new modularity MK after moving the node jv to each target community is calculated respectively. 1 ; If MK 1 >MK, then move node jv to the community with the largest new modularity. If MK 1 ≤MK, then keep the position of node jv unchanged; All nodes in the knowledge graph are evaluated and moved in sequence until all nodes cannot be moved to increase their modularity further or the predetermined maximum number of iterations is reached; Step 75: For the community division after moving the node, each community is regarded as a community node, a new knowledge graph is constructed, and steps S72 to S74 are repeated on the new knowledge graph; Step 76: Continue iterating until the modularity position cannot be improved, and obtain the optimal community division result, including the community to which each node belongs and the modularity value of the entire knowledge graph; According to the optimal community division results, based on each community, the nodes and edges within the community are arranged in sequence into a string form to generate a unique community code, and then the pre-trained Word2Vec model is used to convert each community code into a feature vector to obtain the subgraph features of each community code.

8. The statistical data association mining platform based on knowledge graph according to claim 7 is characterized in that: The method of using the comprehensive feature vector as input and using the Apriori algorithm to obtain the potential pattern in the data includes: Step 81: Collect the comprehensive feature vectors corresponding to the ctp time periods, and according to the comprehensive feature vectors, use the cpth percentile as a threshold, mark the features before cp% in the comprehensive feature vector as 1, and mark the features after cp% as 0, to obtain a binary feature vector of the comprehensive feature vector; Step 82: Take the comprehensive feature vector of each time period as a sample, and for the binary feature vector of each sample, create an item set, which contains the features marked as 1 in the sample; Combine the item sets of all samples into a transaction database, which contains a transaction corresponding to each sample, and the transaction is the item set of the sample; The ratio of the number of times each item set appears in the transaction database to the total number of transactions is taken as the support of each item set; Set the minimum support and define the item sets with support greater than the minimum support as frequent item sets; Step 83: traverse the transaction database and take all item sets whose support is greater than the minimum support as the initial frequent 1-item sets; Step 84: Based on all frequent (pk-1)-itemsets, all frequent (pk-1)-itemsets are combined in pairs to generate initial candidate pk-itemsets; wherein pk represents the order of the current iteration; Step 85: For each newly generated candidate pk-item set, check whether all (pk-1)-item subsets therein are in the frequent (pk-1)-item set. If not, remove the candidate pk-item set. If so, retain the candidate pk-item set. Step 86: traverse the transaction database, calculate the support of each retained candidate pk-item set, and retain the candidate pk-item sets whose support is greater than or equal to the minimum support threshold as frequent pk-item sets; Step 87: Repeat steps S84 to S86 until no new frequent item sets can be generated, and use all the obtained frequent item sets as potential patterns of the knowledge graph; Step 88: For each potential pattern corresponding to the frequent itemsets , generating rules ,in, If ,but , express is a non-empty subset of and , Indicates from Remove The remaining part after For each rule , calculate its confidence, where the confidence is calculated by calculating and The proportion of transactions to the total number of transactions and The support ratio of is calculated to get the rule confidence level; Set a minimum confidence threshold and retain all rules that are greater than or equal to the minimum confidence threshold as statistical associations of the knowledge graph.

9. The statistical data association mining platform based on knowledge graph according to claim 8, characterized in that: The method of generating an explanation report, displaying the potential patterns in the knowledge graph and the mined statistical associations, and increasing the transparency of the pattern explanation includes: For each retained rule, calculate its lift, where the lift is calculated by and The proportion of transactions to the total number of transactions is and The support product of is used to calculate the ratio, and the rule is obtained. The degree of improvement; According to the lift of each retained rule, if the lift is greater than 1, the rule is judged to be and There is a positive correlation between them. If the lift is less than 1, then the judgment rule and There is a negative correlation between them; The chi-square test was used to test the and independence between them, and obtain statistically significant rules; Draw a heat map and use color gradients to represent the support between different item sets; Generate a visual explanation report, which includes all rules verified by the chi-square test and all the frequent item sets they contain.

10. A statistical data association mining method based on knowledge graph, which is implemented based on the statistical data association mining platform based on knowledge graph according to any one of claims 1 to 9, characterized in that: include: S1: Collect raw data from various heterogeneous data sources, perform preprocessing, and integrate the preprocessed raw data to generate a comprehensive data set; S2: Map the comprehensive dataset into a predefined knowledge graph to form a network structure of nodes and edges, and update the knowledge graph in real time; S3: Extract features from the updated knowledge graph, extract path features and subgraph features, and generate a comprehensive feature vector; S4: Using the comprehensive feature vector as input, the Apriori algorithm is used to obtain the latent patterns in the data, and statistical associations are mined through the latent patterns; S5: Generate an explanation report to display the potential patterns in the knowledge graph and the mined statistical associations, and increase the transparency of pattern explanations.

Citation Information

Cited By

  • Intelligent decision graph construction method based on dynamic time sequence event data

    CN120316271A

  • Security management knowledge graph construction and dynamic updating method based on association modeling

    CN120450018A

  • Tool management appliance storage layout optimization method based on artificial intelligence

    CN120746453A

  • Artificial intelligence-based tool management tool warehouse layout optimization method

    CN120746453B

  • Hierarchical data classification processing and deep correlation analysis system and method based on knowledge graph

    CN120974343A