Data analysis method and system based on artificial intelligence
By adopting artificial intelligence-based methods in data analysis, including building sensitive rule recognition engines, multimodal feature extraction and incremental aggregation calculations, the accuracy and efficiency of data analysis in the prior art are solved, and deeper data insights and value extraction are achieved.
Patent Information
- Application Number
- CN202510600527.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing technology has data one-sidedness and incompleteness in data analysis, making it difficult to deeply explore the complex relationships between data, cannot provide deep insights and value information, and the analysis is inefficient.
Using artificial intelligence-based data analysis method, we construct a sensitive rule recognition engine to perform abnormal detection, perform multimodal feature extraction and self-attention calculation, realize incremental aggregation calculation, and self-supervised learning of graph structure through the correlation map between data, and finally data optimization is carried out on real-time aggregation analysis data.
It improves the accuracy and efficiency of data analysis, can more accurately identify sensitive information in the data, deeply explore the complex relationships between data, and provide deep insights and value information.
Smart Images

Figure CN120123700A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and in particular, to a data analysis method and system based on artificial intelligence. Background Art
[0002] With the development of the Internet, Internet of Things, and other digital technologies, a large amount of data has been continuously generated and accumulated, and data analysis based on artificial intelligence has become the core driving force in the field of modern data analysis.
[0003] Most of the existing technologies for data analysis collect business requirement information from multiple channels through various data collection tools and platforms, and perform preliminary data cleaning and integration. Traditional statistical analysis and machine learning methods are used to analyze the collected data to generate preliminary analysis reports.
[0004] Although the existing solutions meet the basic analysis requirements of data analysis to a certain extent, the existing solutions often rely on a single data source and lack comprehensive integration of multi-source data, resulting in one-sidedness and incompleteness of the data; traditional statistical analysis and machine learning methods are difficult to deeply explore the complex relationships between data and cannot provide in-depth insights and valuable information; the existing solutions usually adopt a batch processing method and cannot achieve real-time data processing and dynamic adjustment, resulting in limited data analysis efficiency.
[0005] Therefore, how to improve data analysis efficiency and data analysis accuracy has become an urgent problem to be solved. Summary of the Invention
[0006] The present invention provides a data analysis method based on artificial intelligence, and its main purpose is to solve the problems of low data analysis accuracy and low data analysis efficiency.
[0007] In a first aspect, to achieve the above object, a data analysis method based on artificial intelligence provided by the present invention includes: Obtain historical analysis data, and construct a sensitive rule recognition engine according to the historical analysis data; Obtain data to be analyzed, and perform anomaly detection processing on the data to be analyzed according to the sensitive rule recognition engine to obtain first target analysis data; Perform multi-modal feature extraction on the first target analysis data to obtain a multi-modal feature set, and perform self-attention calculation on the multi-modal feature set to obtain target features of the first target analysis data; Perform incremental aggregation calculation on the first target analysis data according to the target features to obtain real-time aggregation analysis data; Construct an association graph between data according to the data to be analyzed, and perform graph structure self-supervised learning on the data to be analyzed according to the association graph between data to obtain a data analysis target; Optimize the real-time aggregated analysis data according to the data analysis objective to obtain second target analysis data.
[0008] In a second aspect, the present invention also provides an artificial intelligence-based data analysis system, which includes: A rule engine construction module, configured to obtain historical analysis data and construct a sensitive rule recognition engine according to the historical analysis data; An anomaly detection and processing module, configured to obtain data to be analyzed, perform anomaly detection and processing on the data to be analyzed according to the sensitive rule recognition engine, and obtain first target analysis data; A multi-modal attention calculation module, configured to perform multi-modal feature extraction on the first target analysis data to obtain a multi-modal feature set, and perform self-attention calculation on the multi-modal feature set to obtain target features of the first target analysis data; An incremental aggregation module, configured to perform incremental aggregation calculation on the first target analysis data according to the target features to obtain real-time aggregated analysis data; A self-supervised learning module, configured to construct an inter-data association graph according to the data to be analyzed, perform graph structure self-supervised learning on the data to be analyzed according to the inter-data association graph, and obtain a data analysis objective; A data optimization module, configured to optimize the real-time aggregated analysis data according to the data analysis objective to obtain second target analysis data.
[0009] By constructing a sensitive rule recognition engine, the present invention can accurately identify sensitive information in data, greatly reducing the time cost of manual screening and analysis. Through anomaly detection and processing, potential security risks can be effectively filtered out, ensuring the security of data during transmission, storage, and use, and improving the accuracy of analysis. Through multi-modal feature extraction, feature information of multiple modalities such as images and texts can be simultaneously captured from the first target analysis data, forming a comprehensive feature representation, integrating complementary information between different modalities, avoiding the limitations of single-modal features, and making the target features more expressive and generalization-capable. Performing incremental aggregation on the first target analysis data avoids recalculating the entire data set, significantly reducing the computational resources and time cost. By constructing an inter-data association graph, entities and their relationships in the data to be analyzed are presented in an intuitive graphical manner, improving data visualization. Graph structure self-supervised learning can deeply mine potential features in the graph, more accurately capture the internal laws between data, and reduce analysis biases caused by ignoring data associations, significantly improving the accuracy of data analysis. Description of the Drawings
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0011] Figure 1 A schematic flowchart of a data analysis method based on artificial intelligence provided by an embodiment of the present invention; Figure 2 A schematic flowchart of performing anomaly detection processing on the data to be analyzed according to the sensitive rule recognition engine provided by an embodiment of the present invention; Figure 3 A schematic flowchart of performing incremental aggregation calculation on the first target analysis data according to the target feature provided by an embodiment of the present invention; Figure 4 A schematic flowchart of performing graph structure self-supervised learning on the data to be analyzed according to the data interrelationship graph provided by an embodiment of the present invention; Figure 5 A schematic block diagram of a data analysis system based on artificial intelligence provided by an embodiment of the present invention; The implementation, functional features, and advantages of the objectives of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0012] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, and to fully understand and implement how the present disclosure uses technical means to solve technical problems and achieve corresponding technical effects, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The embodiments of the present disclosure and the various features in the embodiments can be combined with each other without conflict, and the formed technical solutions are all within the protection scope of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0013] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0014] An embodiment of the present application provides a data analysis method based on artificial intelligence. The data analysis method based on artificial intelligence can be executed by software or hardware installed on a terminal device or a server device. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0015] Refer to Figure 1 As shown, it is a schematic flowchart of a data analysis method based on artificial intelligence provided by an embodiment of the present invention. In this embodiment, the data analysis method based on artificial intelligence includes: S1. Obtain historical analysis data, and construct a sensitive rule recognition engine according to the historical analysis data.
[0016] In an embodiment of the present invention, the historical analysis data may refer to structured or unstructured data related to sensitive information recognition accumulated in the past, and is used to construct, train, and verify a sensitive rule recognition engine. The historical analysis data may include user behavior logs, sensitive event reports, data classification labels, manually labeled sensitive data categories, etc.
[0017] In an embodiment of the present invention, the historical analysis data can be collected through an internal system. For example, operation logs of servers, databases, and application systems are collected through a log management tool such as ELK Stack, or known sensitive rule data is obtained through external data access.
[0018] In an embodiment of the present invention, constructing a sensitive rule recognition engine according to the historical analysis data includes: Extract sensitive features from the historical analysis data to obtain sensitive information features; Screen the rules in the preset sensitive information recognition rule set according to the sensitive information features to obtain the corresponding initial sensitive information recognition rules; Optimize the combination of the initial sensitive information recognition rules to obtain the target sensitive information recognition rules; Fill the preset rule engine template according to the target sensitive information recognition rules to obtain a sensitive rule recognition engine.
[0019] In the embodiments of the present invention, technologies such as natural language processing and regular expressions can be used to extract patterns or structures that may represent sensitive information from historical analysis data. For example, by performing data word segmentation processing on the historical analysis data, multiple strings are obtained, that is, the continuous text sequence is converted into a structured word sequence, and each string is matched one by one through a predefined regular expression based on sensitive information features to obtain sensitive information features. The sensitive information features may include feature data with fixed formats such as ID card numbers, mobile phone numbers, and bank card numbers.
[0020] Among them, the similarity between the sensitive information features and the sensitive information recognition rules can be calculated through similarity algorithms such as cosine similarity and Jaccard similarity to obtain the initial sensitive information recognition rules.
[0021] Specifically, the screening of the rules in the preset sensitive information recognition rule set according to the sensitive information features to obtain the corresponding initial sensitive information recognition rules includes: Extract rule keywords from the sensitive information recognition rule set to obtain rule keywords; Calculate the similarity between the rule keywords and the sensitive information features to obtain a rule matching degree; Compare the rule matching degree with a preset matching threshold, determine the target rule keywords corresponding to the sensitive information features, and select the sensitive information recognition rules corresponding to the target rule keywords as the initial sensitive information recognition rules.
[0022] In detail, keywords can be directly extracted from the sensitive information recognition rules according to the preset keyword list rules, such as keywords like ID card numbers and bank cards. The sensitive information recognition rules are logical patterns used to detect, match, and identify sensitive information in text, including text matching rules, keyword filtering rules, etc.
[0023] Among them, the text matching rule is a logical pattern based on string matching or semantic similarity calculation, and the keyword filtering rule is to scan the data through a predefined sensitive word library to identify text fragments containing sensitive words, which can be used for bad information in the data.
[0024] Specifically, convert the sensitive information features and rule keywords into vectors, calculate the cosine value of the vector angle to obtain the rule matching degree. If the rule matching degree is greater than the preset matching threshold, it is considered that the sensitive information recognition rule is related to the sensitive information features, record the rule keyword with the highest matching degree, and select the sensitive information recognition rule corresponding to the target rule keyword as the initial recognition rule.
[0025] In the embodiment of the present invention, the sensitive information recognition rules are combined to generate different rule combinations. Specifically, by traversing all the sensitive information recognition rules in the rule set, enumerate all possible rule combination methods. For example, based on the logical relationship of the rules, construct combination expressions of "AND", "OR", and "NOT" to form an initial rule combination set.
[0026] The present invention can use the simulated annealing algorithm to iteratively optimize the rule combination to optimize the performance of the rule combination. Randomly select an initial combination from the initial rule combination set as the current solution, and calculate its fitness value. For example, based on the recall rate, false alarm rate, or comprehensive evaluation index of the rule combination, on the basis of the current solution, generate a neighborhood solution by randomly adjusting the rule combination, such as adding or deleting rules, modifying rule weights or logical relationships. If the fitness of the neighborhood solution is better than the current solution, directly accept it; if it is worse than the current solution, accept it with a certain probability to jump out of the local optimum. As the number of iterations increases, reduce the probability of accepting the inferior solution, and finally converge to the global optimum or approximate optimum solution to obtain the target sensitive recognition rule.
[0027] Among them, use the existing rule engine according to the optimized target sensitive recognition rule, and apply the rule to the data to be analyzed to identify the sensitive information therein. For example, if the rule combination is "ID number matching rule AND mobile phone number matching rule AND non-specific whitelist users", the rule engine will sequentially perform ID number matching and mobile phone number matching on the data, and combine the whitelist filtering logic to output the final sensitive information recognition result.
[0028] S2. Obtain the data to be analyzed, and perform anomaly detection processing on the data to be analyzed according to the sensitive rule recognition engine to obtain the first target analysis data.
[0029] In the embodiment of the present invention, the data to be analyzed includes structured database records, unstructured log files, and multi-source heterogeneous data such as user behavior data and transaction records. By performing structured parsing on the data features and extracting sensitive information features, a rule-based sensitive information recognition engine is constructed to accurately identify the sensitive data in the data to be analyzed. Finally, the generated first target analysis data has been significantly improved in terms of data integrity and accuracy.
[0030] The present invention can also obtain the data to be analyzed from a public dataset. The public dataset can be a language dataset, which contains the experimental data of analysis experiments conducted on multiple experimental platforms. For example, for the business scenario of data compliance analysis of a customer company, a multilingual text dataset can be selected as the data source. This dataset contains the experimental data of multiple experimental platforms in the fields of natural language processing, text classification, etc. By extracting text fragments related to the target language from this dataset, a dataset to be analyzed covering a multilingual environment can be quickly constructed, improving the data acquisition efficiency.
[0031] As Figure 2 shown, in the embodiment of the present invention, the abnormal detection process of the data to be analyzed by the sensitive rule recognition engine to obtain the first target analysis data includes: Identifying sensitive data in the data to be analyzed according to the sensitive rule recognition engine to obtain a set of sensitive data; Performing a verification process on the set of sensitive data to obtain a set of target sensitive data; Calculating the allocation weight of each target sensitive data in the set of target sensitive data; Calculating the allocation weight of each target sensitive data by using the following formula: Wherein, represents the allocation weight of the th target sensitive data , represents the preset sensitive center weight of the set of target sensitive data, represents the transpose, represents the th preset sensitive center offset of the target sensitive data, represents the total number of target sensitive data; Performing regression calculation on the target sensitive data according to the allocation weight to obtain the sensitivity coefficient of the target sensitive data; The sensitivity coefficient of each target sensitive data can be calculated by using the following formula: Wherein, represents the sensitivity coefficient, represents the preset intercept, represents the th target sensitive data 's allocation weight, represents the preset error factor; Taking the target sensitive data with the sensitivity coefficient less than the preset sensitivity threshold as the first target analysis data.
[0032] Specifically, the data to be analyzed is scanned line by line using the rules defined in the sensitive rule recognition engine, such as regular expressions and pattern matching rules, to identify the data that conforms to sensitive features. The identified sensitive data is classified and labeled according to types such as ID numbers, mobile phone numbers, bank card numbers, etc., to form a sensitive data set.
[0033] The present invention verifies the initially identified sensitive data, removes false alarm data, and ensures the accuracy and reliability of the data. This can be achieved by checking whether the sensitive data conforms to predefined format specifications, such as an ID number should be 18 digits and a mobile phone number should be 11 digits.
[0034] Specifically, weighted regression models such as weighted linear regression and weighted logistic regression can be used. The target sensitive data is used as the independent variable, and weights are assigned as regression coefficients for regression calculation to obtain sensitive coefficients. Data with a lower sensitivity level is selected as the first target analysis data according to a preset sensitivity threshold.
[0035] In the embodiment of the present invention, by constructing a sensitive rule recognition engine, sensitive information in the data can be accurately identified, greatly reducing the time cost of manual screening and analysis. Through anomaly detection processing, potential security risks can be effectively filtered out, ensuring the security of the data during transmission, storage, and use, and improving the accuracy of analysis.
[0036] S3. Perform multi-modal feature extraction on the first target analysis data to obtain a multi-modal feature set, and perform self-attention calculation on the multi-modal feature set to obtain the target features of the first target analysis data.
[0037] In the embodiment of the present invention, the first target analysis data includes multi-modal data of picture type and text type. By performing feature extraction in different dimensions and self-attention calculation on data of different modalities, more comprehensive target features can be obtained.
[0038] In the embodiment of the present invention, performing multi-modal feature extraction on the first target analysis data to obtain a multi-modal feature set includes: Performing modal type recognition on the first target analysis data to obtain a modal classification result; If the modal classification result is image modal data, extract the position features and grid features of each pixel point in the image modal data, and perform weighted splicing on the position features and grid features to obtain the image features of the image modal data, and use the image features as the target features; If the modal classification result is text modal data, calculate the word embedding feature sequence of the text data, and perform feature encoding on the word embedding feature sequence to obtain the text features of the text modal data, and use the text features as the target features; Aggregate the target features to obtain a multi-modal feature set of the first target analysis data.
[0039] In an embodiment of the present invention, a pre-trained modal classification model, such as a deep learning-based classifier or a simple rule matching method, can be used to identify the modal type of the first target analysis data. According to the detection result, the data can be classified into an image modality, a text modality, etc.
[0040] Specifically, in an embodiment of the present invention, the pre-trained R-CNN (Regions with Convolutional Neural Networks) can be used to detect pixel points of image data, and the position features of each pixel point in the image data can be obtained. However, since the object position features do not capture multi-granularity feature information such as color and scene in the image data, the image data can be evenly divided into multiple grids, and the pre-trained residual network can be used to perform residual convolution on the grids to obtain the grid features of the image data.
[0041] Specifically, for text modality data, text tokenization is performed, the word embedding feature vector of each token is calculated, and a word embedding feature sequence is obtained. For example, the GloVe can be used to calculate the word embedding feature vector of each token, and then the word embedding feature sequence is input into a pre-constructed recurrent neural network encoder for feature encoding to extract deeper semantic information and obtain the text features of the text modality data.
[0042] In an embodiment of the present invention, the self-attention calculation of the multi-modal feature set to obtain the target features of the first target analysis data includes: Perform scale adjustment on the multi-modal feature set to obtain a normalized feature set; Perform cross-modal feature fusion on the normalized feature set to obtain a modal fusion feature; Perform a convolution operation on the modal fusion feature to obtain an enhanced feature representation, and perform pooling dimensionality reduction processing on the enhanced feature representation to obtain a compressed feature representation; Perform an activation function process on the compressed feature representation to obtain an activation feature representation, and perform multi-layer feature fusion on the activation feature representation to obtain a fusion feature representation; Calculate the self-attention matrix of the fusion feature representation, and calculate the attention weighted features of different attention heads according to the self-attention matrix by using a pre-constructed multi-head attention mechanism; Use the following formula to calculate the attention weighted features of different attention heads: Among them, represents the attention weighted feature, represents the fusion feature representation, Represents a preset scaling factor. Represents the query matrix of the th attention head in the self-attention matrix. Represents the th attention head's key matrix in the attention matrix. Represents the th attention head's value matrix in the attention matrix. Represents transpose. Represents a preset attention factor. Represents activation function. Performs a fully connected dimensionality reduction process on the attention-weighted features to obtain the target features of the first target analysis data.
[0043] In the embodiments of the present invention, the multi-modal feature set can be scaled through normalization techniques such as Z-Score normalization to obtain a feature representation with unified scale. The cross-modal feature fusion includes temporal feature fusion and spatial feature fusion. The temporal feature fusion can use dynamic time warping (DTW) or time interpolation method to unify multi-source timestamps so as to map the normalized feature set to a unified vector space; the spatial feature fusion can eliminate the modality gap through a feature mapping network such as adversarial domain adaptation DAE or graph attention (GAT) mechanism, thereby constructing a joint feature space of the first target analysis data.
[0044] Specifically, the feature weights can be calculated according to the multi-head self-attention mechanism, where the multi-head mechanism captures different feature space associations in parallel, thereby mining the deep semantic associations of multi-source modal features and generating highly reliable target features.
[0045] In the embodiments of the present invention, the self-attention matrix includes a query matrix, a key matrix, and a value matrix. The self-attention matrix can be obtained by multiplying a preset parameter matrix with the fused feature representation. The attention calculation is performed on the self-attention matrix by each attention head in the multi-head attention mechanism to obtain the attention-weighted features of each attention head. The attention features of each attention head can jointly focus on the feature information from other regions, which is beneficial to enhancing the feature representation. For example, in text data, there are usually some words without semantics. Through the multi-head attention mechanism, the words with specific semantics can be strengthened and the words without specific meaning can be weakened to obtain more accurate target features.
[0046] In the embodiments of the present invention, through multi-modal feature extraction, the feature information of multiple modalities such as images and texts can be simultaneously captured from the first target analysis data to form a comprehensive feature representation, integrating the complementary information between different modalities, avoiding the limitations of single-modal features, and making the target features more expressive and generalizable.
[0047] S4. Perform incremental aggregation calculation on the first target analysis data according to the target feature to obtain real-time aggregated analysis data.
[0048] In the embodiments of the present invention, the first target analysis data is incrementally updated according to the target feature to obtain data in real time, and then the incremental data is aggregated and analyzed, such as operations such as summarization and statistics, to update the aggregated analysis result, avoiding repeated calculation of all data and improving the analysis efficiency.
[0049] As Figure 3 shown, in the embodiments of the present invention, the performing incremental aggregation calculation on the first target analysis data according to the target feature to obtain real-time aggregated analysis data includes: Obtain an incremental data packet, perform data parsing on the incremental data packet to obtain incremental data fields; Perform field mapping on the incremental data fields and the target feature to obtain an incremental feature mapping table; Perform feature matching on the first target analysis data according to the incremental feature mapping table to obtain valid incremental data that meets the preset feature threshold; Perform time series analysis on each data stream in the valid incremental data to obtain a timestamp sequence; Perform data interpolation processing on the timestamp sequence to obtain a target incremental data set, and segment the target incremental data set to obtain a time window data set; Perform aggregation calculation on the time window data set to obtain real-time aggregated analysis data; Perform aggregation calculation on the time window data set using the following formula: where represents the real-time aggregated analysis data, represents a preset window, represents the total number of the time window data sets, represents the th time window data set within the window.
[0050] In the embodiments of the present invention, the incremental data packet refers to a scenario where data is continuously updated, and new data will continuously arrive. The new data exists in the form of data packets and may come from various data sources, such as sensors, business systems, etc.
[0051] Among them, the incremental data packet is usually encapsulated in a specific data format. In order to extract useful information, it is necessary to parse the incremental data packet, disassemble the incremental data packet according to a predetermined rule, and a parsing library such as the json module of Python or Jackson of Java can be used to parse the data string in the incremental data packet into a dictionary or an object, and extract the required fields such as user ID and transaction amount from the parsed data.
[0052] Specifically, establish a corresponding relationship between the incremental data fields and the target feature fields to ensure that the incremental data can be correctly input into the preset target system and avoid errors caused by field mismatch. Static mapping or dynamic mapping can be used. The static mapping is based on predefined field mapping rules such as Excel tables and configuration files to clarify the corresponding relationship between the incremental fields and the target fields, and store them in the form of tables or configuration files, thereby generating an incremental feature mapping table.
[0053] Among them, the dynamic mapping automatically infers the mapping relationship according to the data context or metadata (such as field name similarity and data type matching). Through natural language processing (NLP) technology, identify the semantic association between the incremental field and the target feature field, and use the mapping rule to generate the corresponding relationship between the incremental field and the target feature field, thereby obtaining an incremental feature mapping table. The incremental feature mapping table facilitates quickly locating and operating on data related to the target feature in the subsequent data processing process. For example, if the target feature is "product category", the mapping table will indicate which field in the incremental data represents the product category and information such as the value range of this field.
[0054] Specifically, associate and compare the incremental data with the existing first target analysis data, and filter out the data that matches the target feature and meets the preset feature threshold. For example, in sales data analysis, it may be set that only the data with a sales amount greater than a certain value is considered valid incremental data. Through feature matching, the valid incremental data that meets this threshold can be filtered out from the incremental data.
[0055] In the embodiment of the present invention, the time series analysis is a statistical method used to study the law of data changing over time. By analyzing the valid incremental data, the time information corresponding to each data point can be extracted to form a time stamp sequence. For example, in a meteorological dataset, each data point contains information such as temperature and humidity at a certain moment. Through time series analysis, the time stamp sequence of these data points can be obtained to understand the changing trend of meteorological data over time.
[0056] Specifically, since there may be uneven time intervals or missing data during the data collection process, it is necessary to perform data interpolation on the time stamp sequence. The data interpolation estimates the values of unknown data points through mathematical methods based on the values of known data points, so as to make the time series more complete and continuous. For example, in stock price data, if the data at certain moments is missing, the stock prices at these moments can be estimated through the interpolation method.
[0057] Among them, the time window refers to dividing the continuous time series into several time periods of a fixed length. The data within each time period constitutes a time window. Aggregating and calculating the time window data set is to summarize and statistically analyze multiple data points according to certain rules to obtain higher-level data information.
[0058] In the embodiment of the present invention, incremental aggregation of the first target analysis data is performed, which avoids repeated calculation of the entire data set, greatly reduces the computing resources and time costs, and at the same time obtains a more rich and comprehensive aggregated analysis data.
[0059] S5. Construct a data interrelationship graph based on the data to be analyzed, and perform graph structure self-supervised learning on the data to be analyzed according to the data interrelationship graph to obtain a data analysis target.
[0060] In the embodiment of the present invention, the entities in the data to be analyzed are abstracted as nodes in the graph, and the relationships between entities are transformed into edges to construct a data interrelationship graph, which intuitively shows the complex relationships between the data and provides a structured basis for subsequent analysis. Using the constructed data interrelationship graph for graph structure self-supervised learning can enable the self-supervised learning model to learn the internal patterns and rules of the data and realize the extraction of deep semantic features of the data.
[0061] In the embodiment of the present invention, the constructing a data interrelationship graph based on the data to be analyzed includes: Performing attribute parsing on the data to be analyzed to obtain data entity attributes and data interaction relationships, and generating a structured data set according to the data entity attributes and the data interaction relationships; Extracting the association features in the structured data set, and generating a data entity set and an association rule set of the data to be analyzed according to the association features; Taking each entity in the data entity set as a node and taking each association rule in the association rule set as an edge weight to construct a data interrelationship graph.
[0062] In the embodiments of the present invention, by deeply analyzing the data to be analyzed, identifying the entities and entity attributes contained therein, the data fields of the data to be analyzed can be extracted, and data entity attributes can be identified from the data fields. For example, in e-commerce transaction data, the entities can be users, products, orders, etc. The user entity has attributes such as age, gender, and purchase preferences, and the product entity has attributes such as name, price, and category. The data interaction relationships can use preset association rule mining algorithms (such as Apriori, FP-Growth) to discover the association relationships between different attributes in the data, such as which products are often purchased together.
[0063] Specifically, the parsed entity attributes and interaction relationships are sorted and organized according to a predetermined structure to form a structured data set, which is convenient for subsequent storage, query, and analysis. For example, the above e-commerce transaction data can be sorted into a table form, with each row representing a record and containing fields such as user ID, product ID, and transaction time.
[0064] Based on predefined association rules, the structured data set is scanned and analyzed to extract association features that meet the rules. All entities involved in the structured data set are summarized into a data entity set, and the extracted association features are sorted into an association rule set. The data entity set contains all entities participating in the association analysis, and the association rule set clarifies the association methods and strengths between entities.
[0065] Among them, each entity in the data entity set is used as a node in the data association graph, and each association rule in the association rule set is used to define the edge weights between the nodes. According to the definitions of the nodes and edge weights, the nodes and edges are connected to obtain a data association graph.
[0066] As Figure 4 shown, in the embodiments of the present invention, the graph structure self-supervised learning of the data to be analyzed based on the data association graph to obtain a data analysis target includes: Performing a graph traversal on the data association graph to extract all graph node pairs in the data association graph; Performing node pair screening on the graph node pairs to obtain a first node pair and a second node pair, and using the first node pair and the second node pair as positive sample node pairs and negative sample node pairs respectively; Generating a contrast loss function according to the positive sample node pairs, the negative sample node pairs, and a preset contrast factor; Using the following formula to generate a contrast loss function according to the positive sample node pairs, the negative sample node pairs, and a preset contrast factor: Among them, represents the A positive sample node pair, Indicating the th positive sample node pair, Indicating the th negative sample node pair, Indicating a preset comparison factor, Indicating the total number of negative sample node pairs, Indicating the said contrast loss function; Minimize the said contrast loss function to obtain a minimum loss function; Minimize the said contrast loss function using the following formula to obtain a minimum loss function: Wherein, Indicating the th positive sample node pair, Indicating the th positive sample node pair, Indicating the th negative sample node pair, Indicating a preset comparison factor, Indicating the total number of negative sample node pairs, Indicating the said minimum loss function; Optimize a preset self-supervised learning model according to the said minimum loss function to obtain a graph analysis model; Construct a computable inference graph of the data to be analyzed according to the said graph analysis model; Perform graph inference on the said computable inference graph to obtain a data analysis objective.
[0067] In the embodiments of the present invention, breadth-first search (BFS) or depth-first search (DFS) can be used to traverse the association graph between data to obtain all node pair combinations. For example, for a graph G=(V, E), where V is the set of nodes and E is the set of edges, traverse all possible node pairs (u, v), where u, v ∈ V.
[0068] Among them, the node pairs are divided into two categories, namely the first node pair and the second node pair. The first node pair refers to the nodes connected by an edge, and the second node pair refers to the remaining nodes not directly connected by an edge. According to the nodes connected by an edge, positive sample node pairs are formed, and the remaining nodes not directly connected by an edge form negative sample node pairs. Specifically, the positive sample node pairs usually represent node pairs with similar semantics or close associations in the graph. For example, in a social network, friend node pairs with frequent interactions can be regarded as positive samples; negative sample node pairs are node pairs with weaker semantics or associations, such as two randomly selected uncorrelated user node pairs.
[0069] Among them, contrastive learning is a self-supervised learning method. By comparing the similarities between positive samples and negative samples, it can learn the discriminative features between nodes and enhance the understanding of the intrinsic structure of the data. Optimization algorithms such as stochastic gradient descent and Adam can be used to iteratively optimize the contrastive loss function. In each iteration, the parameters of the model are adjusted according to the gradient information of the loss function so that the value of the loss function gradually decreases.
[0070] In detail, the gradient information corresponding to the minimum loss function is used to update the weight parameters of the preset self-supervised learning model through the back-propagation algorithm. After multiple iterative optimizations, the model gradually converges to obtain a graph analysis model that can accurately process graph structure data. This model has the ability to extract useful information from the graph and perform reasoning analysis, such as classifying nodes and predicting the link relationship between nodes.
[0071] Among them, the data to be analyzed can be remapped into a trained graph analysis model, and a computable reasoning graph can be constructed using the node embedding and graph structure information learned by the model. The computable reasoning graph not only contains the structural information of the data, but also incorporates the semantic and pattern information learned by the model. Compared with the association graph between the original data, the computable reasoning graph has stronger expression and computing capabilities.
[0072] In detail, the corresponding graph algorithms and reasoning rules are applied on the computable reasoning graph according to the specific data analysis tasks. For example, in knowledge graph reasoning, reasoning can be performed based on the relationships between entities to predict new entity relationships.
[0073] In an embodiment of the present invention, by constructing a data association graph, the entities and their relationships in the data to be analyzed are presented in an intuitive graphical manner, thereby improving data visualization. Graph structure self-supervised learning can deeply explore the potential features in the graph, learn the deep semantic information of entities and relationships, and more accurately capture the inherent laws between data, reduce analysis bias caused by ignoring data associations, and significantly improve the accuracy of data analysis.
[0074] S6. Optimize the real-time aggregate analysis data according to the data analysis target to obtain second target analysis data.
[0075] In an embodiment of the present invention, data optimization rules are formulated based on key optimization indicators and constraints in the data analysis target, and data selection, data cleaning, data conversion and other processing are performed on the real-time aggregate analysis data according to the data optimization rules to obtain second target analysis data.
[0076] In the embodiment of the present invention, the step of optimizing the real-time aggregate analysis data according to the data analysis target to obtain second target analysis data includes: Perform parameter analysis on the data analysis objective to obtain key optimization indicators; Use the key optimization indicators and preset constraint conditions as data optimization conditions, and generate corresponding data optimization rules according to the data optimization conditions; Perform data adjustment on the real-time aggregated analysis data according to the data optimization rules to obtain second target analysis data.
[0077] In an embodiment of the present invention, the parameter analysis is to convert the data analysis objective into quantifiable and operable indicators. The key optimization indicators are the core criteria for measuring whether the data analysis result reaches the target, and the constraint conditions are the restrictive conditions that must be followed during the data optimization process, such as data privacy protection requirements, reliability requirements of data sources, timeliness requirements of data processing, etc.
[0078] Specifically, combine the key optimization indicators and preset constraint conditions to form data optimization conditions, which jointly determine the direction and scope of data optimization. Based on the key optimization indicators and constraint conditions, determine the overall framework of data optimization, and thus select the required data optimization rules from the preset data optimization rule set. The data optimization rule set includes optimization rules such as data selection rules, data cleaning rules, data conversion rules, and data integration rules.
[0079] Among them, the data selection rule stipulates from which data sources to select data, as well as the criteria and conditions for selecting data; the data cleaning rule clarifies how to handle problems such as missing values, outliers, and duplicate values in the data; the data conversion rule determines how to perform operations such as encoding, normalization, or standardization on the data; the data integration rule stipulates how to merge and fuse data from different sources and different formats.
[0080] The present invention can ensure the orderliness and standardization of the data optimization process, and improve the efficiency and quality of data optimization by formulating data optimization rules.
[0081] In an embodiment of the present invention, data that meets the requirements can be screened out from the real-time aggregated analysis data according to the data selection rule, and the screened data can be processed according to the data cleaning rule. For missing values, interpolation methods, mean filling methods, or model prediction-based methods can be used for filling; for outliers, methods such as box plot methods can be used for identification and processing; for duplicate values, duplicate removal operations are required.
[0082] Specifically, perform encoding, normalization, or standardization processing on the cleaned data according to the data conversion rule. Finally, merge and fuse the processed data according to the data integration rule. Through the above data adjustment operations, second target analysis data is obtained, improving data quality and data availability, and at the same time significantly enhancing the accuracy of data analysis.
[0083] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0084] As Figure 5 shown, it is a functional module diagram of a data analysis system based on artificial intelligence provided by an embodiment of the present invention.
[0085] In the embodiments of the present disclosure, a data analysis system based on artificial intelligence is provided. This data analysis system based on artificial intelligence corresponds one-to-one with the data analysis method based on artificial intelligence in the above embodiments. As Figure 5 shown, the data analysis system 100 based on artificial intelligence includes a rule engine construction module 101, an anomaly detection processing module 102, a multi-modal attention calculation module 103, an incremental aggregation module 104, a self-supervised learning module 105, and a data optimization module 106. The detailed description of each functional module is as follows: The rule engine construction module 101 is used to obtain historical analysis data and construct a sensitive rule recognition engine according to the historical analysis data; The anomaly detection processing module 102 is used to obtain data to be analyzed and perform anomaly detection processing on the data to be analyzed according to the sensitive rule recognition engine to obtain first target analysis data; The multi-modal attention calculation module 103 is used to perform multi-modal feature extraction on the first target analysis data to obtain a multi-modal feature set, and perform self-attention calculation on the multi-modal feature set to obtain the target feature of the first target analysis data; The incremental aggregation module 104 is used to perform incremental aggregation calculation on the first target analysis data according to the target feature to obtain real-time aggregation analysis data; The self-supervised learning module 105 is used to construct an inter-data association graph according to the data to be analyzed, and perform graph structure self-supervised learning on the data to be analyzed according to the inter-data association graph to obtain a data analysis target; The data optimization module 106 is used to optimize the real-time aggregation analysis data according to the data analysis target to obtain second target analysis data.
[0086] In one embodiment, when the rule engine construction module 101 executes constructing the sensitive rule recognition engine according to the historical analysis data, it is used to: Extract sensitive features from the historical analysis data to obtain sensitive information features; Perform rule screening on a preset set of sensitive information recognition rules according to the sensitive information features to obtain corresponding initial sensitive information recognition rules; Optimize the rule combination of the initial sensitive information recognition rule to obtain the target sensitive information recognition rule; Fill the preset rule engine template according to the target sensitive information recognition rule to obtain the sensitive rule recognition engine.
[0087] In one embodiment, when the rule engine construction module 101 executes the rule screening of the preset sensitive information recognition rule set according to the sensitive information feature to obtain the corresponding initial sensitive information recognition rule, it is used for: Extract the rule keywords from the sensitive information recognition rule set to obtain the rule keywords; Calculate the similarity between the rule keywords and the sensitive information feature to obtain the rule matching degree; Compare the rule matching degree with the preset matching threshold, determine the target rule keywords corresponding to the sensitive information feature, and select the sensitive information recognition rule corresponding to the target rule keywords as the initial sensitive information recognition rule.
[0088] In one embodiment, when the anomaly detection processing module 102 executes the anomaly detection processing of the data to be analyzed according to the sensitive rule recognition engine to obtain the first target analysis data, it is used for: Identify sensitive data from the data to be analyzed according to the sensitive rule recognition engine to obtain a sensitive data set; Perform verification processing on the sensitive data set to obtain a target sensitive data set; Calculate the allocation weight of each target sensitive data in the target sensitive data set; Calculate the allocation weight of each target sensitive data using the following formula: Where, Represents the th target sensitive data 's allocation weight, Represents the preset sensitive center weight of the target sensitive data set, Represents the transpose, Represents the th preset sensitive center offset of the target sensitive data, Represents the total number of target sensitive data; Perform regression calculation on the target sensitive data according to the allocation weight to obtain the sensitivity coefficient of the target sensitive data; The sensitivity coefficient of each target sensitive data can be calculated using the following formula: Where, Represents the sensitivity coefficient, Represents the preset intercept, Represents the a target sensitive data The assigned weight of represents a preset error factor; Take the target sensitive data with the sensitive coefficient less than the preset sensitive threshold as the first target analysis data.
[0089] In one embodiment, when the multi-modal attention calculation module 103 performs multi-modal feature extraction on the first target analysis data to obtain a multi-modal feature set, it is used for: Perform modal type recognition on the first target analysis data to obtain a modal classification result; If the modal classification result is image modal data, extract the position feature and grid feature of each pixel point in the image modal data, and perform weighted splicing on the position feature and grid feature to obtain the image feature of the image modal data, and use the image feature as the target feature; If the modal classification result is text modal data, calculate the word embedding feature sequence of the text data, and perform feature encoding on the word embedding feature sequence to obtain the text feature of the text modal data, and use the text feature as the target feature; Collect the target features to obtain a multi-modal feature set of the first target analysis data.
[0090] In one embodiment, when the multi-modal attention calculation module 103 performs self-attention calculation on the multi-modal feature set to obtain the target feature of the first target analysis data, it is used for: Perform scale adjustment on the multi-modal feature set to obtain a normalized feature set; Perform cross-modal feature fusion on the normalized feature set to obtain a modal fusion feature; Perform a convolution operation on the modal fusion feature to obtain an enhanced feature representation, and perform pooling dimensionality reduction processing on the enhanced feature representation to obtain a compressed feature representation; Perform activation function processing on the compressed feature representation to obtain an activation feature representation, and perform multi-layer feature fusion on the activation feature representation to obtain a fusion feature representation; Calculate the self-attention matrix of the fusion feature representation, and calculate the attention weighted features of different attention heads according to the self-attention matrix using a pre-constructed multi-head attention mechanism; Calculate the attention weighted features of different attention heads using the following formula: Where represents the attention weighted feature, represents the fusion feature representation, represents a preset scaling factor, represents the The query matrix of one attention head, indicating the key matrix of the th attention head in the attention matrix, indicating the value matrix of the th attention head in the attention matrix, indicating transpose, indicating a preset attention factor, indicating the activation function; Perform a fully connected dimensionality reduction process on the attention weighted features to obtain the target features of the first target analysis data.
[0091] In one embodiment, when the incremental aggregation module 104 performs incremental aggregation calculation on the first target analysis data according to the target features to obtain real-time aggregation analysis data, it is used for: Obtain an incremental data packet, perform data parsing on the incremental data packet to obtain incremental data fields; Perform field mapping on the incremental data fields and the target features to obtain an incremental feature mapping table; Perform feature matching on the first target analysis data according to the incremental feature mapping table to obtain valid incremental data that meets the preset feature threshold; Perform time series analysis on each data stream in the valid incremental data to obtain a timestamp sequence; Perform data interpolation processing on the timestamp sequence to obtain a target incremental data set, and segment the target incremental data set to obtain a time window data set; Perform aggregation calculation on the time window data set to obtain real-time aggregation analysis data; Perform aggregation calculation on the time window data set using the following formula: Where, represents the real-time aggregation analysis data, represents a preset window, represents the total number of the time window data sets, represents the th time window data set within the window.
[0092] In one embodiment, when the self-supervised learning module 105 constructs an association graph between data according to the data to be analyzed, it is used for: Perform attribute parsing on the data to be analyzed to obtain data entity attributes and data interaction relationships, and generate a structured data set according to the data entity attributes and the data interaction relationships; Extract the associated features in the structured dataset, and generate a data entity set and an association rule set for the data to be analyzed according to the associated features; Construct a data association graph with each entity in the data entity set as a node and each association rule in the association rule set as an edge weight.
[0093] In one embodiment, when the self-supervised learning module 105 performs graph structure self-supervised learning on the data to be analyzed according to the data association graph to obtain a data analysis target, it is used for: Perform a graph traversal on the data association graph to extract all graph node pairs in the data association graph; Perform node pair screening on the graph node pairs to obtain a first node pair and a second node pair, and use the first node pair and the second node pair as a positive sample node pair and a negative sample node pair respectively; Generate a contrast loss function according to the positive sample node pairs, the negative sample node pairs, and a preset contrast factor; Use the following formula to generate a contrast loss function according to the positive sample node pairs, the negative sample node pairs, and a preset contrast factor: where represents the th positive sample node pair, represents the th positive sample node pair, represents the th negative sample node pair, represents the preset contrast factor, represents the total number of negative sample node pairs, represents the contrast loss function; Minimize the contrast loss function to obtain a minimum loss function; Use the following formula to minimize the contrast loss function to obtain a minimum loss function: where represents the th positive sample node pair, represents the th positive sample node pair, represents the th negative sample node pair, represents the preset contrast factor, represents the total number of negative sample node pairs, represents the minimum loss function; Optimize a preset self-supervised learning model according to the minimum loss function to obtain a graph analysis model; Construct a computable inference graph for the data to be analyzed according to the graph analysis model; Perform graph reasoning on the computable inference graph to obtain a data analysis objective.
[0094] In one embodiment, when the data optimization module 106 performs data optimization on the real-time aggregated analysis data according to the data analysis objective to obtain second target analysis data, it is used for: Perform parameter parsing on the data analysis objective to obtain key optimization indicators; Use the key optimization indicators and preset constraint conditions as data optimization conditions, and generate corresponding data optimization rules according to the data optimization conditions; Perform data adjustment on the real-time aggregated analysis data according to the data optimization rules to obtain second target analysis data.
[0095] In the present invention, specific limitations on an artificial intelligence-based data analysis system can refer to the limitations on an artificial intelligence-based data analysis method in the above text, which will not be elaborated here. Each module in the above artificial intelligence-based data analysis system can be implemented in whole or in part by software, hardware, and their combinations. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0096] In the embodiments provided by the present invention, it should be understood that the disclosed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.
[0097] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0098] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claimed rights.
[0099] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-mentioned exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.
[0100] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0101] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.
[0102] In the embodiments provided in the present disclosure, it should be understood that the disclosed systems and methods can also be implemented in other ways. The system embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of systems, methods, and computer program products according to multiple embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0103] It should be noted that in the present disclosure, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements not only includes those elements but also includes other elements that are not explicitly listed, or further includes elements that are inherent to such process, method, article, or device. Without further limitation, the elements defined by the statement "including one..." do not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.
[0104] The above-described embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.
Claims
1. A data analysis method based on artificial intelligence, characterized in that: The method comprises: Acquire historical analysis data, and build a sensitive rule recognition engine based on the historical analysis data; Acquire the data to be analyzed, and perform anomaly detection processing on the data to be analyzed according to the sensitive rule recognition engine to obtain first target analysis data; Performing multimodal feature extraction on the first target analysis data to obtain a multimodal feature set, and performing self-attention calculation on the multimodal feature set to obtain target features of the first target analysis data; Performing incremental aggregation calculation on the first target analysis data according to the target characteristics to obtain real-time aggregate analysis data; Constructing a data association graph based on the data to be analyzed, and performing graph structure self-supervised learning on the data to be analyzed based on the data association graph to obtain a data analysis target; The real-time aggregate analysis data is optimized according to the data analysis target to obtain second target analysis data.
2. The artificial intelligence-based data analysis method according to claim 1, characterized in that: The performing incremental aggregation calculation on the first target analysis data according to the target feature to obtain real-time aggregate analysis data includes: Obtaining an incremental data packet, performing data parsing on the incremental data packet, and obtaining an incremental data field; Performing field mapping between the incremental data field and the target feature to obtain an incremental feature mapping table; Performing feature matching on the first target analysis data according to the incremental feature mapping table to obtain valid incremental data that meets a preset feature threshold; Performing time series analysis on each data stream in the valid incremental data to obtain a timestamp sequence; Performing data interpolation processing on the timestamp sequence to obtain a target incremental data set, and segmenting the target incremental data set to obtain a time window data set; Performing aggregation calculation on the time window data set to obtain real-time aggregate analysis data; The following formula is used to perform aggregation calculation on the time window data set: in, represents the real-time aggregated analysis data, Indicates the default window. represents the total number of data sets in the time window, Indicates A time window dataset within a window.
3. The artificial intelligence-based data analysis method according to claim 1, characterized in that: The step of performing graph structure self-supervised learning on the data to be analyzed according to the data association graph to obtain a data analysis target includes: Performing graph traversal on the data association graph to extract all graph node pairs in the data association graph; Performing node pair screening on the graph node pairs to obtain a first node pair and a second node pair, and using the first node pair and the second node pair as a positive sample node pair and a negative sample node pair, respectively; Generate a contrast loss function according to the positive sample node pair, the negative sample node pair and a preset contrast factor; The contrast loss function is generated according to the positive sample node pair, the negative sample node pair and the preset contrast factor using the following formula: in, Indicates Positive sample node pairs, Indicates Positive sample node pairs, Indicates negative sample node pairs, Indicates the preset contrast factor, represents the total number of negative sample node pairs, represents the contrast loss function; Minimize the contrast loss function to obtain a minimum loss function; The contrast loss function is minimized using the following formula to obtain the minimum loss function: in, Indicates Positive sample node pairs, Indicates Positive sample node pairs, Indicates negative sample node pairs, Indicates the preset contrast factor, represents the total number of negative sample node pairs, represents the minimum loss function; Optimizing the preset self-supervised learning model according to the minimum loss function to obtain a graph analysis model; Constructing a computable reasoning graph of the data to be analyzed according to the graph analysis model; Graph reasoning is performed on the computable reasoning graph to obtain a data analysis target.
4. The artificial intelligence-based data analysis method according to claim 1, characterized in that: The performing self-attention calculation on the multimodal feature set to obtain the target feature of the first target analysis data includes: Performing scale adjustment on the multimodal feature set to obtain a normalized feature set; Performing cross-modal feature fusion on the normalized feature set to obtain modal fusion features; Performing a convolution operation on the modal fusion feature to obtain an enhanced feature representation, and performing a pooling dimensionality reduction process on the enhanced feature representation to obtain a compressed feature representation; Performing activation function processing on the compressed feature representation to obtain an activated feature representation, and performing multi-layer feature fusion on the activated feature representation to obtain a fused feature representation; Calculating a self-attention matrix represented by the fused features, and calculating attention weighted features of different attention heads using a pre-built multi-head attention mechanism according to the self-attention matrix; The attention weighted features of different attention heads are calculated using the following formula: in, represents the attention weighted feature, represents the fusion feature representation, Indicates the preset scaling factor, represents the first The query matrix of the attention head, Indicates the first The key matrix of the attention head, Indicates the first The value matrix of the attention head, represents transpose, represents the preset attention factor, express Activation function; Perform full-connection dimensionality reduction processing on the attention weighted features to obtain target features of the first target analysis data.
5. The artificial intelligence-based data analysis method according to claim 1, characterized in that: The performing anomaly detection processing on the data to be analyzed according to the sensitive rule recognition engine to obtain first target analysis data includes: Performing sensitive data identification on the data to be analyzed according to the sensitive rule identification engine to obtain a sensitive data set; Performing verification processing on the sensitive data set to obtain a target sensitive data set; Calculating the allocation weight of each target sensitive data in the target sensitive data set; The allocation weight of each target sensitive data is calculated using the following formula: in, Indicates Target sensitive data The distribution weight of Represents the preset sensitive center weight of the target sensitive data set, represents transpose, Indicates The preset sensitive center bias of target sensitive data, Indicates the total number of target sensitive data; Performing regression calculation on the target sensitive data according to the assigned weights to obtain a sensitivity coefficient of the target sensitive data; The sensitivity coefficient of each target sensitive data can be calculated using the following formula: in, represents the sensitivity coefficient, represents the preset intercept, Indicates Target sensitive data The distribution weight of Indicates the preset error factor; The target sensitive data whose sensitivity coefficient is less than a preset sensitivity threshold is used as the first target analysis data.
6. The artificial intelligence-based data analysis method according to claim 1, characterized in that: The step of constructing a sensitive rule recognition engine according to the historical analysis data includes: Extracting sensitive features from the historical analysis data to obtain sensitive information features; Filtering a preset sensitive information identification rule set according to the sensitive information characteristics to obtain a corresponding initial sensitive information identification rule; Performing rule combination optimization on the initial sensitive information identification rule to obtain a target sensitive information identification rule; The preset rule engine template is filled in according to the target sensitive information identification rule to obtain a sensitive rule identification engine.
7. The artificial intelligence-based data analysis method according to claim 1, characterized in that: The extracting multimodal features from the first target analysis data to obtain a multimodal feature set includes: Performing modal type identification on the first target analysis data to obtain a modal classification result; If the modality classification result is image modality data, extract the position feature and grid feature of each pixel in the image modality data, and perform weighted splicing on the position feature and grid feature to obtain the image feature of the image modality data, and use the image feature as the target feature; If the modality classification result is text modality data, calculating a word embedding feature sequence of the text data, and performing feature encoding on the word embedding feature sequence to obtain text features of the text modality data, and using the text features as target features; The target features are aggregated to obtain a multimodal feature set of the first target analysis data.
8. The artificial intelligence-based data analysis method according to claim 1, characterized in that: The step of constructing a data association graph based on the data to be analyzed includes: Performing attribute analysis on the data to be analyzed to obtain data entity attributes and data interaction relationships, and generating a structured data set according to the data entity attributes and the data interaction relationships; Extracting correlation features from the structured data set, and generating a data entity set and an association rule set of the data to be analyzed according to the correlation features; A data association graph is constructed with each entity in the data entity set as a node and each association rule in the association rule set as an edge weight.
9. The artificial intelligence-based data analysis method according to claim 1, characterized in that: The step of optimizing the real-time aggregate analysis data according to the data analysis target to obtain second target analysis data includes: Perform parameter analysis on the data analysis target to obtain key optimization indicators; The key optimization indicators and the preset constraint conditions are used as data optimization conditions, and corresponding data optimization rules are generated according to the data optimization conditions; The real-time aggregate analysis data is adjusted according to the data optimization rule to obtain second target analysis data.
10. A data analysis system based on artificial intelligence, characterized in that: The system comprises: A rule engine construction module, used to obtain historical analysis data and construct a sensitive rule recognition engine based on the historical analysis data; An anomaly detection processing module is used to obtain the data to be analyzed, and perform anomaly detection processing on the data to be analyzed according to the sensitive rule recognition engine to obtain first target analysis data; a multimodal attention calculation module, configured to perform multimodal feature extraction on the first target analysis data to obtain a multimodal feature set, and perform self-attention calculation on the multimodal feature set to obtain target features of the first target analysis data; An incremental aggregation module, configured to perform incremental aggregation calculation on the first target analysis data according to the target characteristics to obtain real-time aggregate analysis data; A self-supervised learning module is used to construct a data association map based on the data to be analyzed, and perform graph structure self-supervised learning on the data to be analyzed based on the data association map to obtain a data analysis target; The data optimization module is used to optimize the real-time aggregate analysis data according to the data analysis target to obtain second target analysis data.
Citation Information
Patent Citations
Time sequence characteristic analysis method and system based on multi-dimensional data
CN119202656A
Ship storage demand prediction method and system based on machine learning
CN119886733A
Cited By
Server data processing method and system based on deep learning
CN120386637A
A server data processing method and system based on deep learning
CN120386637B
Intelligent bid invitation risk control early warning method
CN121120222A
An intelligent bidding risk control early warning method
CN121120222B