An artificial intelligence-based big data analysis system and method

By extracting distributed query information from the big data platform under a distributed architecture and performing cross-validation of query gradient constraints and static hierarchical attributes, the problems of task scheduling imbalance and state lag in distributed queries are solved, and efficient distributed query adaptability and accuracy are achieved.

CN120596551BActive Publication Date: 2025-10-10GUIZHOU VOCATIONAL & TECH COLLEGE OF ECONOMICS & TRADE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511097009.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-10-10
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

Existing AI-based big data analysis methods cannot effectively perceive the dynamic evolution characteristics of data among distributed nodes under a distributed architecture, resulting in unbalanced task scheduling, delayed status judgment, and inability to adapt to node load fluctuations in a timely manner during the query process, affecting the reliability of the query results.

Method used

The distributed query information of the big data platform is read through a distributed node cluster, the query gradient constraints are extracted, the parallel matching of the query evolution sequence is determined, cross-validation is performed in combination with static hierarchical attributes, and distributed queries are performed using dynamic parallel labels and cross-trusted trajectories to achieve global perception verification.

Benefits of technology

It improves the adaptability of distributed queries, solves the problems of query status lag and node load imbalance, and improves the accuracy and execution efficiency of query scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596551B_ABST
    Figure CN120596551B_ABST
Patent Text Reader

Abstract

The application provides a big data analysis system and method based on artificial intelligence, relates to the technical field of data analysis, and reads distributed query information of to-be-processed data in a big data platform through a distributed node cluster and extracts query gradient constraints on all nodes in the distributed node cluster; performs parallel matching on a query evolution sequence to obtain a dynamic parallel label corresponding to a query search state; cross-verification is performed on the query gradient constraints on all nodes in the distributed node cluster and static hierarchical attributes corresponding to to-be-queried data items to obtain a cross-trust trajectory corresponding to a query fragmentation state; and distributed query is performed on the to-be-queried data items in the big data platform according to the dynamic parallel label corresponding to the query search state and the cross-trust trajectory corresponding to the query fragmentation state. The application can perform global perception verification on to-be-processed data under a distributed architecture to improve the adaptability of distributed query in big data analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data analysis technology, and more specifically, to an artificial intelligence-based big data analysis system and method. Background Art

[0002] Data analysis is the systematic processing of large-scale, multi-dimensional, and heterogeneous data acquired from various information systems to uncover implicit structural features, association patterns, and evolutionary trends within the data, thereby providing a basis for decision support, business optimization, and intelligent response. In the big data environment, data analysis is no longer limited to static statistics. Instead, it requires the integration of real-time processing capabilities, parallel computing frameworks, and intelligent algorithm models to handle the nonlinear patterns and complex relationships present in massive amounts of data.

[0003] However, existing AI-based big data analysis methods generally rely on centralized feature extraction and static model reasoning. This makes it difficult to effectively perceive the dynamic evolution of data across distributed nodes when querying large-scale, heterogeneous data in real-time scenarios. This leads to unbalanced task scheduling, delayed state judgment, and inability to adapt to node load fluctuations, thus affecting the reliability of query results. Therefore, how to globally perceive and verify the data to be processed in a distributed architecture to improve the adaptability of distributed queries in big data analysis is a challenge facing the industry. Summary of the Invention

[0004] The present application provides an artificial intelligence-based big data analysis system and method, which can perform global perception verification on the data to be processed in a distributed architecture to improve the adaptability of distributed queries in big data analysis.

[0005] In a first aspect, the present application provides a big data analysis method based on artificial intelligence, the analysis method comprising the following steps:

[0006] Read the distributed query information of the data to be processed in the big data platform through the distributed node cluster, and extract the query gradient constraints on all nodes in the distributed node cluster;

[0007] Determining a query evolution sequence in a data query behavior according to the distributed query information, performing parallel matching on the query evolution sequence, and obtaining a dynamic parallel tag corresponding to a query retrieval state;

[0008] Extract the static hierarchical attributes corresponding to the query data item in the distributed architecture, perform cross-validation based on the query gradient constraints on all nodes in the distributed node cluster and the static hierarchical attributes corresponding to the query data item, and obtain the cross-trusted trajectory corresponding to the query shard status;

[0009] A distributed query is performed on the data items to be queried in the big data platform based on the dynamic parallel tags corresponding to the query retrieval status and the cross-trusted tracks corresponding to the query sharding status.

[0010] In this embodiment, the distributed query information refers to a collection of statements, paths, and resources generated by each distributed node during the query process.

[0011] In this embodiment, extracting the query gradient constraints on all nodes in the distributed node cluster specifically includes:

[0012] Determining local gradient features on all nodes according to the distributed node cluster;

[0013] Determine query constraints during data query based on local gradient features on all nodes;

[0014] The query gradient constraints on all nodes in the distributed node cluster are determined by the query constraint conditions.

[0015] In this embodiment, determining the query evolution sequence in the data query behavior according to the distributed query information specifically includes:

[0016] Extracting query operation logs on each node according to the distributed query information;

[0017] Construct the corresponding query state transition record based on the query operation log on each node;

[0018] Generate the query evolution sequence in the data query behavior based on all query state transition records.

[0019] In this embodiment, the query retrieval status refers to the specific running status of the query task in different stages and links during the data query process.

[0020] In this embodiment, extracting static hierarchical attributes corresponding to the data item to be queried under a distributed architecture specifically includes:

[0021] Building hierarchical attribute indexes in distributed architectures;

[0022] Parsing associated feature tags of the data item to be queried based on the hierarchical attribute index;

[0023] The static hierarchical attribute corresponding to the data item to be queried is determined according to the associated feature tag.

[0024] In this embodiment, the query shard status refers to the comprehensive performance between the execution status of each data query task on different shards and the current operation characteristics during the distributed query process.

[0025] In this embodiment, performing a distributed query on the data items to be queried in the big data platform based on the dynamic parallel tags corresponding to the query retrieval status and the cross-trusted traces corresponding to the query sharding status specifically includes:

[0026] determining a dynamic execution period of a distributed query according to the dynamic parallel tag and the cross trusted trace;

[0027] By matching the dynamic execution cycle with the query load threshold of the big data platform, a candidate execution node set for distributed query is screened;

[0028] A distributed query is executed on the data item to be queried according to the candidate execution node set.

[0029] In this embodiment, the distributed query refers to a query method in which a query request is divided into multiple subtasks, which are executed in parallel on a distributed set of nodes, and the query results are finally aggregated and returned.

[0030] In a second aspect, the present application provides an artificial intelligence-based big data analysis system for executing an artificial intelligence-based big data analysis method, the analysis system comprising:

[0031] The data reading module is used to read the distributed query information of the data to be processed in the big data platform through the distributed node cluster and extract the query gradient constraints on all nodes in the distributed node cluster;

[0032] A parallel matching module is used to determine the query evolution sequence in the data query behavior according to the distributed query information, perform parallel matching on the query evolution sequence, and obtain a dynamic parallel tag corresponding to the query retrieval state;

[0033] The cross-validation module is used to extract the static hierarchical attributes corresponding to the query data item in a distributed architecture, and perform cross-validation based on the query gradient constraints on all nodes in the distributed node cluster and the static hierarchical attributes corresponding to the query data item to obtain the cross-trusted trajectory corresponding to the query shard status;

[0034] The distributed query module is used to perform distributed queries on the data items to be queried in the big data platform based on the dynamic parallel tags corresponding to the query retrieval status and the cross-trusted tracks corresponding to the query sharding status.

[0035] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects:

[0036] The distributed query information of to-be-processed data in the big data platform is read through the distributed node cluster, and query gradient constraints on all nodes in the distributed node cluster are extracted; a query evolution sequence in a data query behavior is determined according to the distributed query information, parallel matching is performed on the query evolution sequence, and a dynamic parallel label corresponding to a query retrieval state is obtained; a static hierarchical attribute corresponding to a to-be-queried data item is extracted under a distributed architecture, cross verification is performed on the query gradient constraints on all nodes in the distributed node cluster and the static hierarchical attribute corresponding to the to-be-queried data item, and a cross-trust trajectory corresponding to a query sharding state is obtained; and distributed query is performed on the to-be-queried data item in the big data platform according to the dynamic parallel label corresponding to the query retrieval state and the cross-trust trajectory corresponding to the query sharding state.

[0037] It can be seen that in the present application, the accuracy of query scheduling can be improved in the case of query state lag and node load imbalance in distributed big data query; wherein, by extracting the distributed query information of to-be-processed data in the distributed node cluster and extracting the query gradient constraints on all nodes, the dynamic distribution state of the query features of each node in the big data platform is accurately identified and uniformly modeled; by determining the query evolution sequence according to the distributed query information and performing parallel matching, the evolution trend identification and retrieval state labeling of the multi-node query behavior path can be realized, so as to construct a dynamic query label system that is responsive and adaptive in scheduling in a multi-node high-concurrency environment, effectively solving the problems of state identification lag and scheduling inconsistency; by extracting the static hierarchical attribute of the to-be-queried data item and cross verifying in combination with the query gradient constraints, the trust evaluation capability of the query sharding state in the distributed environment can be enhanced, and the consistency and traceability of the sharding query result can be improved through cross matching of the feedback label and the sharding rule; by performing distributed query according to the dynamic parallel label and the cross-trust trajectory, the linkage scheduling of the query behavior features and the data sharding state is realized, the execution node and the query load are accurately matched, and the execution efficiency of the whole-process distributed query in big data analysis is improved.

[0038] In summary, the technical solution adopted by the present application can perform global perception verification on to-be-processed data under a distributed architecture to improve the adaptability of distributed query in big data analysis. BRIEF DESCRIPTION OF DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only the embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0040] Figure 1This is an exemplary flow chart of an artificial intelligence-based big data analysis method provided in this application;

[0041] Figure 2 is a schematic diagram of a process for determining dynamic parallel tags provided by this application;

[0042] Figure 3 is a schematic diagram of a process for determining cross-trusted trajectories provided in this application;

[0043] Figure 4 This is a module structure diagram of an artificial intelligence-based big data analysis system provided in this application. DETAILED DESCRIPTION

[0044] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0045] The embodiment of the present application provides an artificial intelligence-based big data analysis system and method, the core of which is to read the distributed query information of the data to be processed in the big data platform through a distributed node cluster, and extract the query gradient constraints on all nodes in the distributed node cluster; determine the query evolution sequence in the data query behavior according to the distributed query information, perform parallel matching on the query evolution sequence, and obtain the dynamic parallel label corresponding to the query retrieval state; extract the static hierarchical attributes corresponding to the data items to be queried under the distributed architecture, cross-validate according to the query gradient constraints on all nodes in the distributed node cluster and the static hierarchical attributes corresponding to the data items to be queried, and obtain the cross-trusted trajectory corresponding to the query sharding state; and perform distributed query on the data items to be queried in the big data platform based on the dynamic parallel label corresponding to the query retrieval state and the cross-trusted trajectory corresponding to the query sharding state.

[0046] Example 1: In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods. Figure 1 As shown in FIG, this figure is an exemplary flow chart of an artificial intelligence-based big data analysis method according to this embodiment of the present application, and the analysis method includes the following steps:

[0047] In step S1, the distributed query information of the data to be processed in the big data platform is read through the distributed node cluster, and the query gradient constraints on all nodes in the distributed node cluster are extracted.

[0048] In a specific implementation, the distributed node cluster can be used to read the distributed query information of the data to be processed in the big data platform in the following manner: a lightweight query log collection agent (such as Fluentd or Apache Flink) can be deployed on each node in the distributed node cluster to listen to the data query events generated in the database or query engine (such as Hive, Presto, and SparkSQL) in real time. Each agent packages the collected query statement, execution time, data scanning path, resource usage, and other information into a structured log, and transmits the log to a centralized query analysis engine through a message queue (such as Kafka) in an asynchronous manner. The analysis engine uses a rule template or a semantic parser to standardize the extraction of the log content, constructs a query meta-information set, and uses the query meta-information set as the distributed query information of the data to be processed in the big data platform. In other embodiments, other methods can also be used to read the distributed query information of the data to be processed in the big data platform, which will not be described here.

[0049] It should be noted that in this application, the distributed node cluster refers to a system cluster composed of multiple independent but interconnected and cooperative computing nodes, which is responsible for data storage, query processing, and computing tasks to achieve horizontal expansion of resources and high availability; the distributed query information refers to the collection of statements, paths, and resources generated by each distributed node during the query process.

[0050] In this embodiment, the query gradient constraint on all nodes in the distributed node cluster can be achieved by the following steps:

[0051] According to the distributed node cluster, the local gradient features on all nodes are determined.

[0052] According to the local gradient features on all nodes, the query constraint condition during data query is determined.

[0053] The query gradient constraint on all nodes in the distributed node cluster is determined through the query constraint condition.

[0054] In specific implementation, first, deploy a performance collection agent on each node, such as using Prometheus NodeExporter or a custom probe, to regularly collect key indicators during query operations, including CPU usage, memory usage, disk I / O rate, query latency, cache hit rate, etc., and combine the operation logs of the query engine (such as Presto, Apache Hive, or Apache Spark) to extract the execution statement of each query, the data tables involved, the filter conditions, grouping operations, data volume, and execution time. Use a window sliding analysis strategy to model the operation logs, and use gradient calculation methods (such as first-order derivative approximation or linear fitting slope) to obtain the changing trends of each indicator in different time periods, and finally generate local gradient features on all nodes. Then, a K-means or DBSCAN clustering algorithm is used to cluster local gradient features, identifying three common node status categories: high load, medium load, and low load. Based on the typical query types and resource usage of each node category over the past period, a node-query constraint template mapping is established. For example, if a node is in a medium load state, the constraint is "a single query cannot scan more than 50MB of partitioned data, and the number of concurrent queries cannot exceed 3." Based on the constraint template, a rule inference engine (such as Drools) is invoked to automatically generate currently acceptable query constraints for each node based on the current state of the distributed node cluster. Finally, a lightweight prediction module (which can be based on ARIMA time series or LSTM deep neural networks) is deployed on each node to predict future short-term resource fluctuations. Based on the current query constraints and resource forecast curves, an acceptable query intensity range for each node is integrated and used as the query gradient constraint. For example, if node A currently limits the maximum query data volume to 100MB, but predicts that I / O will decrease within 10 minutes, the query gradient constraint is [80MB, 120MB].

[0055] It should be noted that in this application, local gradient characteristics refer to the trajectory of query performance indicators of each node changing with resource usage, data access pattern and query complexity during the query execution process within a period of time; query constraints refer to the execution restrictions on query scheduling, concurrency, data shard size, and response time limit derived from the query performance status of the node; query gradient constraints refer to the dynamic restrictions imposed by each node on the query operation in a specific time window.

[0056] In step S2, a query evolution sequence in the data query behavior is determined according to the distributed query information, and the query evolution sequence is matched in parallel to obtain a dynamic parallel tag corresponding to the query retrieval state.

[0057] In this embodiment, determining the query evolution sequence in the data query behavior according to the distributed query information can be achieved by using the following steps:

[0058] Extracting query operation logs on each node according to the distributed query information;

[0059] Construct the corresponding query state transition record based on the query operation log on each node;

[0060] Generate the query evolution sequence in the data query behavior based on all query state transition records.

[0061] To implement this, first deploy a unified log collection process on each distributed node, using, for example, Apache Flume or a custom collection program. Connect to the query engine log output port (e.g., Spark History Server or Presto Query Logger) to output query operation logs through the query engine log output port. Collected query operation logs are uniformly structured, with fields including query unique identifier, user identifier, query start time, end time, execution stage identifier, accessed tables and columns, filter conditions, aggregation method, and query status. Next, the query lifecycle is divided into several typical state nodes, such as "query plan generation," "data preprocessing," "execution scheduling," "result merging," and "abnormal interruption." Each state node represents a stage of query behavior. A log parsing program (using Python or Java) is then written to map fields such as timestamps, operation stages, and execution results from the query operation logs to typical state nodes, constructing a chronological sequence of state nodes. Weighted edges or labeled transition relationships are then established between adjacent state nodes, describing the conditions (e.g., success or timeout) and costs (e.g., time consumption or resource consumption) for state transitions, ultimately forming a complete record of query state transitions. Finally, query state transition records generated on all nodes are aggregated into a centralized analysis engine (e.g., using Apache Kafka for transfer and then processed by Flink or Spark Streaming) and categorized and aggregated by query identifier. Each query is then sorted chronologically to form a linear evolution path. The state types in the linear evolution path are encoded as state codes (e.g., PLAN → PREPROCESS → EXECUTE → MERGE → SUCCESS). This encoded linear evolution path serves as the query evolution sequence in the data query behavior.

[0062] It should be noted that in the present application, the data query behavior refers to the whole process operation and state change of initiating, executing and completing the query task in the big data platform; the query operation log refers to the detailed event track recorded by the node in the query process; the query state transition record refers to the data structure of the state change of the query behavior with the time advancing; and the query evolution sequence refers to the time sequence sequence describing the dynamic evolution path of the query behavior in the query life cycle.

[0063] Preferably, in the present embodiment, the query evolution sequence is matched in parallel to obtain the dynamic parallel label corresponding to the query retrieval state, and the dynamic parallel label is determined according to the query evolution sequence, the query operation log and the query state transition record. Figure 2 As shown in the figure, the figure is a flowchart for determining the dynamic parallel label in some embodiments of the present application, and the dynamic parallel label in the present embodiment can be realized by the following steps:

[0064] In step S21, the query sub-sequence fragments in the data parallel processing are determined according to the query evolution sequence;

[0065] In step S22, the state transition rules matched with the query sub-sequence fragments are generated based on the preset query mode library;

[0066] In step S23, the query sub-sequence fragments are matched with the state transition rules to obtain the query matching trend corresponding to the query retrieval state;

[0067] In step S24, the dynamic parallel label corresponding to the query retrieval state is determined through the query matching trend.

[0068] In specific implementation, first, design a sharding strategy based on time windows or state transition points. For example, a fixed-length sliding window can be set, or key states (such as "execution start", "intermediate merge", and "execution end") can be used as boundaries to dynamically intercept query subsequences during data parallel processing. Programming languages ​​(such as Python and Java) can be used to implement the sharding algorithm, read the query evolution sequence, and automatically generate several consecutive query subsequence shards according to the interception rules. Next, based on historical query evolution sequence data and combined with expert knowledge, machine learning or rule mining technology is used to generate typical state transition patterns, store them in the database, and obtain a query pattern library. The pattern corresponding to the current analysis task is read from the query pattern library to automatically generate state transition rules. Each state transition rule includes fields such as the starting state, target state, transition condition, and weight. Then, a sequence matching algorithm based on a finite state machine (FSM) or dynamic programming is used to sequentially match the states in the query subsequence shards against the state transition rules. A matching score is calculated, and the state nodes of each query subsequence shard are traversed to verify whether they meet the transition conditions in the state transition rules. State transition paths that meet the conditions are assigned high weights, while non-matching state transition paths are assigned low weights or eliminated. Based on the matching results, the most frequent and weighted state transition paths are counted to determine the query matching trend for the query subsequence shards, such as dynamic states like "continuously executing," "about to merge," and "abnormally interrupted." Finally, mapping rules are developed to map different query matching trends to specific dynamic parallel labels, such as "high parallelism," "resource constraints require reduced parallelism," and "waiting for I / O." A dynamic label generation program receives query matching trend input in real time and outputs dynamic parallel labels corresponding to the query retrieval states based on the mapping rules.

[0069] It should be noted that, in this application, the query retrieval state refers to the specific operating state of the query task in different stages and links during the data query process; the query pattern library refers to a pre-established database containing typical query state transition patterns and behavior rules; the query subsequence shard refers to a continuous state sequence segment intercepted from the complete query evolution sequence; the state transition rule represents the association rule of the transition conditions and logical relationships between query states; the query matching trend refers to the representation of the tendency of the query subsequence shard on the state transition path; the dynamic parallel label refers to the parallel processing feature used to mark the current query retrieval state.

[0070] In step S3, the static hierarchical attributes corresponding to the data item to be queried are extracted under the distributed architecture, and cross-validation is performed based on the query gradient constraints on all nodes in the distributed node cluster and the static hierarchical attributes corresponding to the data item to be queried to obtain the cross-trusted trajectory corresponding to the query shard status.

[0071] In the embodiment, the static hierarchical attribute corresponding to the data item to be queried is extracted under the distributed architecture by the following steps:

[0072] Building a hierarchical attribute index in the distributed architecture;

[0073] Resolving the associated feature label of the data item to be queried based on the hierarchical attribute index;

[0074] Determining the static hierarchical attribute corresponding to the data item to be queried according to the associated feature label.

[0075] In the specific implementation, first, a multi-dimensional hierarchical structure is designed according to the business requirements and data characteristics, for example, first layering by region, then layering by time, and finally layering by data category. Then, an index structure suitable for the distributed environment is selected, such as a distributed hash table (DHT), a B+ tree, or an R tree. In combination with a metadata management method, distributed storage and access of the attribute index are implemented. By traversing all the data items stored in the distributed nodes, the corresponding attribute information is extracted. According to the hierarchical design, the attributes are written into the index structure. A distributed coordination service (such as Zookeeper) is used to ensure the consistency and synchronization of the index construction, that is, the hierarchical attribute index in the distributed architecture is obtained. Then, a unified index query interface is designed to support the input of the unique identifier of the data item to be queried, to quickly access the corresponding attribute index structure, and to parse the attribute nodes corresponding to the data item to be queried through index path traversal and matching, to extract the associated feature labels such as category label, time label, and geographic label. A distributed computing framework (such as Apache Spark or Flink) is used to access the node indexes in parallel, to fuse the multi-dimensional and hierarchical extracted feature labels, and to generate the associated feature label of the data item to be queried. Finally, according to the business logic and data model, a mapping rule from the associated feature label to the static hierarchical attribute is formulated, and the rule of the corresponding specific static hierarchical attribute of the associated feature label is determined. Then, the associated feature label is input, and the mapping rule is used for judgment and classification, to output the final static hierarchical attribute.

[0076] It should be noted that, in the present application, the distributed architecture refers to a system structure composed of multiple distributed nodes, which is used to support data storage, calculation, and management, and to ensure high availability and high concurrency of data. The data item to be queried refers to a specific data object that needs to be retrieved, processed, and analyzed in the big data platform. The hierarchical attribute index refers to a structure for organizing and indexing attribute information of the data item in different dimensions or levels (such as category, time, and region). The associated feature label refers to a descriptive identifier of the static feature of the data item to be queried. The static hierarchical attribute refers to the attribute level information inherent to the data item to be queried and not frequently changing with time.

[0077] Preferably, in the present embodiment, the cross-verification is performed according to the query gradient constraint on all nodes in the distributed node cluster and the static hierarchical attribute corresponding to the data item to be queried, to obtain a cross-trust track corresponding to the query shard state, and the cross-trust track is referenced to determine the query shard state. Figure 3 As shown in the figure, the figure is a flowchart for determining the cross-trust track in some embodiments of the present application, and the cross-trust track in the present embodiment can be realized by the following steps:

[0078] In step S31, the shard balance constraint in the query shard state is determined according to the query gradient constraint on all nodes in the distributed node cluster.

[0079] In step S32, the shard feedback label in the query shard state is determined according to the static hierarchical attribute corresponding to the data item to be queried.

[0080] In step S33, the trusted query information when the query shard state is cross-matched is determined according to the shard balance constraint and the shard feedback label.

[0081] In step S34, the cross-trust track corresponding to the query shard state is determined according to the trusted query information.

[0082] In a specific implementation, first, the distributed monitoring system is used to collect the query load, response time and resource occupation of each node in real time, form a unified query gradient constraint data set, use statistical analysis or machine learning model (such as load prediction model, cluster analysis) to evaluate the node load distribution, and use the output result of the load distribution model as the sharding balance constraint rule, for example, limit the single node query task volume to be less than a threshold value, or ensure that the sharding tasks are evenly distributed among the nodes. The sharding balance constraint rule is specified by using a constraint expression language or a rule engine (such as Drools), and the specified result is used as the sharding balance constraint in the query sharding state. Next, the distributed index service is called to retrieve the static hierarchical attribute information of the data items to be queried, set the sharding feedback tag generation rule according to the business requirements, for example, assign different weights or priorities to different hierarchical attributes, and then use the rule engine to map and convert the static hierarchical attributes to generate corresponding sharding feedback tags. The sharding feedback tags can include data category priority, access frequency level, node adaptation degree and other information. Then, a multi-dimensional verification model containing the sharding balance constraint and the sharding feedback tag is designed, the dual constraints of load balancing and data attribute matching are fused, a weighted fusion method is used to integrate the parameters of the sharding balance constraint and the features of the sharding feedback index, a comprehensive score is formed, and according to the comprehensive score, the credibility of each query sharding state is calculated. The credibility is used as the credible query information when the query sharding states are cross-matched. Finally, a cross-credible trajectory data structure containing a timestamp, a query sharding state, a credibility value and the like is defined, the query sharding state trajectory is updated in real time based on the credible query information flow, a sliding window mechanism is used to record the state evolution, a time series analysis algorithm (such as hidden Markov model, Bayesian filtering) is integrated to perform trend prediction and anomaly detection on the updated query sharding state trajectory, and the results of the trend prediction and anomaly detection are used as the cross-credible trajectory corresponding to the query sharding state.

[0083] It should be noted that in this application, the query sharding state refers to the comprehensive performance between the execution of each data query task on different shards and the current running characteristics in the distributed query process; the sharding balance constraint refers to the balance rule that must be followed when assigning tasks to each query shard; the sharding feedback tag refers to an index that measures the query adaptation degree of the data in the query sharding state; the credible query information refers to the execution information of the query task after multi-dimensional verification and cross-matching; and the cross-credible trajectory refers to the state reflecting the evolution of the query sharding state over time.

[0084] In step S4, the cross-credible trajectory corresponding to the query sharding state and the dynamic parallel tag corresponding to the query retrieval state are used to perform distributed query on the data items to be queried in the big data platform.

[0085] In this embodiment, a distributed query of the data items to be queried in the big data platform based on the dynamic parallel tags corresponding to the query retrieval status and the cross-trusted traces corresponding to the query sharding status can be implemented by the following steps:

[0086] determining a dynamic execution period of a distributed query according to the dynamic parallel tag and the cross trusted trace;

[0087] By matching the dynamic execution cycle with the query load threshold of the big data platform, a candidate execution node set for distributed query is screened;

[0088] A distributed query is executed on the data item to be queried according to the candidate execution node set.

[0089] In its implementation, the dynamic parallelism labels corresponding to the query retrieval status and the cross-trusted trajectories corresponding to the query sharding status are first read. The dynamic parallelism labels and cross-trusted trajectories are jointly modeled along the time dimension to form a multi-dimensional query scheduling condition matrix. Based on the query concurrency requirements and the changing trends of the cross-trusted trajectories, a sliding time window method is used to calculate feasible query execution windows. The window length can be dynamically adjusted based on system resource fluctuations. A dynamic execution cycle data structure containing start time, end time, and execution priority is constructed, thus obtaining the dynamic execution cycle of the distributed query. This dynamic execution cycle is then linked to the scheduler to control the initiation and parallel triggering of specific query tasks. Then, through a resource management system (such as Kubernetes, YARN, or Mesos), metrics such as CPU utilization, memory utilization, and I / O load are obtained from each distributed node in real time. During the dynamic execution cycle, the load curve of each node is compared and analyzed to filter out nodes that can maintain a query load threshold within the dynamic execution cycle. Multi-dimensional filtering conditions, including node availability, historical query success rate, and network latency, are used to form the final set of candidate execution nodes. Finally, based on the static hierarchical attributes of the data items to be queried and their data locations, the query request is logically divided to generate corresponding sub-query tasks. A distributed scheduling engine (such as Apache Flink, Spark, Presto, or a self-developed scheduler) is used to evenly distribute the sub-tasks to each node in the candidate execution node set. The query sub-task is then started in each node, and its execution status is monitored in real time. The status and intermediate results are reported through the heartbeat mechanism, and the query results of each node are summarized. The final calculation and integration are performed in the master node or the intermediate result merging node to return a complete query response.

[0090] It should be noted that in the present application, the dynamic execution period refers to a time scheduling window for controlling query start, pause and redistribution; the query load threshold refers to the upper limit value of the node resource utilization and the task density preset in the big data platform; the candidate execution node set refers to the node set that meets the query load threshold limit condition and has query execution capability under the current dynamic execution period; the distributed query refers to a query mode in which a query request is divided into multiple sub-tasks, executed in parallel on a distributed node set, and finally the query results are aggregated and returned.

[0091] It can be seen that in the present application, the accuracy of query scheduling can be improved in the case of query state lag and node load imbalance in distributed big data query; wherein, by extracting the distributed query information of the to-be-processed data in the distributed node cluster and extracting the query gradient constraint on all nodes, the dynamic distribution state of the query features of each node in the big data platform is accurately identified and uniformly modeled; by determining the query evolution sequence according to the distributed query information and performing parallel matching, the evolution trend identification and retrieval state labeling of the multi-node query behavior path can be realized, so as to construct a dynamic query label system with accurate response and adaptive scheduling in a multi-node high-concurrency environment, effectively solving the problems of state identification lag and scheduling inconsistency; by extracting the static hierarchical attributes of the to-be-queried data items and cross- verifying combined with the query gradient constraint, the reliable evaluation capability of the query shard state in the distributed environment can be enhanced, and through cross-matching of the feedback label and the shard rule, the consistency and traceability of the shard query result can be improved; by performing distributed query according to the dynamic parallel label and the cross-trust trajectory, the linkage scheduling of the query behavior features and the data shard state is realized, and the execution node and the query load are accurately matched, so as to improve the execution efficiency of the whole-process distributed query in big data analysis.

[0092] In summary, the technical scheme adopted in the present application can perform global perception verification on to-be-processed data under a distributed architecture to improve the adaptability of distributed query in big data analysis.

[0093] Embodiment two, the present application provides a big data analysis system based on artificial intelligence, as shown in Figure 4 The figure is a module structure diagram of a big data analysis system based on artificial intelligence according to the embodiment of the present application, the analysis system comprises:

[0094] The data reading module 100 is used for reading the distributed query information of the to-be-processed data in the big data platform through the distributed node cluster, and extracting the query gradient constraint on all nodes in the distributed node cluster;

[0095] The parallel matching module 200 is configured to determine a query evolution sequence in the data query behavior according to the distributed query information, perform parallel matching on the query evolution sequence, and obtain a dynamic parallel label corresponding to the query retrieval state;

[0096] The cross verification module 300 is configured to extract a static hierarchical attribute corresponding to the data item to be queried under a distributed architecture, perform cross verification on the query gradient constraint on all nodes in the distributed node cluster and the static hierarchical attribute corresponding to the data item to be queried, and obtain a cross-trust trajectory corresponding to the query fragmentation state.

[0097] The distributed query module 400 is configured to perform distributed query on the data item to be queried in the big data platform according to the dynamic parallel label corresponding to the query retrieval state and the cross-trust trajectory corresponding to the query fragmentation state.

[0098] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device realize functions specified in the flowcharts and / or block diagrams. Figure 1 The device for implementing the functions specified in one flow or multiple flows and / or blocks Figure 1 The device for implementing the functions specified in one flow or multiple flows and / or blocks

[0099] Those skilled in the art can understand that all or part of the steps of various methods in the above embodiments can be completed by instructing the relevant hardware by means of a program, and the program can be stored in a computer readable storage medium, including Read-Only Memory (ROM), Random Access Memory (RAM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), One-time Programmable Read-Only Memory (OTPROM), Electrically-Erasable Programmable Read-Only Memory (EEPROM), Compact Disc Read-Only Memory (CD-ROM) or other optical disk storage, magnetic disk storage, magnetic tape storage, or any other medium that can be used to carry or store data in a computer readable manner.

[0100] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements, but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising a" does not exclude the existence of additional identical elements in the process, method, article, or apparatus that includes the element.

Claims

1. A big data analysis method based on artificial intelligence, characterized in that: The analysis method comprises the following steps: Read the distributed query information of the data to be processed in the big data platform through the distributed node cluster, and extract the query gradient constraints on all nodes in the distributed node cluster; Extracting query gradient constraints on all nodes in a distributed node cluster specifically includes: Determining local gradient features on all nodes based on the distributed node cluster; wherein the local gradient features refer to the trajectory of query performance indicators of each node changing with resource usage, data access mode, and query complexity during query execution over a period of time. By extracting the execution statement of each query, the data tables involved, the filter conditions, the grouping operation, the data volume, and the execution time, the gradient calculation method is used to obtain the change trend of each indicator in different time periods to generate local gradient features on all nodes; Determine query constraints during data query based on local gradient features on all nodes; Determine the query gradient constraints on all nodes in the distributed node cluster using the query constraint conditions; The query gradient constraint refers to the dynamic restriction imposed by each node on the query operation in a specific time window; Determining a query evolution sequence in a data query behavior according to the distributed query information, performing parallel matching on the query evolution sequence, and obtaining a dynamic parallel tag corresponding to a query retrieval state; Determining the query evolution sequence in the data query behavior according to the distributed query information specifically includes: Extracting query operation logs on each node according to the distributed query information; Construct the corresponding query state transition record based on the query operation log on each node; Generate the query evolution sequence in the data query behavior based on all query state transition records; The query evolution sequence refers to a time series sequence that describes the dynamic evolution path of query behavior during the query life cycle; The parallel matching of the query evolution sequence to obtain the dynamic parallel tag corresponding to the query retrieval state specifically includes: Determining query subsequence sharding during data parallel processing according to the query evolution sequence; Generate state transition rules that match query subsequence shards based on a preset query pattern library; Matching the query subsequence fragments with the state transition rules to obtain a query matching trend corresponding to the query retrieval state; Determine a dynamic parallel tag corresponding to a query retrieval state based on the query matching trend; The dynamic parallel tag refers to the parallel processing feature used to mark the current query retrieval state, the state transition rule represents the association rule of the transition conditions and logical relationships between query states, and the query matching trend refers to the representation of the tendency of the query subsequence fragments on the state transition path; Extract the static hierarchical attributes corresponding to the query data item in the distributed architecture, perform cross-validation based on the query gradient constraints on all nodes in the distributed node cluster and the static hierarchical attributes corresponding to the query data item, and obtain the cross-trusted trajectory corresponding to the query shard status; Extracting static hierarchical attributes corresponding to the data item to be queried under a distributed architecture specifically includes: Building hierarchical attribute indexes in distributed architectures; Parsing associated feature tags of the data item to be queried based on the hierarchical attribute index; Determining the static hierarchical attribute corresponding to the data item to be queried according to the associated feature tag; The static hierarchical attributes refer to the attribute hierarchical information that is inherent in the data item to be queried and does not change frequently over time; Among them, cross-validation is performed based on the query gradient constraints on all nodes in the distributed node cluster and the static hierarchical attributes corresponding to the queried data items. The cross-trusted trajectory corresponding to the query shard status is obtained, which specifically includes: Determine the shard balance constraint in the query shard state based on the query gradient constraint on all nodes in the distributed node cluster; Determine the shard feedback label in the query shard status according to the static hierarchical attribute corresponding to the data item to be queried; Determining trusted query information when querying shard status cross-matching according to the shard balance constraint and the shard feedback tag; Determine a cross-trusted track corresponding to the query shard status according to the trusted query information; The cross-trusted trajectory refers to the state that reflects the evolution of query shard status changes over time. The shard balance constraint refers to the balance rule that must be followed when assigning tasks to each query shard. The shard feedback label is an indicator that measures the query adaptability of the data in the query shard status. The trusted query information refers to the execution information of the query task after multi-dimensional verification and cross-matching. Perform distributed queries on the data items to be queried in the big data platform based on the dynamic parallel tags corresponding to the query retrieval status and the cross-trusted tracks corresponding to the query sharding status; The distributed query of the data items to be queried in the big data platform based on the dynamic parallel tags corresponding to the query retrieval status and the cross-trusted tracks corresponding to the query sharding status specifically includes: determining a dynamic execution period of a distributed query according to the dynamic parallel tag and the cross trusted trace; By matching the dynamic execution cycle with the query load threshold of the big data platform, a candidate execution node set for distributed query is screened; A distributed query is executed on the data item to be queried according to the candidate execution node set.

2. The big data analysis method based on artificial intelligence according to claim 1, characterized in that: The distributed query information refers to a collection of statements, paths, and resources generated by each distributed node during the query process.

3. The big data analysis method based on artificial intelligence according to claim 1, characterized in that: The query retrieval status refers to the specific running status of the query task in different stages and links during the data query process.

4. The big data analysis method based on artificial intelligence according to claim 1, characterized in that: The query shard status refers to the comprehensive performance between the execution status of each data query task on different shards and the current operation characteristics during the distributed query process.

5. The big data analysis method based on artificial intelligence according to claim 1, characterized in that: The distributed query is a query method that divides a query request into multiple subtasks, executes them in parallel on a distributed set of nodes, and finally summarizes and returns the query results.

6. An artificial intelligence-based big data analysis system, used to execute the artificial intelligence-based big data analysis method according to any one of claims 1 to 5, characterized in that: The analysis system comprises: The data reading module is used to read the distributed query information of the data to be processed in the big data platform through the distributed node cluster and extract the query gradient constraints on all nodes in the distributed node cluster; A parallel matching module is used to determine the query evolution sequence in the data query behavior according to the distributed query information, perform parallel matching on the query evolution sequence, and obtain a dynamic parallel tag corresponding to the query retrieval state; The cross-validation module is used to extract the static hierarchical attributes corresponding to the query data item in a distributed architecture, and perform cross-validation based on the query gradient constraints on all nodes in the distributed node cluster and the static hierarchical attributes corresponding to the query data item to obtain the cross-trusted trajectory corresponding to the query shard status; The distributed query module is used to perform distributed queries on the data items to be queried in the big data platform based on the dynamic parallel tags corresponding to the query retrieval status and the cross-trusted tracks corresponding to the query sharding status.

Citation Information

Patent Citations

  • Enterprise financial document unified management system and method based on distributed storage

    CN120430878A

  • Methods and systems for a database

    WO2018170276A2