An information analysis method, device and medium based on artificial intelligence and big data
Patent Information
- Application Number
- CN202511484619.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2045-10-17
AI Technical Summary
[0005]因此,本发明提供了一种基于人工智能和大数据的信息分析方法解决性能分析效率低下和缺乏优化闭环问题
[0016] The beneficial effects of this invention are as follows: by using a dynamic performance evaluation and anomaly early warning mechanism, a dynamic performance benchmark is established by adopting a time-effect decay algorithm, and anomaly detection technology is combined to identify deviation nodes and generate multi-level early warning signals, which effectively responds to changes in data distribution and reduces the false alarm rate; it improves analysis efficiency and reduces the cost of manual intervention, providing strong technical support for intelligent decision-making in complex environments.
Smart Images

Figure CN121388802B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data analytics, and in particular to an information analysis method, device, and medium based on artificial intelligence and big data. Background Technology
[0002] Artificial intelligence and big data analytics are at the core of modern information processing. Multimodal data processing technology integrates various data types such as text, images, and audio, and utilizes deep learning models to achieve cross-modal feature extraction and semantic fusion, improving the comprehensiveness and accuracy of data understanding. Distributed computing frameworks support parallel processing and dynamic resource scheduling of large-scale data, improving analysis efficiency and method scalability. Anomaly detection algorithms can automatically identify abnormal patterns and outliers in data, enhancing reliability and real-time monitoring capabilities. Visual analysis tools transform complex data into intuitive insights through interactive dashboards and graphical reports, assisting users in decision-making.
[0003] In the current field of artificial intelligence and big data analytics, despite progress in multimodal data processing, distributed computing frameworks, and anomaly detection algorithms, two limitations still constrain further performance improvements. Firstly, dynamic performance evaluation mechanisms are inadequate; existing methods generally rely on static performance benchmarks or fixed thresholds for anomaly detection, lacking real-time adaptability to changes in data distribution and concept drift. Secondly, closed-loop optimization processes suffer from structural deficiencies; existing solutions often terminate at the stage of generating visualization reports, lacking in-depth root cause analysis and automated optimization capabilities. This not only reduces processing efficiency but also easily introduces human error, making it difficult for methods to achieve real-time response and continuous improvement in high-speed data stream scenarios. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides an information analysis method based on artificial intelligence and big data to solve the problems of low efficiency and lack of optimization loop in performance analysis.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides an information analysis method based on artificial intelligence and big data, comprising: collecting a big data task package to be analyzed and performing fusion processing through multimodal collaborative analysis to generate a multi-source heterogeneous dataset; analyzing the multi-source heterogeneous dataset by constructing a distributed analysis link to generate a preliminary analysis dataset; calculating the dynamic performance benchmark of the distributed analysis link based on the preliminary analysis dataset, identifying deviation nodes using an anomaly detection algorithm to generate anomaly node identifiers, and generating a performance anomaly warning signal by quantifying the deviation value; performing feature importance ranking and decision path node backtracking on the performance anomaly warning signal to generate a visual diagnostic report containing root cause data chains and decision defect nodes; constructing an information analysis optimization model based on the visual diagnostic report, analyzing and optimizing the root cause data chains and decision defect nodes to generate corrective data features, and generating analysis optimization instructions through decision logic reconstruction; executing the analysis optimization instructions to verify and re-analyze the visual diagnostic report to obtain a strategy implementation effect evaluation report.
[0007] As a preferred embodiment of the information analysis method based on artificial intelligence and big data described in this invention, the big data task package to be analyzed includes task target features and data to be analyzed; The collected big data task package is fused through multimodal collaborative analysis to generate a multi-source heterogeneous dataset. A distributed analysis pipeline is then constructed to analyze this multi-source heterogeneous dataset, generating a preliminary analysis dataset. The specific steps are as follows: Based on the characteristics of the task objectives and the data to be analyzed, feature fusion is performed through multimodal collaborative analysis to generate multi-source heterogeneous datasets; Analyze the data modes of multi-source heterogeneous datasets and construct resource allocation schemes by calculating data complexity; Based on the resource allocation scheme, a distributed analysis link is constructed by dynamically deploying image computing nodes and text computing nodes; By sharing a semantic space mapping, image computing nodes and text computing nodes are used to generate a preliminary analysis dataset.
[0008] As a preferred embodiment of the information analysis method based on artificial intelligence and big data described in this invention, the following steps are taken: Based on the preliminary analysis dataset, a dynamic performance benchmark of the distributed analysis link is calculated; an anomaly detection algorithm is used to identify deviation nodes and generate anomaly node identifiers; and a performance anomaly warning signal is generated by quantifying the deviation value. Based on the preliminary analysis dataset, the time-degradation performance evaluation algorithm is used to perform exponential degradation calculation on the distributed analysis link to generate a dynamic performance benchmark. Based on the dynamic performance benchmark, an anomaly detection algorithm is used to identify deviation nodes in image computing nodes and text computing nodes. The deviation nodes are marked to obtain the abnormal node identifiers, and the performance abnormality warning signal is generated by calculating the node deviation value.
[0009] As a preferred embodiment of the information analysis method based on artificial intelligence and big data described in this invention, the specific steps for performing feature importance ranking and decision path node backtracking on the performance anomaly warning signal are as follows: Based on the performance anomaly warning signal, time-series features, spatial correlation features, and performance index features are extracted through multimodal signal analysis to form a signal feature set; Based on the signal feature set, a feature importance ranking map is generated by calculating the contribution of signal features; Based on the feature importance ranking map, a time-series backtracking algorithm is used to track the propagation trajectory of abnormal early warning signals and generate defect node location information.
[0010] As a preferred embodiment of the information analysis method based on artificial intelligence and big data described in this invention, the step of generating a visual diagnostic report containing a root cause data chain and decision defect nodes refers to constructing a root cause data chain by combining a feature importance ranking graph with defect node location information, locating decision defect nodes through a graph structure analysis algorithm, and generating a visual diagnostic report.
[0011] As a preferred embodiment of the information analysis method based on artificial intelligence and big data described in this invention, the specific steps for constructing the information analysis optimization model based on the visualized diagnostic report are as follows: Based on the visual diagnostic report, a multimodal fusion analysis method is used to extract performance indicators and abnormal patterns to generate an optimized dataset; Based on the optimized dataset, a multi-objective optimization algorithm is used to optimize the accuracy, efficiency, and resource consumption of decision-making defect nodes, and preliminary optimization parameters are generated. Cross-validation was performed on the initial optimized parameters, and an information analysis optimization model was constructed by calculating performance scores.
[0012] As a preferred embodiment of the information analysis method based on artificial intelligence and big data described in this invention, the steps of analyzing and optimizing the root cause data chain and decision defect nodes to generate corrective data features, and generating analysis and optimization instructions through decision logic reconstruction, are as follows: Based on the information analysis optimization model, the correlation characteristics and impact paths of the root cause data chain are extracted to generate a root cause analysis report; Based on the root cause analysis report, the impact of decision-making defect nodes is assessed to generate corrective data features, and then the decision logic is reconstructed to generate analysis and optimization instructions.
[0013] As a preferred embodiment of the information analysis method based on artificial intelligence and big data described in this invention, the steps of executing analysis and optimization instructions to verify and re-analyze the visualized diagnostic report to obtain a strategy implementation effect evaluation report are as follows. The analysis and optimization instructions are transformed into an execution operation sequence using a bidirectional instruction mapping method, and a verification environment is built. Temporal feature extraction is performed in the verification environment, and the reconstructed data feature set is obtained by dynamically adjusting the sequence of execution operations. Anomaly detection is performed on the reconstructed data feature set to generate anomaly distribution maps, and the effectiveness of the strategy implementation is evaluated by quantifying the strategy implementation effect.
[0014] In a second aspect, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein when the computer program is executed by the processor, it implements any step of the information analysis method based on artificial intelligence and big data as described in the first aspect of the present invention.
[0015] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the information analysis method based on artificial intelligence and big data as described in the first aspect of the present invention.
[0016] The beneficial effects of this invention are as follows: by using a dynamic performance evaluation and anomaly early warning mechanism, a dynamic performance benchmark is established by adopting a time-effect decay algorithm, and anomaly detection technology is combined to identify deviation nodes and generate multi-level early warning signals, which effectively responds to changes in data distribution and reduces the false alarm rate; it improves analysis efficiency and reduces the cost of manual intervention, providing strong technical support for intelligent decision-making in complex environments. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of an information analysis method based on artificial intelligence and big data.
[0019] Figure 2 A flowchart for generating a preliminary analysis dataset.
[0020] Figure 3 A flowchart for generating performance anomaly warning signals.
[0021] Figure 4A flowchart for obtaining a strategy implementation effectiveness evaluation report. Detailed Implementation
[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0025] Reference Figures 1-4 This is one embodiment of the present invention, which provides an information analysis method based on artificial intelligence and big data, including the following steps: S1: Collect the big data task package to be analyzed, perform multimodal collaborative analysis to generate multi-source heterogeneous datasets, and analyze the multi-source heterogeneous datasets by constructing a distributed analysis link to generate preliminary analysis datasets.
[0026] S1.1: The big data task package to be analyzed includes the task objective features and the data to be analyzed; Specifically, the target features of the big data task package to be analyzed come from the analysis requirements submitted by the user or the instructions input from the upper-level application. They usually include analysis dimensions (such as image recognition, text classification, time series prediction, etc.), accuracy requirements (such as accuracy threshold, recall standard), processing time limit (such as real-time response or batch processing), and resource constraints (such as computing resource quota, memory limit).
[0027] The role of task objective features is to provide a clear guiding framework for the entire analysis process, ensuring that subsequent processing steps meet the expected goals. For example, setting image recognition accuracy requirements can guide the depth of feature extraction, or processing time constraints (such as real-time response time requirements or batch processing time limits) can dynamically adjust the resource allocation strategy of distributed nodes.
[0028] The data to be analyzed comes from diverse and heterogeneous data sources, including image data (such as surveillance video frames and medical images), text data (such as log files and user comments), sensor data (such as temperature readings and device status signals), and time-series data (such as transaction records and network traffic). Data is collected and preprocessed in a standardized manner through a unified access interface to form a structured dataset available for analysis.
[0029] The role of the data to be analyzed is to serve as the core processing object, providing raw materials for multimodal collaborative analysis. For example, image data is used to extract visual features through convolutional neural networks, text data is used to extract semantic features through natural language processing, and sensor data is used to capture dynamic patterns (such as trend changes in temperature sensors, periodic fluctuations in heart rate monitoring, and abnormal detection patterns of equipment vibration) through time series analysis. A unified representation is formed through cross-modal fusion.
[0030] S1.2: Based on the task target features and the data to be analyzed, feature fusion is performed through multimodal collaborative analysis to generate a multi-source heterogeneous dataset; It should be noted that the image data in the data to be analyzed; visual features are obtained through color distribution analysis, texture pattern recognition (e.g., analyzing pixel intensity variations to distinguish smooth, rough, or regular textures) and shape contour detection (e.g., outlining the outer contour and inner boundaries of objects) in the analysis dimension of the task objective features; semantic features are obtained through word frequency statistics (e.g., listing high-frequency and rare words), semantic association analysis (e.g., analyzing word collocation patterns in sentences) and topic classification processing (e.g., based on keyword matching and content grouping) in the text data in the data to be analyzed based on the accuracy requirements of the task objective features; dynamic features are obtained through time series smoothing, trend line fitting, and outlier labeling in the sensor data in the data to be analyzed based on the processing time limit of the task objective features.
[0031] By aligning visual features, semantic features, and dynamic features to a unified time dimension, a multi-source heterogeneous dataset containing a unified representation of multiple modalities is generated.
[0032] S1.3: Analyze the data modalities of multi-source heterogeneous datasets and construct a resource allocation scheme by calculating data complexity; It should be noted that identifying the data types contained in the multi-source heterogeneous dataset, such as image data, text data, and sensor data; for image data, features such as resolution and number of color channels are determined by examining the image data structure and format; for text data, data modality analysis determines features such as character set and vocabulary type by analyzing encoding methods and content structure; for sensor data, data modality analysis determines features by evaluating data acquisition parameters (e.g., number of data points and sampling frequency).
[0033] The data complexity of image data is the product of its resolution and the number of color channels; the data complexity of text data is evaluated by the total number of characters in the character set and the diversity of vocabulary types (such as nouns, verbs, and adjectives); the data complexity of sensor data is the product of the number of sensor data points and the sampling frequency.
[0034] Allocate more image computing node resources to image data with higher data complexity; ensure that computing resources match data complexity to generate a resource allocation scheme.
[0035] S1.4: Based on the resource allocation scheme, a distributed analysis link is constructed by dynamically deploying image computing nodes and text computing nodes; It should be noted that, based on the data complexity and image computing node resources in the resource allocation scheme, the deployment process for image computing nodes includes configuring the software environment required to process image data, such as deploying containerized instances to run image processing libraries and deep learning frameworks to support color distribution analysis, texture pattern recognition, and shape contour detection operations.
[0036] For text computing nodes, the deployment process involves setting up memory allocation for processing text data and a natural language processing (NLP) toolchain to support semantic feature analysis. This includes, for example, integrating NLP libraries for word frequency statistics, semantic association analysis, and topic classification.
[0037] Visual features (color distribution features, texture pattern features, and shape contour features) output by image computing nodes and semantic features (word frequency features, semantic association features, and topic classification features) output by text computing nodes are temporally aligned in a unified dimensional space to construct a distributed analysis link.
[0038] S1.5: Generate a preliminary analysis dataset by sharing a semantic space mapping between image computing nodes and text computing nodes.
[0039] It should be noted that the visual features output by the image computing node after processing image data include color distribution features, texture pattern features, and shape contour features; the semantic features output by the text computing node after processing text data include word frequency features, semantic association features, and topic classification features.
[0040] By mapping a shared semantic space, a unified feature representation framework is established, mapping visual and semantic features to the same dimensional space. This allows color distribution features in the visual features to correspond spatially with topic classification features in the semantic features, generating a shared semantic space. Furthermore, texture pattern features are associated with semantic association features, and shape contour features are paired with word frequency features. Within this shared semantic space, visual and semantic features are fused to eliminate modal differences and generate a preliminary analysis dataset.
[0041] S2: Based on the preliminary analysis dataset, calculate the dynamic performance benchmark of the distributed analysis link, and use an anomaly detection algorithm to identify deviation nodes and generate anomaly node identifiers, and generate performance anomaly warning signals by quantifying deviation values.
[0042] S2.1: Based on the preliminary analysis dataset, the time-degradation performance evaluation algorithm is used to perform exponential degradation calculation on the distributed analysis link to generate a dynamic performance benchmark; It should be noted that, based on the preliminary analysis dataset, which includes performance benchmark values (such as processing latency, resource utilization, and throughput) for each node in the distributed analysis chain, a time-degradation performance evaluation algorithm is used to process the performance index data through an exponential decay function. The decay factor is objectively set based on the frequency of historical data changes; for example, the decay factor value is automatically obtained by analyzing the volatility and time intervals of data points, ensuring that recent data naturally has a higher impact and older data naturally has a lower impact. An exponential decay calculation is performed to generate a dynamic performance benchmark, expressed as: ; in, As a dynamic performance benchmark, The total number of data points. The index value of the data point. It is an exponentially decaying weighting function. It is an exponential function. As the attenuation factor, For the current time and data point The absolute difference between times For the first The performance metric values for each data point.
[0043] S2.2: Based on the dynamic performance benchmark, an anomaly detection algorithm is used to identify deviation nodes in image computing nodes and text computing nodes; It should be noted that the performance benchmark values of the distributed links in the dynamic performance benchmark (such as processing latency, resource utilization, and throughput) are used to ensure data timeliness by monitoring and collecting the current performance index values of image computing nodes and text computing nodes in real time. Next, the current performance index value of each node is compared with the corresponding performance benchmark value one by one. Image computing nodes and text computing nodes that deviate from the performance benchmark value are identified and designated as deviation nodes.
[0044] S2.3: Mark the deviation nodes to obtain the abnormal node identifiers, and generate a performance abnormality warning signal by calculating the node deviation value.
[0045] It should be noted that the deviation nodes are spatiotemporally bound to generate a unique code, and a timestamp is recorded to identify the abnormal nodes. For example, a unique identifier containing the node type abbreviation and sequence number is generated for each deviation node, and the current timestamp is attached to ensure uniqueness and traceability.
[0046] The difference between the current performance index value and the performance benchmark value is used as the node deviation value. Warning levels are generated through signal classification mapping. For example, a low-risk warning signal is generated when the node deviation value is in the range of 10%-25%, a medium-risk warning signal is generated when it is in the range of 25%-50%, and a high-risk warning signal is generated when it is above 50%. The warning level and the corresponding node deviation value are integrated into a performance abnormality warning signal.
[0047] S3: Perform feature importance ranking and decision path node backtracking on performance anomaly warning signals to generate a visual diagnostic report containing root cause data chains and decision defect nodes.
[0048] S3.1: Based on the performance anomaly warning signal, time-series features, spatial correlation features, and performance index features are extracted through multimodal signal analysis to form a signal feature set; It should be noted that the performance anomaly warning signal includes node deviation value and warning level; Based on the timestamp of the performance anomaly warning signal, the signal triggering time interval, duration, and occurrence frequency pattern are recorded as timing features. Record the signal propagation trajectory in the performance anomaly warning signal, and obtain the hop count and transmission delay between the abnormal node and the root cause node as the communication path distance, which serves as a spatial correlation feature. At the same time, performance index data is directly read from the performance anomaly warning signal, and performance index features are extracted, including the degree of processing latency deviation, the degree of resource utilization deviation, and the degree of throughput deviation. Finally, the temporal features, spatial correlation features, and performance index features are integrated into a signal feature set.
[0049] S3.2: Generate a feature importance ranking map by calculating the contribution of signal features based on the signal feature set; It should be noted that the signal feature set includes time-series features, spatial correlation features, and performance index features. The product of time-series features and their weight coefficients, the product of spatial correlation features and their weight coefficients, and the product of performance index features and their weight coefficients are respectively used as the weighted values of time-series features, spatial correlation features, and performance index features. The weighting coefficients for temporal features, spatial correlation features, and performance index features are determined by analyzing the frequency of different features' influence on anomaly diagnosis, with values ranging from [value range missing]. For example, the weight of time series features is often set to 0.4, the weight of spatial correlation features is 0.3, and the weight of performance index features is 0.3. The specific values need to be adjusted according to the actual data distribution.
[0050] The contribution values are the ratios of the weighted values of time-series features, spatial correlation features, and performance index features to the total weighted value. The time-series features, spatial correlation features, and performance index features are sorted from high to low according to their contribution values to generate a feature importance ranking map. S3.3: Based on the feature importance ranking map, a time-series backtracking algorithm is used to track the propagation trajectory of abnormal early warning signals and generate defect node location information; It should be noted that the feature importance ranking map includes a list sorted by feature contribution and corresponding contribution values; the time-series backtracking algorithm analyzes the timestamp sequence of the anomaly warning signal to identify the initial anomaly occurrence time point, gradually backtracks to the preceding time node, records the signal strength change trajectory, and reverses the signal propagation path to obtain the anomaly propagation path map; it prioritizes tracking the propagation link corresponding to the feature with the highest contribution, locates the defect node that caused the anomaly, generates a defect node identifier, and generates defect node location information including the defect node identifier, the anomaly propagation path map, the timestamp sequence, and the feature contribution ranking list.
[0051] S3.4: Construct a root cause data chain by combining the feature importance ranking map with the defect node location information, and locate the decision defect node through graph structure analysis algorithm to generate a visual diagnostic report.
[0052] It should be noted that the feature importance ranking map includes a feature contribution ranking list, and the defect node location information includes defect node identifiers, anomaly propagation paths, and timestamp sequences; the contribution features in the feature importance ranking map are associated and mapped with the propagation paths in the defect node location information to obtain the root cause data chain.
[0053] Based on the root cause data chain, a graph structure analysis algorithm (such as graph neural network and influence propagation) is used to analyze the topological relationship between decision-making defect nodes and identify decision-making defect nodes with logical errors; the root cause data chain and decision-making defect node information are integrated to generate a visual diagnostic report.
[0054] S4: Based on the visual diagnostic report, build an information analysis and optimization model, analyze and optimize the root cause data chain and decision defect nodes to generate corrective data features, and generate analysis and optimization instructions through decision logic reconstruction.
[0055] S4.1: Based on the visual diagnostic report, a multimodal fusion analysis method is used to extract performance indicators and abnormal patterns to generate an optimized dataset; It should be noted that, based on the visual diagnostic report, which contains detailed information on the root cause data chain and decision-making defect nodes, a multimodal fusion parsing method is used to perform data parsing of the text content in the visual diagnostic report and extract performance indicator data, such as processing latency values, resource utilization percentage, and throughput rate. Based on the chart elements in the visual diagnostic report, extract abnormal pattern information, such as error type classification, occurrence frequency statistics, and impact range description; then integrate the extracted performance index data and abnormal pattern information to generate a structured optimization dataset.
[0056] S4.2: Based on the optimized dataset, use a multi-objective optimization algorithm to optimize the accuracy, efficiency, and resource consumption of decision-making defect nodes, and generate preliminary optimization parameters; It should be noted that, based on the optimized dataset, which includes performance metrics extracted from the visual diagnostic report, a multi-objective optimization algorithm is used to perform the optimization operation. The multi-objective optimization algorithm optimizes the accuracy, efficiency, and resource consumption objectives of the defective nodes. For example, the accuracy objective focuses on improving the accuracy of node processing, the efficiency objective focuses on reducing processing time, and the resource consumption objective aims to reduce the usage of memory and computing resources. The optimization process defines objective functions, such as the accuracy maximization function, the efficiency maximization function, and the resource consumption minimization function, and generates a set of parameter settings and performance evaluation data as preliminary optimization parameters.
[0057] S4.3: Perform cross-validation on the initial optimized parameters and construct an information analysis optimization model by calculating performance scores; It should be noted that cross-validation is performed based on the initial optimized parameters. Cross-validation divides the initial optimized parameters into multiple mutually exclusive subsets, using one subset as the test set and the remaining subsets as the training set. The initial optimized parameters are then applied to evaluate their generalization performance on different data subsets. Subsequently, the evaluation results are quantified by calculating the performance score of the initial optimized parameters. Finally, the optimal parameter configuration is integrated based on the performance score, selecting the parameter combination that performs best on most indicators to construct the information analysis optimization model. The expression for calculating the performance score of the initial optimized parameters is as follows: ; in, To initially optimize the performance score of the parameters, The total number of folds for cross-validation. This is the index value of the total number of folds in cross-validation. The total number of performance metrics. This is the index value of the total number of performance metrics. These are the weighting coefficients for performance indicators. For the first Normalization function for each performance metric For the first The first result calculated on the test set The original values of each indicator.
[0058] The weighting coefficients of the performance metrics are derived from the task objective features learned through optimization algorithms, and their values range from [value range missing]. .
[0059] S4.4: Based on the information analysis optimization model, extract the correlation characteristics and impact paths of the root cause data chain to generate a root cause analysis report; It should be noted that the weight coefficients of the performance indicators in the information analysis optimization model correspond to the time-series features, spatial correlation features, and performance indicator features with the highest contribution in the feature importance ranking map, and feature node identifiers are obtained through feature mapping; the feature node identifiers are associated with the performance indicator data in the root cause data chain to obtain the associated features containing feature names, weight values, and data sources.
[0060] Based on the abnormal propagation path and timestamp sequence recorded in the defect node location information in the visual diagnostic report; the defect nodes are arranged according to the timestamp sequence and the propagation direction between defect nodes is marked to generate an impact path containing the path node sequence and propagation timeline, and a root cause analysis report is generated by combining the correlation features.
[0061] S4.5: Based on the root cause analysis report, assess the impact of decision-making defect nodes, generate corrective data features, and generate analysis and optimization instructions through decision logic reconstruction.
[0062] It should be noted that the root cause analysis report includes correlation features and impact paths obtained from the information analysis optimization model; based on the visualization diagnostic report, the correlation strength of the correlation features (such as the thickness of the lines between feature nodes in the visualization diagnostic report and the color depth representing the strength level) is matched with the path criticality of the impact paths (such as the centrality and connection density of the path in the topology), for example, the deeper the color depth of the impact node, the stronger the path criticality of the impact path; according to the impact level results, the resource allocation ratio and response latency limit are adjusted, for example, the resource allocation ratio is adjusted first for high-impact nodes to reduce the risk of failure, and corrective data features are generated; Based on the characteristics of the calibration data, the initial optimization parameters are modified, and the efficiency of the process is optimized by adjusting the priority of the decision path, such as rearranging the path order or setting conditional branches, and analysis optimization instructions are generated.
[0063] S5: Execute the analysis and optimization command to verify and reanalyze the visual diagnostic report, and obtain a strategy implementation effect evaluation report.
[0064] S5.1: Transform the analysis and optimization instructions into an execution operation sequence using bidirectional instruction mapping and build a verification environment; It should be noted that the command type (e.g., parameter adjustment command, path priority command, or resource allocation command) and parameter details (e.g., numerical settings, conditional rules, or configuration options) in the parsing and analysis optimization instructions are transformed into an executable sequence of operation steps. For example, the "optimize node parameters" instruction is mapped to the specific operation commands "adjust processing frequency" and "set memory threshold" to obtain the execution operation sequence. Hardware resource simulation (such as allocating virtual computing nodes and memory space) and software tool deployment (such as installing data analysis libraries and testing frameworks) are performed based on the execution operation sequence to ensure that the environmental conditions match the real scenario and output a verification environment.
[0065] S5.2: Perform temporal feature extraction in the verification environment, and obtain the reconstructed data feature set by dynamically adjusting the execution operation sequence; It should be noted that in the validation environment, time series features are captured by analyzing time series data, such as the timestamp intervals of data points, identifying trends, and detecting periodic fluctuations. Adjust the processing frequency in the execution operation sequence according to the time sequence characteristics to change the resource allocation ratio, obtain the adjusted execution operation sequence to adapt to periodic needs; integrate the update characteristics and time sequence characteristics in the adjusted execution operation sequence to generate a reconstructed data feature set.
[0066] S5.3: Perform outlier detection on the reconstructed data feature set to generate an anomaly distribution map, and obtain a strategy implementation effect evaluation report by quantifying the strategy implementation effect.
[0067] It should be noted that the reconstructed data feature set includes time-series features extracted from the verification environment and updated features generated by dynamically adjusting the execution operation sequence; the total number of data feature values and the total value of data features in the reconstructed data feature set are statistically analyzed, and the ratio of the total value of data features to the total number of data feature values is used as the anomaly baseline value; the data feature values are matched with the anomaly baseline value to identify anomalous data points that exceed the anomaly baseline value, while those that do not exceed the anomaly baseline value are not marked; the anomalous data points in the reconstructed data feature set are statistically analyzed to generate an anomaly distribution map; the anomaly distribution map records the location, frequency, and severity of anomalies in a structured format; and the location, frequency, and severity data of anomalies are integrated to generate a strategy implementation effectiveness evaluation report.
[0068] This embodiment also provides a computer device applicable to information analysis methods based on artificial intelligence and big data, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the information analysis method based on artificial intelligence and big data as proposed in the above embodiment.
[0069] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0070] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the information analysis method based on artificial intelligence and big data as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0071] In summary, this invention, through a dynamic performance evaluation and anomaly early warning mechanism, establishes a dynamic performance benchmark using a time-decrease algorithm, and combines anomaly detection technology to identify deviation nodes and generate multi-level early warning signals, effectively responding to changes in data distribution and reducing false alarm rates; it improves analysis efficiency and reduces the cost of manual intervention, providing strong technical support for intelligent decision-making in complex environments.
[0072] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. An information analysis method based on artificial intelligence and big data, characterized in that: include, The collected big data task packages are fused through multimodal collaborative analysis to generate multi-source heterogeneous datasets. These datasets are then analyzed using a distributed analysis pipeline to generate preliminary analysis datasets. The specific steps are as follows: The big data task package to be analyzed includes task target features and data to be analyzed. Based on the characteristics of the task objectives and the data to be analyzed, feature fusion is performed through multimodal collaborative analysis to generate multi-source heterogeneous datasets; Analyze the data modes of multi-source heterogeneous datasets and construct resource allocation schemes by calculating data complexity; Based on the resource allocation scheme, image computing nodes and text computing nodes are dynamically deployed to build a distributed analysis link; By sharing a semantic space mapping, image computing nodes and text computing nodes are used to generate a preliminary analysis dataset; Based on the preliminary analysis dataset, the dynamic performance benchmark of the distributed analysis link is calculated, anomaly detection algorithm is used to identify deviation nodes and generate anomaly node identifiers, and performance anomaly warning signals are generated by quantifying deviation values. Ranking the importance of performance anomaly warning signals by performance characteristics and backtracking the decision path nodes; The feature importance ranking map and defect node location information are used to construct a root cause data chain, and the decision defect node is located by graph structure analysis algorithm to generate a visual diagnostic report. The specific steps for building an information analysis and optimization model based on visualized diagnostic reports are as follows. Based on the visual diagnostic report, a multimodal fusion analysis method is used to extract performance indicators and abnormal patterns to generate an optimized dataset; Based on the optimized dataset, a multi-objective optimization algorithm is used to optimize the accuracy, efficiency, and resource consumption of decision-making defect nodes, and preliminary optimization parameters are generated. Cross-validation was performed on the initial optimized parameters, and an information analysis optimization model was constructed by calculating performance scores. The initial optimized parameters are subjected to cross-validation, which divides the initial optimized parameters into multiple mutually exclusive subsets. One subset is used as the test set, and the remaining subsets are used as the training set. The initial optimized parameters are then applied to evaluate their generalization performance on different data subsets. Subsequently, the evaluation results are quantified by calculating the performance score of the initial optimized parameters. Finally, the optimal parameter configuration is integrated based on the performance score, and the parameter combination that performs best on most indicators is selected to construct the information analysis optimization model. The expression for calculating the performance score of the initial optimized parameters is as follows: ; in, To initially optimize the performance score of the parameters, The total number of folds for cross-validation. This is the index value of the total number of folds in cross-validation. The total number of performance metrics. This is the index value of the total number of performance metrics. These are the weighting coefficients for performance indicators. For the first Normalization function for each performance metric For the first The first result calculated on the test set The original values of each indicator; The root cause data chain and decision defect nodes are analyzed and optimized to generate corrective data features, and analysis and optimization instructions are generated through decision logic reconstruction. Based on the information analysis optimization model, the process of extracting the correlation features and influence paths of the root cause data chain to generate a root cause analysis report involves obtaining feature node identifiers by mapping the time-series features, spatial correlation features, and performance indicator features with the highest contribution in the feature importance ranking map corresponding to the weight coefficients of performance indicators in the information analysis optimization model; and associating the feature node identifiers with the performance indicator data in the root cause data chain to obtain the correlation features containing feature names, weight values, and data sources. Based on the abnormal propagation path and timestamp sequence recorded in the defect node location information in the visual diagnostic report; arrange the defect nodes according to the timestamp sequence and mark the propagation direction between defect nodes to generate an impact path containing the path node sequence and propagation timeline, and generate a root cause analysis report by combining the correlation features; Based on the root cause analysis report, the impact of decision-making defect nodes is assessed to generate corrective data features. After the decision logic is reconstructed, analysis and optimization instructions are generated. This means that the root cause analysis report includes the correlation features and impact paths obtained from the information analysis and optimization model. Based on the visualization diagnostic report, the correlation strength of the correlation features is matched with the path criticality of the impact paths. According to the impact level results, the resource allocation ratio and response delay limit are adjusted to generate corrective data features. Based on the characteristics of the calibration data, the initial optimization parameters are modified, and by adjusting the priority of the decision path, analysis and optimization instructions are generated. The analysis and optimization commands are executed to validate and re-analyze the visualized diagnostic report, and a strategy implementation effectiveness evaluation report is obtained. The specific steps are as follows. The analysis and optimization instructions are transformed into an execution operation sequence using a bidirectional instruction mapping, and a verification environment is built. Temporal feature extraction is performed in the verification environment, and the reconstructed data feature set is obtained by dynamically adjusting the sequence of execution operations. Anomaly distribution maps are generated by performing outlier detection on the reconstructed data feature set, and an evaluation report on the effectiveness of the strategy implementation is obtained by quantifying the strategy implementation effect. Transforming analysis and optimization instructions into an execution sequence using bidirectional instruction mapping and building a verification environment refers to parsing the command type and parameter details in the analysis and optimization instructions, transforming abstract instructions into an executable sequence of operation steps, and obtaining the execution sequence. Hardware resources are simulated and software tools are deployed based on the execution operation sequence to ensure that the environmental conditions match the real scenario and output the verification environment. Extracting time-series features in a validation environment and obtaining a reconstructed data feature set by dynamically adjusting the execution operation sequence refers to capturing time-series features by analyzing time-series data in the validation environment, adjusting the processing frequency in the execution operation sequence according to the time-series features to change the resource allocation ratio, obtaining an adjusted execution operation sequence to adapt to periodic needs, and integrating the update features and time-series features in the adjusted execution operation sequence to generate a reconstructed data feature set. The process involves outlier detection and anomaly distribution mapping of a reconstructed data feature set, followed by obtaining a strategy implementation effectiveness evaluation report through quantification. This process involves: reconstructing the data feature set, which includes time-series features extracted from the validation environment and updated features generated by dynamically adjusting the execution sequence; statistically analyzing the total number and total value of data feature values in the reconstructed data feature set, using the ratio of the total number to the total number of data feature values as the anomaly baseline value; matching data feature values with the anomaly baseline value to identify outlier data points exceeding the baseline value, while those not exceeding the baseline value are not marked; statistically analyzing the outlier data points in the reconstructed data feature set to generate an anomaly distribution map; recording the location, frequency, and severity of outliers in a structured format in the anomaly distribution map; and integrating the location, frequency, and severity data of outliers to generate a strategy implementation effectiveness evaluation report.
2. The information analysis method based on artificial intelligence and big data as described in claim 1, characterized in that: Based on the preliminary analysis dataset, the dynamic performance benchmark of the distributed analysis link is calculated, and an anomaly detection algorithm is used to identify deviation nodes and generate anomaly node identifiers. Furthermore, a performance anomaly warning signal is generated by quantifying the deviation value. The specific steps are as follows: Based on the preliminary analysis dataset, the time-degradation performance evaluation algorithm is used to perform exponential degradation calculation on the distributed analysis link to generate a dynamic performance benchmark. Based on the dynamic performance benchmark, an anomaly detection algorithm is used to identify deviation nodes in image computing nodes and text computing nodes. The deviation nodes are marked to obtain the abnormal node identifiers, and the performance abnormality warning signal is generated by calculating the node deviation value.
3. The information analysis method based on artificial intelligence and big data as described in claim 2, characterized in that: The specific steps for performing feature importance ranking and decision path node backtracking on the performance anomaly warning signal are as follows: Based on the performance anomaly warning signal, time-series features, spatial correlation features, and performance index features are extracted through multimodal signal analysis to form a signal feature set; Based on the signal feature set, a feature importance ranking map is generated by calculating the contribution of signal features; Based on the feature importance ranking map, a time-series backtracking algorithm is used to track the propagation trajectory of abnormal early warning signals and generate defect node location information.
4. The information analysis method based on artificial intelligence and big data as described in claim 3, characterized in that: The specific steps for analyzing and optimizing the root cause data chain and decision-making defect nodes to generate corrective data features, and then reconstructing the decision logic to generate analysis and optimization instructions, are as follows: Based on the information analysis optimization model, the correlation characteristics and impact paths of the root cause data chain are extracted, and a root cause analysis report is generated. Based on the root cause analysis report, the impact of decision-making defect nodes is assessed to generate corrective data features, and then the decision logic is reconstructed to generate analysis and optimization instructions.
5. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the information analysis method based on artificial intelligence and big data as described in any one of claims 1 to 4.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the information analysis method based on artificial intelligence and big data as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Network fault diagnosis method and system based on 5G communication gateway
CN120343602A
Big data integration system based on artificial intelligence
CN120372283A
Intelligent power distribution network equipment state sensing and abnormity diagnosis system
CN120632742A