AI intelligent decision reasoning method and system based on machine learning

By introducing causal time-series windows and adaptive calibration mechanisms into the machine learning model, the problem of difficulty in identifying the cumulative risk of multiple small events in existing technologies is solved, enabling early identification and adaptive calibration of complex, slow-burning systemic risks, thus improving the accuracy and reliability of identification.

CN120952185AActive Publication Date: 2025-11-14SHENZHEN DIGITAL CONVERGENCE TECH CO LTD

Patent Information

Application Number
CN202511206961.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-11-14
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing machine learning models struggle to identify complex, slow-burning systemic risks that are formed by the gradual accumulation of multiple small, non-critical events over time, and they also suffer from overfitting risks and high computational costs.

Method used

By monitoring the data stream from the sensors, a causal time-series window is opened to generate causal correlation fingerprint vectors and causal noise feature vectors. These are then input into a machine learning model for decision-making and reasoning. When the system fails to identify risks, an adaptive calibration process is triggered to dynamically adjust the time-series window parameters. Finally, a post-decision arbitrator is used to verify logical consistency.

Benefits of technology

It enables early identification and adaptive calibration of systemic risks, improves the accuracy and reliability of identification, reduces the false alarm rate, and is suitable for edge computing devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952185A_ABST
    Figure CN120952185A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning intelligent decision, and discloses an AI intelligent decision reasoning method and system based on machine learning, and the method comprises the steps: monitoring a micro event which does not reach an alarm threshold value to open a time sequence window, coding the time sequence causal association between a plurality of sensors into a fingerprint vector which does not contain an original amplitude value in the window, according to the method, the limitation that traditional machine learning depends on static data snapshots is overcome, the hidden dynamic causal chain is directly coded into identifiable fingerprints, and therefore the recognition accuracy of the fingerprints is improved, and the recognition accuracy of the fingerprints is improved. The system can make accurate early warning in the germination period instead of the outbreak period of a systematic risk formed by accumulation of a plurality of tiny events, and the transformation from post-event diagnosis to beforehand foreseeing is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an AI intelligent decision-making reasoning method and system based on machine learning, belonging to the field of machine learning intelligent decision-making technology. Background Technology

[0002] Currently, in complex systems such as modern industrial control, power grid dispatching, or large-scale data center operation and maintenance, intelligent decision-making systems based on machine learning models have become a key technical support for ensuring their safe and stable operation. The common practice is to analyze real-time data collected by thousands of sensors to identify complex data patterns related to potential faults. The widespread application of this technical approach is based on an industry consensus: compared with traditional fixed threshold alarm methods, machine learning models, due to their powerful nonlinear fitting capabilities, can more effectively process high-dimensional and highly coupled data, thus demonstrating their advantages in identifying known fault modes.

[0003] However, as the scale and complexity of systems continue to increase, a more insidious and destructive risk pattern is becoming increasingly prominent. This pattern involves multiple independent, non-critical, minor anomalies that, over a relatively long timescale, collectively and gradually form a synergistic causal chain, ultimately leading to an irreversible decline in system function or sudden failure. When dealing with such complex, slow-burning systemic risks, current mainstream machine learning applications reveal their inherent fundamental constraints. The core issue is that these models are essentially based on static correlations from single sampling or short-term windows for learning and reasoning. While they can effectively identify strongly correlated events like high temperatures and malfunctions, they cannot effectively capture the hidden causal transmission path spanning hours or even days between three independent events, such as slight fan wear, minor fluctuations in ambient humidity, and a slow decline in coolant quality. This is because the representation of any single event in any single snapshot of data does not reach a level sufficient for the model to classify it as an anomaly.

[0004] To address this issue, a seemingly straightforward approach is to further increase the complexity of the model and the density of data collection, hoping to use more powerful computing power to automatically mine these weak signals from massive amounts of data. However, this approach does not change the essence of analyzing static correlations. Instead, it introduces more irrelevant background noise, leading to a sharp increase in the risk of overfitting during model training. The exponential increase in computational costs also makes its large-scale deployment on universally accessible edge computing devices impractical. Thus, existing technology falls into a fundamental contradiction. Specifically, existing technology has the following shortcomings: 1. The analysis object of machine learning models is an isolated, static snapshot of data, and its architecture itself lacks a mechanism for directly representing and learning the temporal causal relationships between data; 2. The system's perception of small, non-critical anomalies is limited to treating them as useless noise or interference to be filtered, failing to recognize their value as potential starting points of causal chains; 3. The solution to complex problems relies excessively on increasing the complexity of model algorithms and computational resources, while neglecting the possibility of information purification and value reconstruction from the source of data input. Therefore, how to construct a completely new processing method that enables machine learning models to break free from the limitations of static snapshot thinking and directly learn and identify causal fingerprints woven together by multiple micro-events over time, thereby providing accurate early warnings at the nascent stage of systemic risks rather than at their outbreak stage, is the technical problem that this invention aims to solve. Summary of the Invention

[0005] This invention provides an AI intelligent decision-making reasoning method and system based on machine learning. Its main purpose is to solve the problem that existing machine learning methods are unable to identify complex, slow-burning systemic risks formed by the gradual accumulation of multiple small non-critical events over time.

[0006] To achieve the above objectives, this invention provides an AI intelligent decision-making reasoning method based on machine learning, comprising the following steps:

[0007] Step a: Monitor the data stream of at least one sensor, and when a directional micro-event occurs that does not reach the alarm threshold, open a causal timing window with a defined duration;

[0008] Step b: Within the causal time window, based on the sequential timing and direction of change of multiple predetermined associated sensor responses, generate a causal correlation fingerprint vector that does not contain the original amplitude of the sensors, and simultaneously encode causal noise feature vectors from the unexpected fluctuations of the data streams of multiple predetermined associated sensors.

[0009] Step c: Input the causal association fingerprint vector and the causal noise feature vector into the machine learning model. The machine learning model corrects the confidence of the recognition result based on the causal association fingerprint vector according to the causal noise feature vector and outputs decision reasoning.

[0010] Step d: When the machine learning model fails to output decision reasoning within a defined time period and the system experiences a confirmed anomaly, an adaptive calibration process is triggered. This process automatically backtracks to the causal correlation fingerprint vector generated before the anomaly occurred and dynamically adjusts the duration of the causal time series window in step a until the machine learning model can identify the pattern corresponding to the confirmed anomaly.

[0011] Preferably, the method further includes: when the decision reasoning is a high-confidence risk warning, the decision reasoning is submitted to the post-decision arbitrator, and the post-decision arbitrator performs a logical consistency check on the decision reasoning according to the pre-established physical common sense rule library. If the decision reasoning contradicts the physical common sense rule library, the execution of the decision reasoning is blocked and an abnormal alarm is output.

[0012] Preferably, the causal correlation fingerprint vector is a discrete numerical vector, which contains encoding of at least one of the following relational features: whether the micro-events of the pre-determined correlated sensor occur within the causal time window; whether the change direction of the pre-determined correlated sensor micro-events is consistent with that of the micro-events in step a; and whether the time difference between the occurrence of the pre-determined correlated sensor micro-events and the micro-events in step a falls within a determined time interval.

[0013] Preferably, the causal noise feature vector includes encoding at least one of the following unexpected fluctuation features: the variance of sensor readings at the time of the micro-event in step a, and the jitter of the sensor's data stream transmission delay.

[0014] Preferably, the step of dynamically adjusting the duration of the causal time-series window in the adaptive calibration process specifically involves: increasing the duration by a determined step size ΔT to obtain a series of increasing time-series windows T. i And for each time window T i Regenerate the causal correlation fingerprint vector and the causal noise feature vector until a certain time window T is reached. i Below, the confidence level C of the decision reasoning output by the machine learning model. i Condition C is met for the first time i ≥C target Among them, T i =T0 + i·ΔT, where T0 is the initial duration, i is the increment number, and C target Let T be a defined confidence target threshold; and let the time window T that meets the conditions be... i The duration is fed back to step a as the updated duration.

[0015] Preferably, the micro-events in step a that do not reach the alarm threshold and have directionality are defined as: the sensor’s data stream shows a monotonically increasing or monotonically decreasing trend within a continuous sampling period, and the amplitude of the data stream change is within a defined non-alarm value range.

[0016] Preferably, the machine learning model is a shallow neural network.

[0017] Preferably, the machine learning model is a decision tree ensemble model, and the generation of causal association fingerprint vectors and causal noise feature vectors, as well as the inference of the machine learning model, are all completed on the edge computing device.

[0018] Preferably, the physical common sense rule base contains logical rules in the following form: if the decision reasoning is an overheat warning for device A, then check whether the speed of cooling fan B of device A is within the normal operating range; if not, then it is determined to be a logical contradiction.

[0019] An AI-powered intelligent decision-making and reasoning system based on machine learning, comprising:

[0020] The monitoring and triggering module is configured to monitor the data stream of at least one sensor and open a causal timing window of a defined duration when a micro-event that does not reach an alarm threshold and is directional occurs.

[0021] The dual-channel encoding module, connected to the monitoring and triggering module, is configured to generate a causal correlation fingerprint vector that does not contain the original amplitude of the sensors, based on the sequence and direction of change of the responses of multiple predetermined correlated sensors within a causal time window, and simultaneously encode causal noise feature vectors from the unexpected fluctuations of the data streams of multiple predetermined correlated sensors.

[0022] The collaborative decision-making module, connected to the dual-channel encoding module, is configured to receive causal correlation fingerprint vectors and causal noise feature vectors, and to correct the confidence of the recognition results based on the causal correlation fingerprint vectors through an internal machine learning model based on the causal noise feature vectors, so as to output decision reasoning.

[0023] The adaptive calibration module, connected to the collaborative decision-making module and the monitoring and triggering module, is configured to trigger the adaptive calibration process when the collaborative decision-making module fails to output decision reasoning within a defined time period and the system receives an external confirmation of an anomaly signal. This process traces back to the causal correlation fingerprint vector generated by the dual-channel encoding module before the anomaly occurred and dynamically adjusts the duration of the causal timing window to update the monitoring and triggering module with the adjusted duration.

[0024] Compared with the prior art, the beneficial effects of the present invention are:

[0025] 1. The method provided by this invention does not start from an abnormal state of the system, but rather takes a slight, non-alarm-level directional change of any sensor as the starting point for causal detection. A short time window is then opened, during which the system no longer records the specific values ​​of the sensor readings, but continuously listens for minute changes in other related sensors. Based on the relationship characteristics such as the order of occurrence, time interval, and consistency of the direction of change of these events, a low-dimensional, discrete fingerprint vector is generated that directly represents the multi-point temporal causal relationship. The machine learning model directly processes this fingerprint stream that has already encoded process information. Its identification target changes from a complex fit between the original data and the fault result to the direct identification of specific causal fingerprint combination patterns. This processing method makes the early identification of systemic risks formed by the gradual accumulation of multiple inconspicuous events over time a deterministic pattern matching process.

[0026] 2. In operation, the method of this invention also takes the logical contradiction event of the machine learning model failing to identify any preset risk pattern within a continuous time period while the system experiences a confirmed actual anomaly as an internal feedback signal. This signal triggers the system to automatically backtrack to the unidentified causal fingerprint vector and its original micro-event data stored before the anomaly occurred, and dynamically stretches and re-encodes the temporal window parameters in the fingerprint generation process with a preset step size until the machine learning model can successfully identify the fingerprint pattern corresponding to the actual anomaly. The system then updates the adjusted temporal window parameters to the baseline parameters under the current operating conditions. This process does not rely on any external intervention, enabling the system to have online adaptive parameter calibration capabilities when facing the slow drift of causal chain propagation time caused by equipment aging or environmental changes, thus maintaining the long-term effectiveness of decision reasoning.

[0027] 3. After the machine learning model outputs a high-confidence decision inference, the decision is not executed immediately. Instead, it is first submitted to a post-decision arbitrator based on a preset physical common sense rule base. This arbitrator is independent of the machine learning model and quickly verifies the rationality of the decision inference according to deterministic logical rules. For example, when the decision is an overheating warning for equipment, it verifies whether the speed of its cooling fan is within the normal range. Only when the decision inference does not contradict any physical common sense rules is it allowed to be executed. If there is a contradiction, the decision execution is blocked and a logical anomaly alarm is output. This step adds a final line of defense based on different principles to the entire intelligent decision-making process, placing the probabilistic pattern recognition capability of machine learning under the supervision of deterministic physical logic. Structurally, it avoids the risk of the model making a logically erroneous decision due to occasional data anomalies, thus improving the application security of the method in scenarios with high reliability requirements. Attached Figure Description

[0028] Figure 1 This is a schematic diagram of the overall process and adaptive calibration closed loop of the present invention;

[0029] Figure 2 This is a schematic diagram of the edge computing application architecture of the system of the present invention;

[0030] Figure 3 This is a schematic diagram illustrating the functions and signal interactions between the modules of the system of the present invention. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be described in detail below. It should be noted that the following embodiments are intended to explain the present invention and not to limit the scope of protection of the present invention.

[0032] This invention provides an AI intelligent decision-making reasoning method and system based on machine learning. Logically, this method and system can be divided into four collaborative modules or stages: First, a monitoring and triggering module continuously analyzes sensor data streams to identify micro-events that serve as the starting point for causal detection; then, a dual-channel encoding module generates, in parallel, a causal correlation fingerprint vector representing temporal correlation and a causal noise feature vector representing signal quality within a causal time-series window dynamically opened by the aforementioned modules; subsequently, a collaborative decision-making module receives the dual-channel vector and uses its internal machine learning model for decision-making reasoning and confidence correction; finally, when a specific logical contradiction arises between the system output and the actual situation, an adaptive calibration module is activated to dynamically adjust the causal time-series window parameters online.

[0033] In a typical engineering implementation, the system of the present invention can be deployed on an edge computing device. The monitoring and triggering module, the dual-channel encoding module, the collaborative decision-making module, and the adaptive calibration module are all independent software functional modules or service processes running on the processor of the edge computing device. The monitoring and triggering module acquires sensor data streams by reading the device's local data bus or network interface. When it detects a micro-event and opens a causal timing window, it sends the window opening signal and associated sensor identifiers to the dual-channel encoding module via inter-process communication. Upon receiving the signal, the dual-channel encoding module is activated and encodes the corresponding data streams according to the aforementioned identifiers within the window duration. The generated causal correlation fingerprint vector and causal noise feature vector are then transmitted to the collaborative decision-making module for inference via shared memory or a message queue. The output of the collaborative decision-making module can be used to output decisions externally and is continuously monitored by the adaptive calibration module. When the adaptive calibration module's own logic is triggered, it directly transmits the updated causal timing window duration parameter T to the monitoring and triggering module via function calls or control signaling. iThis allows for closed-loop control of system behavior.

[0034] In complex systems such as large-scale data center operations, a common technical challenge is that systemic risks often arise from the accumulation and propagation of multiple independent, minor anomalies that do not reach alarm thresholds over time. Traditional machine learning models based on static data snapshots struggle to capture such causal chains spanning time. To address this challenge, the monitoring and triggering module in this invention is configured to monitor the real-time data stream of at least one sensor in the system. A directional micro-event that does not reach an alarm threshold is defined as the starting point for causal detection. Specifically, this micro-event is deterministically defined as: the sensor's data stream exhibiting a monotonic upward trend within a continuous sampling period. A micro-event is defined as a temperature change that either rises or falls monotonically, with the overall amplitude falling within a defined non-alarm value range. For example, if the alarm threshold of a coolant temperature sensor is set to 85℃ and its normal fluctuation range is 60℃±2℃, then a change in temperature that monotonically rises within 10 consecutive sampling cycles with a total increase between 0.5℃ and 1.5℃ can be defined as a micro-event. Once such a micro-event is detected, the monitoring and triggering module immediately opens a causal timing window with a defined initial duration T0. The opening of this window signifies that the system has shifted from passive state monitoring to active causal relationship detection, thereby providing the necessary time reference for subsequently capturing a series of related responses triggered by the micro-event.

[0035] Within the opened causal time-series window, simply recording the raw amplitudes of each sensor results in excessively high data dimensionality, and the causal relationships contained within remain implicit. Machine learning models would require significant computational resources for fitting. Therefore, the dual-channel encoding module employs an information purification and feature extraction procedure. First, based on the temporal sequence and direction of change of multiple pre-determined correlated sensors within the window, the system generates a causal fingerprint vector that does not contain the raw amplitudes. This vector is a discrete numerical vector that encodes at least one relational feature, such as whether the micro-event of the correlated sensor occurred within the window, whether its direction of change is consistent with the initial micro-event, and the occurrence of both. Whether the time difference falls within a defined time interval can be represented by a three-dimensional vector, indicating that the related events have occurred in different directions but the time interval meets the requirements. This encoding method directly transforms dynamic process information into a static feature vector that can be efficiently identified by machine learning models. Secondly, to address the issue that unexpected fluctuations such as sensor background noise or data transmission delays may affect fingerprint reliability, the dual-channel encoding module synchronously encodes causal noise feature vectors from unexpected fluctuations in the data stream. These vectors include features such as the variance of sensor readings when the initial micro-event occurs, as well as data stream transmission delay jitter. Together, these two vectors constitute a complete representation of a causal detection event.

[0036] Subsequently, the collaborative decision-making module receives the aforementioned causal fingerprint vectors and causal noise feature vectors, and inputs them together into an internal machine learning model. This model can be a simplified shallow neural network or a decision tree ensemble model, suitable for inference on edge computing devices. The main task of the model is to identify specific combinations of causal fingerprint vectors. These patterns are learned during the offline training phase by studying the fingerprint streams corresponding to a large number of labeled historical risk events. In other words, when performing pattern recognition, the model corrects the confidence of the recognition result based on the synchronously input causal noise feature vectors. Specifically, when a potential risk fingerprint pattern is accompanied by a low-noise feature, such as a vector with low volatility variance and low transmission jitter, the model outputs a high-confidence decision inference. Conversely, if it is accompanied by a high-noise feature, the confidence of its output will be lowered accordingly. This mechanism allows the model's decision-making process to incorporate a quantitative evaluation of the quality of the input information, thereby improving the reliability of the decision.

[0037] Considering that equipment aging or environmental changes may cause a slow drift in the actual causal propagation time in the system, thereby invalidating the preset causal timing window duration T0, this invention sets up an online adaptive calibration process. The trigger condition for this process is: when the collaborative decision-making module does not output any decision reasoning within a defined time period, and the system simultaneously experiences an actual anomaly confirmed by external signals or manual intervention, this logical contradiction event will activate the adaptive calibration module. The module automatically backtracks to the causal association fingerprint vector generated by the dual-channel encoding module before the confirmed anomaly occurred, but which was not recognized by the model, and initiates a dynamic adjustment procedure for the causal timing window duration. This procedure increments the duration by a defined step size ΔT, resulting in a series of incremental timing windows T. i , among which, T i =T0 + i·ΔT, where i is the incrementing number for each new time window T i The system regenerates causal correlation fingerprint vectors and causal noise feature vectors from the original micro-event data, and inputs them into the machine learning model for re-identification until a certain time window T is reached. i Below, the confidence level C of the decision reasoning output by the model is... i Condition C is met for the first time i ≥C target , where C target Let T be a defined confidence target threshold; at this point, the system will select the time window T that meets the conditions. i The duration of the updated baseline duration is fed back to the monitoring and triggering module. This closed-loop calibration process does not require manual intervention and is used to maintain the long-term effectiveness of decision reasoning under dynamic operating conditions.

[0038] Finally, to address the potential risk of logical fallacies in machine learning models due to occasional data anomalies, this invention adds a security verification step based on different principles at the final stage of decision execution. Specifically, when the collaborative decision-making module outputs a high-confidence risk warning, the decision reasoning is not executed immediately. Instead, it is first submitted to a post-decision arbitrator. This arbitrator performs a logical consistency check on the decision reasoning based on a pre-established physical common sense rule base based on deterministic logic. For example, the rule base may contain logical rules in the following form: if the decision reasoning is an overheat warning for device A, then it must be verified whether the speed of the cooling fan B of device A is within the normal operating range. If the speed is not within the normal range, it is determined to be a logical contradiction. Only when the decision reasoning does not contradict any rule in the rule base is it allowed to be executed. If a contradiction exists, the execution of the decision is blocked and an anomaly alarm is output. This step places the probabilistic pattern recognition of machine learning under the supervision of deterministic physical logic, thereby improving the application security of the method in scenarios with high reliability requirements from the system architecture perspective.

[0039] Example 1: In a continuously operating large power grid dispatch center, there is a long-standing and unresolved operation and maintenance problem in a high-voltage substation it monitors. This problem stems from the combined effect of three independent factors—slight insulator contamination, gradual increase in circuit breaker contact resistance, and slow decline in transformer cooling system efficiency—which, at any single moment, do not constitute alarm conditions. This is a complex, slow-burning systemic risk. Existing monitoring systems or machine learning models based on data snapshots, because they analyze isolated static values, cannot identify the causal transmission paths between these minute changes that span hours or even days, thus lacking the ability to provide early warning for such risks.

[0040] During an actual operation, the monitoring and triggering module first detected a sustained unidirectional temperature rise of 0.6°C from the transformer coolant outlet temperature sensor reading. This change, as a micro-event that did not reach the alarm threshold, triggered the system to open a causal timing window with an initial duration of T0. Within this window, the dual-channel encoding module continuously monitored other preset associated sensors and successively detected a small directional increase from the high-voltage insulator surface leakage current sensor and a slight unidirectional rise from the main circuit breaker contact voltage drop sensor. The system then generated a specific causal fingerprint vector based on the sequence, time interval, and consistency of the direction of change of these three micro-events. This vector does not contain any original temperature or voltage amplitude but directly encodes this dynamic causal chain into a set of discrete features that can be directly recognized by machine learning models. At the same time, to evaluate the signal quality of the detected causal chain, the dual-channel encoding module also simultaneously extracted the reading fluctuation variance and data transmission delay jitter during the micro-event period from the data streams of the three sensors and encoded them into a causal noise feature vector.

[0041] Subsequently, these two vectors are jointly input into the machine learning model of the collaborative decision-making module. The collaborative working method here is as follows: the causal correlation fingerprint vector provides the model with the pattern signal to be identified, while the causal noise feature vector provides the model with a quantitative basis for judging the credibility of the signal. This approach enables the system to effectively distinguish between real causal transmission events and pseudo-patterns formed by background noise, thereby maintaining high early warning sensitivity while avoiding the problem of increased false alarm rate that may be caused by introducing more non-critical events as analysis objects. Furthermore, this processing method transforms the complex problem in traditional machine learning, which requires manually mining weak temporal correlations from massive amounts of raw data, into a pattern matching problem for low-dimensional discrete fingerprints that have already encoded process information, making real-time and efficient inference possible on edge computing devices.

[0042] Ultimately, based on the input causal correlation fingerprint vector and the high signal quality evidence provided by the causal noise feature vector, the machine learning model matched a dangerous fingerprint combination pattern highly correlated with historical complex faults and output a high-confidence risk warning after confidence level correction. Based on this, the dispatch center arranged maintenance in advance and intervened after confirming the initial state of the anomaly, thus avoiding a possible chain failure event of main power grid equipment caused by the gradual accumulation of multiple potential risks. The realization of this process does not rely on increasing the complexity of the model algorithm, but on information purification and value reconstruction at the source of data input, transforming the analysis of static correlations into the direct identification of dynamic causal chains.

[0043] Example 2: To objectively verify the effectiveness of the method of the present invention in identifying systemic risks accumulated from multiple minor anomalies, a hardware-in-the-loop simulation test platform for simulating an industrial circulating cooling pump system was built. The data acquisition system of this platform has a sampling rate of not less than 1 kHz, wherein the temperature sensor has a measurement resolution of 0.01℃ and the vibration sensor has a frequency resolution of 0.5 Hz, in order to capture sufficiently fine data changes. The fault data stream injected in the experiment is generated based on a fluid thermodynamics and mechanical dynamics model. This model aims to reproduce the progressive thermal runaway process of the pump body caused by the combined effect of slow deterioration of bearing lubrication and slight contamination of coolant.

[0044] Two groups were set up for comparison in the experiment. The control group used a machine learning model based on the isolated forest algorithm to directly detect anomalies in the raw high-dimensional data streams collected by all sensors. Its warning threshold was set to an anomaly score of 0.6. The experimental group deployed the technical solution of this invention in its entirety. Before the experiment, a key parameter in the experimental group, namely the initial duration T0 of the causal time series window, needed to be set. The setting of this parameter needs to balance the completeness of capturing effective causal chains with the accuracy of avoiding interference from irrelevant events. Its value is based on the physical characteristics of the system. It should be greater than the shortest time for the transmission of key physical quantities in the system, and less than the average time interval of random noise events under fault-free conditions. For the thermodynamic characteristics of this experimental platform, the duration T0 was set to 5.0s.

[0045] After the experiment started, the simulation platform began to inject a pre-set composite slow-heating fault scenario lasting 60 minutes. At the 29th minute of the experiment, the monitoring and triggering module of the experimental group first captured a micro-event that met the preset conditions on a bearing temperature sensor. Starting from this point, within the opened causal timing window, related micro-events were successively captured on the motor vibration spectrum and coolant turbidity sensors. The system then began to generate causal correlation fingerprint vectors and causal noise feature vectors, which were analyzed in real time by the collaborative decision-making module. At the 40th minute of the experiment, the machine learning model of the experimental group had matched a dangerous fingerprint combination pattern that was highly correlated with historical faults and output a risk warning with a confidence level of 96%. At the same time, the anomaly score output by the isolated forest model of the control group was only 0.21, far below its warning threshold of 0.6.

[0046] As the simulation continued, the anomaly score output by the control group model only reached 0.62 for the first time at the 55th minute, triggering a risk warning. This result indicates that the effective warning time of the experimental group was approximately 15 minutes earlier than that of the control group. The reason for this difference is that in the raw data processed by the control group model, the magnitude of changes in a single dimension was insufficient to constitute a significant statistical anomaly in the early stages of the fault. However, the experimental group, through a data preprocessing step of causal fingerprinting, transformed weak but temporally correlated changes in multiple dimensions into a highly discriminative feature vector. The object learned and identified by the machine learning model was no longer the absolute value of data points, but rather the causal pattern itself that characterizes the fault mechanism. The experimental results show that when dealing with complex, slow-burning risk scenarios that gradually accumulate from multiple small anomalies, this invention, by encoding dynamic causal relationships into static fingerprint features, can identify risks in the early development stage of the fault rather than in the near-failure stage. Its technical effectiveness and engineering feasibility have been verified.

[0047] Example 3: This example combines Figures 1 to 3 This section describes an AI-based intelligent decision-making and reasoning method and system based on machine learning, such as... Figure 1 As shown, the process begins with sensor data streams acquired from systems such as industrial control or data centers. The data streams first enter the monitoring and triggering module, which monitors micro-events and opens a causal timing window accordingly. After the window opens, the dual-channel encoding module is activated, generating two outputs: a causal correlation fingerprint vector and a causal noise feature vector. These two vectors are fed into the collaborative decision-making module, where the machine learning model performs inference and confidence correction. If its output is a high-confidence warning, the warning is submitted to the post-decision arbitrator. The arbitrator performs logical verification based on a physical commonsense rule base. If the verification passes, a decision inference is output; if the verification fails, an abnormal alarm indicating a logical contradiction is issued and execution is prevented. On the other hand, if the collaborative decision-making module does not output a decision within a specific period and receives an external confirmation of the abnormality, the adaptive calibration module is triggered. This module dynamically adjusts the duration of the causal timing window online and feeds the updated duration parameters back to the monitoring and triggering module, thus forming a closed-loop adaptive adjustment circuit.

[0048] like Figure 2As shown, the core of this system is deployed on an edge computing device, which integrates a monitoring and triggering module, a dual-channel encoding module, a collaborative decision-making module, an adaptive calibration module, and a post-decision arbitrator. It also includes a local machine learning model database. The system's input comes from an external sensor array, which provides the edge computing device with real-time local bus or network data streams. At the same time, the human-machine interface (HMI) can provide the system with manually confirmed abnormal signals to trigger the adaptive calibration process. The system's output is risk warnings and decisions, which are transmitted to the upper-level industrial control system or data center to achieve intelligent monitoring and intervention of the entire system.

[0049] like Figure 3 As shown, the data stream from sensor group sensor 1 to sensor N is continuously monitored and triggered by the micro-event detection module. Upon detecting a micro-event, a timing window is opened, and the micro-event trigger signal is sent to the dual-channel encoding module. This module then generates a causal correlation fingerprint vector and a causal noise feature vector, and inputs these two vectors into the machine learning model in the collaborative decision-making module for inference. If the model outputs a high-confidence warning, the warning first undergoes logical verification and rule validation by the post-arbitrator. If the validation passes, a decision is executed; if there is a contradiction, an anomaly alarm is output. In addition, when the collaborative decision-making module fails to identify an anomaly, this state, together with the confirmation signal from the outside, serves as the input to the adaptive calibration module. This module performs parameter backtracking and timing window adjustment, and feeds back the updated parameters to the monitoring and triggering module, completing the dynamic optimization of system parameters.

[0050] Example 4: When the method and system of the present invention are first deployed on a new large industrial centrifugal air compressor unit that lacks historical fault data, the key algorithm parameters and models of the system need to be initially calibrated in order to enable it to operate effectively under specific working conditions. This process aims to provide objective basis for the definition of micro-events, the generation of causal fingerprints, and the training of machine learning models.

[0051] To perform this calibration, an offline parameter setting procedure is first executed. The input to this procedure is all sensor data collected by the compressor unit during continuous operation for at least 24 hours under normal conditions. The first step is to calculate the statistical baseline noise of the data stream under steady-state conditions for each sensor. Specifically, this involves calculating the standard deviation σ of the first-order difference sequence of its readings. Furthermore, the micro-event that triggers the causal timing window is defined as: when the sensor reading exhibits a monotonically changing change for at least 5 consecutive sampling points, and the total change amplitude ΔV within those 5 sampling points satisfies the condition 3σ < ΔV < 0.2·V. alarm At that time, V alarmThe procedure provides a reproducible quantitative method for identifying micro-events by setting a preset alarm threshold for the sensor. The second step involves processing the same 24-hour baseline data based on the calibrated micro-event definition to determine the initial duration T0 of the causal time window. This process aims to identify non-fault-related causal events caused by routine operations under normal conditions. It calculates the distribution of the time interval Δt between all event pairs triggered by micro-event A and subsequently followed by micro-event B from another related sensor by statistically analyzing all observed Δt values. max As the initial duration T0 of the causal time series window, if Δt is obtained through statistical calculation... max The initial value of T0 is set to 4.2s, and the adjustment step size ΔT in the adaptive calibration process is set to a fixed proportion of 10% of T0, i.e., 0.42s. The third step involves initializing and training the machine learning model. Given the lack of real fault samples, a semi-supervised learning approach is used to generate the training dataset. This involves using non-risk fingerprint vectors generated from a large amount of normal operation data as negative samples, and combining them with fingerprint vectors generated from a few typical fault causal chains manually labeled based on historical equipment maintenance logs as positive samples. Using this initially constructed training set, a decision tree ensemble model is trained. The goal of the model training is to learn the decision boundary that can distinguish between these two types of samples. After training, the model is tested on a reserved validation set to determine the initial confidence target threshold C. target The standard for this threshold is that the recall rate of the model for known positive samples is not less than 99.5%. This step completes the cold start process of the model.

[0052] After the aforementioned offline calibration and model training procedures, the system possesses the initial monitoring and inference capabilities for this new compressor unit. All its core parameters and models are generated driven by the device's own operating data, thus transforming the parameter and model selection process into an engineering implementation process that follows specific steps and is based on specific data input. This provides a benchmark starting point for adaptive calibration during subsequent online operation of the system and ensures the reproducibility of the technical solution in different application scenarios.

[0053] Example 5: Before the system is deployed in a specific application scenario, an offline rule extraction and verification procedure needs to be executed to establish the physical common sense rule base required for the post-decision arbitrator. This procedure first systematically sorts out the combinations of equipment states with deterministic causal relationships based on the design documents and process flow diagrams of the controlled system. Then, these combinations are transformed into logical rules in the form of if-only. Before all generated rules are loaded into the rule base, they also need to be checked for logical consistency through formal verification tools to identify and eliminate potential conflicts between rules, thereby improving the judgment accuracy of the arbitrator when running online.

[0054] After the system is put into online operation, in order to handle novel causal fingerprint vectors that did not appear during the training phase, the system has built a new online recording and correlation analysis mechanism for new patterns. When the dual-channel encoding module generates a fingerprint vector that does not match any known risk or security pattern, the vector, its corresponding timestamp, and the original micro-event data will be automatically recorded in a database to be analyzed. This process does not trigger any immediate alarms and is only used as an observation record of a new pattern. If the system receives an externally confirmed abnormal signal within a subsequent preset time period, the adaptive calibration module will backtrack and query this database to be analyzed while performing its calibration process. It will also temporally correlate the new pattern fingerprint vector recorded before the anomaly occurred with the anomaly event and mark it as a potential risk pattern to be reviewed for use in subsequent model iteration training.

[0055] Example 6: In order to apply the method and system of the present invention to a new large industrial centrifugal air compressor unit that lacks historical fault data, a standardized pre-deployment procedure needs to be executed before the system is put into formal operation. This procedure aims to establish a baseline model and parameters for the specific equipment that can characterize its healthy operating status.

[0056] The procedure first determines the specific dimensional composition of the causal fingerprint vector for the compressor unit. This step combines data-driven analysis and engineering knowledge base verification. Specifically, it uses 24-hour baseline data collected during the offline calibration phase of the system to calculate and analyze the information flow intensity between the data streams of each sensor through transfer entropy calculation. It then selects sensor pairs that are statistically in the top 10% in terms of information transfer volume, forming a candidate causal relationship list. At the same time, based on the failure mode and effect analysis (FMEA) engineering document of the compressor unit, it extracts sensor combinations with physical causal relationships. Finally, the parts in the aforementioned candidate causal relationship list that intersect with the content of the FMEA document are determined as the core relationships that need to be encoded in the causal fingerprint vector for this specific scenario, thus providing a dual basis of data and mechanism for its dimensional composition.

[0057] Subsequently, the procedure configures a defined signal input path for the trigger condition of system confirmation anomaly during the adaptive calibration process. The system integrates a dedicated external event input channel at the hardware interface level. This channel is configured to receive signals from two sources: one is a Boolean confirmation signal manually input by maintenance personnel through the human-machine interface after physically inspecting the equipment and confirming the fault on-site; the other is a switch signal output from a traditional alarm system based on a fixed high threshold integrated within the controlled system itself. When the adaptive calibration module receives any valid signal from this channel, it uses the timestamp of that signal as a reference to start the backtracking calibration procedure. This design provides the necessary, clearly sourced external feedback for the entire adaptive learning loop to close.

[0058] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0059] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. An AI-based intelligent decision-making and reasoning method based on machine learning, characterized in that, Includes the following steps: Step a: Monitor the data stream of at least one sensor, and when a directional micro-event occurs that does not reach the alarm threshold, open a causal timing window with a defined duration; Step b: Within the causal time window, based on the sequential timing and direction of change of multiple predetermined associated sensor responses, generate a causal correlation fingerprint vector that does not contain the original amplitude of the sensors, and simultaneously encode causal noise feature vectors from the unexpected fluctuations of the data streams of multiple predetermined associated sensors. Step c: Input the causal association fingerprint vector and the causal noise feature vector into the machine learning model. The machine learning model corrects the confidence of the recognition result based on the causal association fingerprint vector according to the causal noise feature vector and outputs decision reasoning. Step d: When the machine learning model fails to output decision reasoning within a defined time period and the system experiences a confirmed anomaly, an adaptive calibration process is triggered. This process automatically backtracks to the causal correlation fingerprint vector generated before the anomaly occurred and dynamically adjusts the duration of the causal time series window in step a until the machine learning model can identify the pattern corresponding to the confirmed anomaly.

2. The AI ​​intelligent decision-making and reasoning method based on machine learning according to claim 1, characterized in that, The method further includes: when the decision reasoning is a high-confidence risk warning, the decision reasoning is submitted to the post-decision arbitrator, which performs a logical consistency check on the decision reasoning based on a pre-established physical common sense rule library. If the decision reasoning contradicts the physical common sense rule library, the execution of the decision reasoning is blocked and an abnormal alarm is output.

3. The AI ​​intelligent decision-making and reasoning method based on machine learning according to claim 1, characterized in that, The causal association fingerprint vector is a discrete numerical vector, which contains encoding of at least one of the following relationship features: whether the micro-events of the pre-determined associated sensor occur within the causal time window; whether the change direction of the pre-determined associated sensor micro-events is consistent with that of the micro-events in step a; and whether the time difference between the occurrence of the pre-determined associated sensor micro-events and the micro-events in step a falls within a determined time interval.

4. The AI ​​intelligent decision-making and reasoning method based on machine learning according to claim 1, characterized in that, The causal noise feature vector contains encoding of at least one of the following unexpected fluctuation features: the variance of sensor readings when the micro-event occurs in step a, and the jitter of sensor data stream transmission delay.

5. The AI ​​intelligent decision-making and reasoning method based on machine learning according to claim 1, characterized in that, The specific steps for dynamically adjusting the duration of the causal time-series window in the adaptive calibration process are as follows: the duration is increased by a determined step size ΔT to obtain a series of increasing time-series windows T. i And for each time window T i Regenerate the causal correlation fingerprint vector and the causal noise feature vector until a certain time window T is reached. i Below, the confidence level C of the decision reasoning output by the machine learning model. i Condition C is met for the first time i ≥C target Among them, T i =T0 + i·ΔT, where T0 is the initial duration, i is the increment number, and C target Let T be a defined confidence target threshold; and let the time window T that meets the conditions be... i The duration is fed back to step a as the updated duration.

6. The AI ​​intelligent decision-making and reasoning method based on machine learning according to claim 1, characterized in that, In step a, a micro-event that does not reach the alarm threshold and has directionality is defined as: the sensor's data stream exhibits a monotonically increasing or monotonically decreasing trend within a continuous sampling period, and the amplitude of the data stream change is within a defined non-alarm value range.

7. The AI ​​intelligent decision-making and reasoning method based on machine learning according to claim 1, characterized in that, The machine learning model is a shallow neural network.

8. The AI ​​intelligent decision-making and reasoning method based on machine learning according to claim 1, characterized in that, The machine learning model is a decision tree ensemble model. The generation of causal association fingerprint vectors and causal noise feature vectors, as well as the inference of the machine learning model, are all completed on edge computing devices.

9. The AI ​​intelligent decision-making and reasoning method based on machine learning according to claim 2, characterized in that, The physics common sense rule base contains logical rules in the following form: If the decision reasoning is an overheat warning for device A, then check whether the speed of cooling fan B of device A is within the normal operating range. If not, it is determined to be a logical contradiction.

10. An AI intelligent decision-making and reasoning system based on machine learning, characterized in that, include: The monitoring and triggering module is configured to monitor the data stream of at least one sensor and open a causal timing window of a defined duration when a micro-event that does not reach an alarm threshold and is directional occurs. The dual-channel encoding module, connected to the monitoring and triggering module, is configured to generate a causal correlation fingerprint vector that does not contain the original amplitude of the sensors, based on the sequence and direction of change of the responses of multiple predetermined correlated sensors within a causal time window, and simultaneously encode causal noise feature vectors from the unexpected fluctuations of the data streams of multiple predetermined correlated sensors. The collaborative decision-making module, connected to the dual-channel encoding module, is configured to receive causal correlation fingerprint vectors and causal noise feature vectors, and to correct the confidence of the recognition results based on the causal correlation fingerprint vectors through an internal machine learning model based on the causal noise feature vectors, so as to output decision reasoning. The adaptive calibration module, connected to the collaborative decision-making module and the monitoring and triggering module, is configured to trigger the adaptive calibration process when the collaborative decision-making module fails to output decision reasoning within a defined time period and the system receives an external confirmation of an anomaly signal. This process traces back to the causal correlation fingerprint vector generated by the dual-channel encoding module before the anomaly occurred and dynamically adjusts the duration of the causal timing window to update the monitoring and triggering module with the adjusted duration.

Citation Information

Patent Citations

  • Electromechanical system fault diagnosis system based on deep learning

    CN120163069A

  • Power plant intelligent maintenance method and system based on multi-modal dynamic graph learning

    CN120494806A

Cited By

  • Broaching tool machining data acquisition method and system

    CN121131866A

  • A method and system for collecting machining data of a broaching tool

    CN121131866B