Multi-modal equipment integrated management system and method based on intelligent AI
By introducing intelligent AI technology into the device integrated management system, real-time causal modeling and self-supervised anomaly detection of multimodal data, generating multi-strategy execution paths, and optimizing management through cognitive feedback paths, the problems of lack of explicitness in causal relationship modeling and lack of causal chain support for strategy generation are solved, and the efficiency and accuracy of fault tracing, abnormal detection and policy decision-making are improved.
Patent Information
- Application Number
- CN202510705294.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The causal modeling of existing equipment integrated management systems lacks explicitness and interpretability, resulting in fault tracing relying on manual experience, abnormal data is not isolated, strategy generation lacks causal chain support, control logic fragmentation, strategy generalization ability is weak, and it is unable to cope with complex working conditions or cascading failures.
The multimodal device integrated management system based on intelligent AI is adopted, including the autonomous perception module for real-time causal modeling, the liquid feature factory module extracts time series features, the path decision module generates multi-strategy execution paths, and the execution control module realizes collaborative control and cognitive reflection coordination module for optimization management.
Through real-time causal modeling and self-supervised abnormality detection, the efficiency and accuracy of fault traceability and abnormality detection are improved; through causal enhancement of embedding mechanisms, the coverage and adaptability of strategic decisions are improved; through cognitive feedback paths, the long-term stability and intelligence of the system are enhanced, and accurate perception, efficient decision-making, precise execution and intelligent optimization management of multimodal devices are achieved.
Smart Images

Figure CN120217271A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of device integration, and more specifically, to a multi-modal device integration management system and method based on intelligent AI. Background Art
[0002] A patent with the publication number CN115098156A discloses a network modality management system and a management method. The system includes a multi-modal network integration development environment and a multi-modal network distributed compilation and deployment environment; the multi-modal network integration development environment further includes a network modality deployment file packaging tool for packaging network modality source files and corresponding configuration files into network modality deployment files; the multi-modal network distributed compilation and deployment environment includes a multi-modal network modality program package manager deployed on a controller server and a network node device network modality program package manager deployed on a network node device. The present invention realizes the unified management of network modalities in a multi-modal network, automatically distributes deployment files to all network node devices in the network, and uniformly coordinates the compilation and deployment work of source code files on each network node device and different target forwarding modules on the device, thereby significantly improving the management efficiency of network modalities.
[0003] Existing device integration management systems and methods mainly have the following problems: The causal relationship modeling lacks explicitness and interpretability. Traditional device integration management systems generally rely on black-box data correlation modeling methods, which can only identify "what is related" but cannot answer "why it is related"; and the causal relationship is often hidden in complex model parameters and lacks a clear structural representation. This limitation of the lack of explicit causal representation directly leads to the following series of chain problems: Fault tracing relies on manual experience and lacks an automated analysis path: Since the system cannot clearly express and trace the causal path between events structurally, once an anomaly or fault occurs, it can only rely on the experience of operation and maintenance personnel to locate problems, making it difficult to achieve automatic fault reasoning and efficient location based on the causal chain, reducing the system response efficiency and accuracy; Abnormal data is not isolated, interfering with the evolution of the causal structure: In the absence of causal explicitness, the system also lacks a structural recognition and isolation mechanism for "abnormal behaviors". When abnormal data enters the system, it often directly participates in model training or graph structure update, resulting in incorrect learning and edge weight distortion, and then generating false causal accumulation in long-term operation, seriously affecting the system stability and judgment accuracy.
[0004] The lack of causal chain support in strategy generation and the fragmentation of control logic: Based on the absence of causal structure, the decision-making logic of the system often only stays at "data correlation-driven", ignoring the actual links of "functional collaboration" and "causal dependence" between devices. For example, when the root cause of the motor speed drop is insufficient lubrication of the gearbox, the system may only adjust the motor parameters and cannot trigger a joint strategy involving the lubrication system, resulting in fragmented strategies and limited execution effects. Weak strategy generalization ability and inability to handle complex working conditions or cascading failures: A strategy system lacking causal support can often only handle predefined standard scenarios and lacks effective strategy coverage for complex working conditions such as multi-modal asynchronous anomalies and chain reactions between devices, requiring a large amount of manual intervention. Moreover, the system cannot perform causal closed-loop evaluation from execution feedback, resulting in the lack of dynamic adaptability of strategy adjustment and being easily misled by abnormal data or sensing noise.
[0005] In view of this, the present invention proposes a multi-modal device integrated management system and method based on intelligent AI to solve the above problems. Summary of the Invention
[0006] To overcome the above defects of the prior art and to achieve the above object, the present invention provides the following technical solution: A multi-modal device integrated management system based on intelligent AI, comprising: An autonomous perception module that performs real-time causal modeling on multi-modal data using embedded causal analysis; combines a self-supervised anomaly detection mechanism to output a device causal graph and a data stream with confidence labels. A liquid feature factory module that, based on the device causal graph and the data stream with confidence labels, uses a gated temporal convolutional network to extract time series features, reconstructs the feature path through a dynamic feature pooling mechanism, and combines a sparse regularization method to compress and select high-dimensional features, thereby generating a feature tensor. A path decision module that generates multi-strategy execution paths based on the feature tensor and the device causal graph; constructs a strategy behavior tree according to the multi-strategy execution paths through a multi-agent curriculum learning mechanism; uses a distributed consensus mechanism to verify the consistency of the execution paths of the strategy behavior tree and outputs a multi-path strategy instruction set. An execution control module for parsing the multi-path strategy instruction set to generate a preliminary group execution action plan; correcting the preliminary group execution action plan according to the data stream with confidence labels to generate execution feedback data and execution status data, and realizing collaborative control between each execution node by combining a group topology optimization algorithm. A cognitive reflection coordination module for constructing a cognitive feedback path including multi-modal state tracking and execution evaluation according to the execution feedback data and the execution status data, and performing integrated optimization management of multi-modal devices through an adaptive feedback mechanism.
[0007] Preferably, the method for real-time causal modeling includes: Collect heterogeneous data from each sensing terminal. According to the local clock synchronization signal, align the heterogeneous data to a unified sampling time axis in an interpolation manner. At the same time, use the spatial nearest neighbor interpolation method to complete the missing values, and perform noise filtering on the completed heterogeneous data, thereby completing the unified spatio-temporal standardization of the heterogeneous data and obtaining multimodal data; Preset the time window length, and collect multimodal data through a sliding time window mechanism to form a multimodal data set; the multimodal data set consists of multimodal data collected at all time points within the initial time window; construct a full-variable lag prediction model, use the multimodal data set as the input, perform Granger causality test, and obtain a set of directed causal edges; Input the multimodal data set into the structure inference layer of the variational autoencoder, encode the distribution characteristics of each multimodal data, aim to minimize the reconstruction error and the KL divergence of the prior structure, calculate and obtain the connection weights of the directed causal edges, and obtain a set of edge weights; use the data contained in the multimodal data set as nodes, combine the set of directed causal edges and the set of edge weights, and output the device causal graph.
[0008] Preferably, the method for the self-supervised anomaly detection mechanism includes: Receive the multimodal data collected through the sliding window mechanism, and perform zero-mean unit-variance standardization to eliminate the difference in numerical dimensions, thereby obtaining the standardized multimodal data; introduce two pre-text tasks to learn normal behavior patterns in an unlabeled environment; the two pre-text tasks are reconstruction pre-text and prediction pre-text respectively; Reconstruct the standardized multimodal data through the reconstruction pre-text and calculate the reconstruction error. Use the reconstruction error as the loss function, and minimize the loss function through backpropagation and gradient descent to obtain the reconstruction error score, and automatically learn the normal cooperation pattern between multimodal data; Through the prediction pre-text, based on the standardized multimodal data of the last time steps, predict the multimodal data of the next time step; use the prediction error as the loss function and minimize it through backpropagation and gradient descent to obtain the prediction error score, and capture the temporal dynamic characteristics between multimodal data; Take 50% weights of the calculated reconstruction error score and prediction error score respectively, and obtain the joint anomaly score through weighted fusion; preset the joint anomaly score threshold, and define a confidence function, calculate and obtain the confidence according to the joint anomaly score; mark the obtained confidence on the corresponding multimodal data, thereby obtaining a data stream with confidence labels; Preset a confidence threshold, and mark the confidence below the preset confidence threshold as an abnormal confidence; for the multimodal data with abnormal confidence, trigger a freezing operation to temporarily suppress the update of the edge weights of the data in the device causal graph; for each freezing operation, dynamically calculate the freezing duration according to the current abnormal confidence.
[0009] Preferably, the method for obtaining the feature tensor includes: Perform multi-scale convolution processing on the data stream with confidence labels through a gated temporal convolutional network, introduce a gating mechanism to control the information flow, and extract time series features; construct a feature fusion path based on the device causal graph and confidence labels, and aggregate the time series features of different modalities in the device causal graph in a graph path-guided manner to obtain a feature fusion vector; apply the L1 norm constraint to the feature fusion vector to eliminate redundant features, and perform feature dimension compression through sparse principal component analysis to generate a feature tensor.
[0010] Preferably, the method for obtaining the multi-strategy execution path includes: Through a causal enhancement embedding mechanism, jointly map the feature tensor and the device causal graph to the policy space to construct a joint representation vector; based on the joint representation vector, use a multi-agent curriculum learning mechanism to identify all policy paths under the current policy goal and generate a candidate policy space; For each candidate policy path in the candidate policy space, use a graph neural network to evaluate its execution utility, and the execution utility is calculated based on the execution success probability of the candidate policy path; preset an execution success probability threshold, evaluate the execution utility of all candidate policy paths, and filter out the candidate policy paths greater than the preset execution success probability threshold to form the final multi-strategy execution path.
[0011] Preferably, the method for obtaining the multi-path policy instruction set includes: Map the multi-strategy execution path to a structured policy behavior tree, define the joint representation vector as the root node of the policy behavior tree, and guided by each policy path, recursively expand the state transition chain on the basis of the root node to convert each policy path into a complete branch of the policy behavior tree; introduce a fault-tolerant distributed consensus protocol to broadcast the policy execution path of the policy behavior tree to all participating devices. When each device receives the policy execution instruction, perform consistency verification using Byzantine fault tolerance to check whether there are conflicts in the multi-strategy execution path, retain the policy execution paths without conflicts to form a consistency path set, and convert it into a structured instruction set to obtain the multi-path policy instruction set.
[0012] Preferably, the method for correcting the action plan for the preliminary population includes: Perform action parsing according to the multi-path policy instruction set, disassemble the policy into device control instructions, select the execution order based on the device causal graph, bind the device control instructions, execution order and corresponding devices, and generate a preliminary group execution action plan; Call the data stream with confidence labels to correct the preliminary group execution action plan, generate correction instructions based on the dynamic correction table, and make an adaptive adjustment to the control instructions of the corresponding devices.
[0013] Preferably, the method for realizing collaborative control between execution nodes includes: An execution control module, configured to parse the multi-path policy instruction set and generate a preliminary group execution action plan; correct the preliminary group execution action plan according to the data stream with confidence labels, generate execution feedback data and execution status data, and realize collaborative control between execution nodes in combination with the group topology optimization algorithm; Make an adaptive adjustment to the control instructions of the corresponding devices. After completing the correction of the preliminary group execution action plan, generate execution feedback data, and at the same time collect the execution status data during the execution of the correction instructions; Based on the execution status data and the device causal graph, construct a device-instruction bipartite graph, bind all control instructions to the corresponding devices, and use the group topology optimization algorithm to minimize the global execution cost as the optimization goal, perform resource sharing and information synchronization between devices, and then realize the system control between execution nodes.
[0014] Preferably, the method for integrated optimization management of the multimodal device includes: The execution feedback data includes execution success / failure flags, action deviation values, feedback confidence trajectories, and fault trigger flags and reasons; the execution status data includes node active status, task completion rate indicators, average confidence and volatility indicators, and inter-node response delay matrices; fuse the execution feedback data and execution status data to form a multimodal tracking data set; Use a time-aware neural network to perform state evolution modeling on the fused multimodal tracking data set to generate a state tracking model covering the evolution path of the device operating state; According to the state tracking model, construct an execution evaluation index system including execution stability, policy matching degree, and multimodal consistency, output the device operation deviation analysis result and the behavior consistency evaluation result, and construct a cognitive feedback path based on the device operation deviation analysis result and the behavior consistency evaluation result; Dynamically adjust the device collaboration mechanism based on the cognitive feedback path to perform integrated optimization management of the multimodal device.
[0015] A multimodal device integrated management method based on intelligent AI includes: S1. Perform real-time causal modeling on multimodal data using embedded causal analysis; combine a self-supervised anomaly detection mechanism to output a device causal graph and a data stream with confidence labels; S2. Based on the device causal graph and the data stream with confidence labels, use a gated temporal convolutional network to extract time series features, reconstruct the feature path through a dynamic feature pooling mechanism, combine a sparse regularization method to compress and select high-dimensional features, and then generate a feature tensor; S3. Generate multi-strategy execution paths based on the feature tensor and the device causal graph; through a multi-agent curriculum learning mechanism, construct a policy behavior tree according to the multi-strategy execution paths; use a distributed consensus mechanism to verify the consistency of the execution paths of the policy behavior tree, and output a multi-path policy instruction set; S4. Parse the multi-path policy instruction set to generate a preliminary group execution action plan; correct the preliminary group execution action plan according to the data stream with confidence labels to generate execution feedback data and execution status data, and combine a group topology optimization algorithm to achieve collaborative control between execution nodes; S5. Construct a cognitive feedback path including multi-modal state tracking and execution evaluation according to the execution feedback data and execution status data, and perform integrated optimization management of multi-modal devices through an adaptive feedback mechanism.
[0016] Compared with the prior art, the present invention has the following beneficial effects: Accurately reflect the causal relationship between the internal operation mechanism of the device and external interference through real-time causal modeling, which helps to deeply understand the operation principle of the device, provides a reliable basis for subsequent decision-making and control. Compared with traditional implicit correlation analysis methods, the causal relationship is highly interpretable, providing a clear logical path for fault tracing and prediction; introducing a self-supervised anomaly detection task does not require manual annotation, automatically learns the intra-modal statistical features, temporal dynamic laws, cross-modal correlation relationships and spatio-temporal consistency of the normal operation of the device, solves the pain point of traditional supervised methods relying on a large number of abnormal samples, and reduces the data annotation cost; at the same time, combined with the multi-modal data-triggered freezing operation of the anomaly confidence, it can effectively prevent the interference of abnormal data on the structure of the device causal graph, and ensure the stability and reliability of the system; Fully consider the causal relationship and policy objectives between devices through a causal enhancement embedding mechanism, generate multi-strategy execution paths including causal dependencies. Compared with single-strategy decision-making, the decision coverage rate in complex scenarios is significantly improved, and the policy adaptability is enhanced; through the policy behavior tree, gradually transition from basic policies to complex collaborative policies, improve the policy exploration efficiency, and avoid the cold start problem of traditional reinforcement learning; Construct a cognitive feedback path including multi-modal state tracking and execution evaluation, perform integrated optimization management of multi-modal devices through an adaptive feedback mechanism, enhance the long-term stability and intelligence of the system, can monitor and evaluate the operation status of the device in real time, and dynamically adjust the device collaboration mechanism according to the evaluation results, improving the integrated management level and operation efficiency of the device; By integrating cutting-edge technologies such as causal reasoning, self-supervised learning, and group topology optimization, an intelligent framework for the entire "perception - feature - decision - execution - feedback" link is constructed, forming the collaborative advantage of interdisciplinary technologies, and achieving precise perception, efficient decision-making, accurate execution, and intelligent optimization management of multi-modal devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic structural diagram of a multi-modal device integrated management system based on intelligent AI according to the present invention; Figure 2 It is a schematic flowchart of a method for integrating and managing multi-modal devices based on intelligent AI according to the present invention; Figure 3 It is a flowchart of a method for a self-supervised anomaly detection mechanism provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0019] Embodiment 1 Please refer to Figure 1 and Figure 3 As shown, this Embodiment 1 further describes a multi-modal device integrated management system based on intelligent AI proposed by the present invention, including: With the rapid development of artificial intelligence technology, a single modality has been difficult to meet the requirements of complex tasks. Multi-modal AI breaks through the limitations of a single modality by integrating various data sources such as text, images, audio, and video, significantly enhancing the expressiveness and generalization ability of the model. In the fields of industrial automation, smart home, medical diagnosis, etc., multi-modal data fusion technology has become the core driving force for realizing intelligent management of devices. For example, an autonomous driving system realizes precise perception of the road environment by fusing camera, radar, and lidar data; a medical diagnosis system improves the accuracy of disease detection by integrating medical images and medical record texts. However, there are significant technical bottlenecks in existing multi-modal device integrated management systems in aspects such as data fusion, causal modeling, intelligent decision-making, and real-time control, which are specifically manifested as follows: In industrial scenarios, devices usually come from different manufacturers, use proprietary communication protocols, and have significant differences in data formats (such as binary sensor signals and unstructured text logs). As a result, when integrating systems, it is necessary to rely on customized gateways or middleware for protocol conversion, which has poor compatibility and high extension costs. For example, traditional solutions lack a unified spatio-temporal alignment mechanism for cross-modal fusion of vibration signals (high-frequency time-series data) and device manuals (unstructured text), often resulting in the loss of key causal information due to inconsistent sampling rates and clock deviations.
[0020] Traditional device management systems mostly rely on correlation analysis or black-box models to mine data associations, and can only identify "data co-occurrence", making it difficult to accurately reflect the causal relationship between the internal operating mechanism of the device and external interference (such as the causal direction between abnormal motor current and bearing wear); this limitation will lead to difficulties in fault tracing (lack of a clear causal logic path, making it difficult to quickly locate the root cause of the fault), low efficiency of anomaly detection (relying on manually labeled anomaly samples, with high data annotation costs and being easily affected by subjective factors), insufficient decision coverage (single-strategy decision-making cannot adapt to complex scenarios, with poor strategy adaptability), and poor system stability; For example, when the temperature of a certain device is abnormal, the traditional system can only trigger an alarm, but it is difficult to locate whether the root cause is excessive upstream vibration or a sudden change in load. At the same time, there are no effective modeling methods for the temporal dynamic laws (such as periodic fluctuations) and cross-modal coupling relationships (such as the coordinated change of rotational speed and vibration) of multi-modal data, resulting in a high false alarm rate for anomaly detection (such as misjudgment triggered by a single-modal threshold), and fault tracing relying on manual experience, with low efficiency.
[0021] Existing multi-modal device integration management systems have significant deficiencies in compatibility, real-time performance, intelligence, and security, and are difficult to meet the complex requirements of industries such as industry, healthcare, and energy. To solve this problem, the present invention aims to provide an efficient and reliable solution for cross-industry device management through systematic innovations in protocol standardization, real-time fusion of multi-modal data, intelligent decision-making, and security and privacy protection.
[0022] Therefore, in order to effectively solve the above problems, the present invention proposes a multi-modal device integration management system based on intelligent AI, including: An autonomous perception module that performs real-time causal modeling on multi-modal data using embedded causal analysis; combined with a self-supervised anomaly detection mechanism, it outputs a device causal graph and a data stream with confidence labels; A liquid feature factory module that, based on the device causal graph and the data stream with confidence labels, uses a gated temporal convolutional network to extract time-series features, reconstructs the feature path through a dynamic feature melting pool mechanism, and combines a sparse regularization method to compress and select high-dimensional features, thereby generating a feature tensor; The path decision-making module generates multi-strategy execution paths based on feature tensors and device causal graphs; constructs a policy behavior tree according to the multi-strategy execution paths through a multi-agent curriculum learning mechanism; uses a distributed consensus mechanism to verify the consistency of the execution paths of the policy behavior tree, and outputs a multi-path policy instruction set; The execution control module is used to parse the multi-path policy instruction set, generate a preliminary group execution action plan; correct the preliminary group execution action plan according to the data stream with confidence labels, generate execution feedback data and execution status data, and realize the cooperative control between execution nodes by combining the group topology optimization algorithm; The cognitive reflection coordination module is used to construct a cognitive feedback path including multi-modal state tracking and execution evaluation according to the execution feedback data and execution status data, and perform integrated optimization management of multi-modal devices through an adaptive feedback mechanism.
[0023] The method of real-time causal modeling includes: Collect heterogeneous data from each sensing terminal, such as images, temperatures, accelerations, vibrations, voltages, currents, etc. generated during the operation of each device. According to the local clock synchronization signal, align the heterogeneous data to a unified sampling time axis by interpolation, and at the same time use the spatial nearest neighbor interpolation method to fill in the missing values, and perform noise filtering on the filled heterogeneous data, so as to complete the unified spatio-temporal standardization of the heterogeneous data and obtain multi-modal data; Preset the time window length, collect multi-modal data through the sliding time window mechanism to form a multi-modal data set; the multi-modal data set is composed of multi-modal data collected at all time points within the initial time window; construct a full-variable lag prediction model, use the multi-modal data set as input, perform Granger causality test, and obtain a set of directed causal edges; Input the multi-modal data set into the structure inference layer of the variational autoencoder, encode the distribution characteristics of each multi-modal data, and calculate and obtain the connection weights of the directed causal edges with the goal of minimizing the reconstruction error and the KL divergence of the prior structure to obtain an edge weight set; use the data included in the multi-modal data set as nodes, combine the directed causal edge set and the edge weight set, and output the device causal graph, that is, a directed graph reflecting the causal relationship between the internal operation mechanism of the device and external interference.
[0024] The method of the self-supervised anomaly detection mechanism includes: Receive the multi-modal data collected through the sliding window mechanism, and perform zero-mean unit-variance standardization to eliminate the difference in numerical dimensions, so as to obtain the standardized multi-modal data; introduce two pre-text tasks to learn normal behavior patterns in an unlabeled environment; the two pre-text tasks are reconstruction pre-text and prediction pre-text respectively; Learning the normal behavior pattern means automatically inducing the data characteristics and temporal patterns of the "device in normal working state" from historical data without any manual annotation (i.e., without "abnormal" or "normal" sample labels), so as to construct an internal cognitive model of the "normal" state of the device, specifically including intra-modal statistical features, temporal dynamic patterns, cross-modal correlation relationships, and spatio-temporal consistency; Intra-modal statistical features: Model the statistical characteristics such as amplitude distribution, variance, and spectral components of heterogeneous data (such as temperature, current, vibration signals) obtained by each sensing terminal during normal operation; for example, the average value of the temperature sensor stabilizes within a certain range during normal periods, and high-frequency mutations do not occur in the vibration signal during normal periods; Temporal dynamic patterns: Capture the evolution patterns of multi-modal data over time, such as periodic fluctuations, trend changes, or sudden jump patterns; by predicting whether the data value at the next time step conforms to the historical trend, to measure whether the current observation deviates from the "normal" dynamics; Cross-modal correlation relationships: Under normal conditions, there are often stable coupling relationships between different modalities; for example, when the motor speed increases, the vibration and current will increase synchronously; when the load is stable, the temperature and current have corresponding co-variations; after these typical co-patterns are learned by the system, once "one modality is normal and the other modality deviates", it can be recognized as abnormal; Spatio-temporal consistency: The data of the device under normal operation has coherent and smooth changes in space (between different devices or sensing points) and time; sudden jumps and isolated abnormal points often mean faults or abnormal events; Reconstruct the pre-text to reconstruct the standardized multi-modal data and calculate the reconstruction error. Use the reconstruction error as the loss function, and minimize the loss function through backpropagation and gradient descent to ensure that the reconstructed data can accurately restore the input as much as possible, obtain the reconstruction error score, and automatically learn the normal co-patterns between multi-modal data; Through prediction pre-text Based on the most recent time steps of the standardized multi-modal data, predict the multi-modal data at the next time step; use the prediction error as the loss function and minimize it through backpropagation and gradient descent to obtain the prediction error score, thereby strengthening the learning ability of multi-modal temporal dynamics and capturing the temporal dynamic characteristics between multi-modal data; How the pre-text task helps with "normal" mode learning:
[0025] Reconstruction pre-text: Let the system learn to "restore" the input data stream. If the system can accurately reconstruct, it is very likely that compression and decoding are performed according to the normal mode; if the reconstruction error is too large, it means that the input data deviates from the normal feature distribution.
[0026] Prediction pre-text: Let the system "predict" the next step given the data of the past few steps. If the prediction error is small, it indicates that the time series follows the normal dynamic pattern learned by the system; if the error is significant, it implies a deviation from the current pattern.
[0027] Through these two tasks, without any abnormal labels, an internal characterization of "normal behavior" can be spontaneously made, including both the statistical characteristics of single modalities and the coupled temporal relationships between multi-modalities. Any input data that deviates from this "learned normal pattern" will generate a high error during the reconstruction or prediction phase and thus be judged as a potential anomaly.
[0028] Take 50% weights for the calculated reconstruction error scores and prediction error scores respectively, and obtain the joint anomaly score through weighted fusion; preset the joint anomaly score threshold and define the confidence function, and calculate the confidence according to the joint anomaly score; the confidence function is , where is the timestamp; is the control factor, used to control the mapping sensitivity between the joint anomaly score and the confidence; is the joint anomaly score; is the preset joint anomaly score threshold; mark the obtained confidence on the corresponding multi-modal data, and then obtain the data stream with confidence labels; Preset the confidence threshold, and record the confidence lower than the preset confidence threshold as the abnormal confidence; for the multi-modal data with abnormal confidence, trigger the freezing operation to temporarily suppress the update of the edge weights of the data in the device causal graph and avoid abnormal data introducing false causal relationships; for each freezing operation, dynamically calculate the freezing duration according to the current abnormal confidence ; where is the preset maximum freezing duration; is the preset minimum freezing duration; is the preset confidence threshold; that is to say: if the abnormal confidence is closer to the preset confidence threshold, then the freezing duration will be close to the minimum value; otherwise, the freezing duration will approach the preset maximum freezing duration; this freezing operation can adaptively adjust the freezing time according to the severity of the anomaly to prevent abnormal data from continuously interfering with the structure of the device causal graph.
[0029] For example: In the operating environment of an intelligent wind farm, device nodes such as fan blades, main shafts, gearboxes, and pitch systems collect multi-modal data streams (including vibration, temperature, current, voltage, wind speed, etc.) through sensors. The system needs to construct a device operation causal graph based on this data to learn the coupling and fault propagation paths between devices. Some abnormal data (such as sensor inaccuracy, electromagnetic interference, instantaneous jumps, etc.) may cause misleading enhancement or incorrect connection in the learning of edge weights in the causal graph, resulting in misjudgment or the generation of false root cause chains.
[0030] The system calculates, for each moment , the reconstruction error score (such as based on the output of the variational autoencoder VAE); the prediction error score (such as based on the time series prediction network); and sets equal weights (50%) for fusion to obtain the joint anomaly score : If the preset joint anomaly score threshold is 0.4 and the control factor is -6, then the confidence ; If the preset confidence threshold is 0.7, since 0.658 < 0.7, it is determined as abnormal confidence data; the preset freezing duration range: the preset shortest freezing duration is 5 seconds, and the preset longest freezing duration is 60 seconds; dynamically calculate the freezing duration seconds; the update of the edge weight of the device node corresponding to this multi-modal data will be frozen within 8.3 seconds to avoid short-term anomaly interference in the learning process.
[0031] The method for obtaining the feature tensor includes: Perform multi-scale convolution processing on the data stream with confidence labels through a gated temporal convolutional network, and introduce a gating mechanism to control the information flow to enhance the response to key change trends and temporal dependence relationships, and extract time series features; construct a feature fusion path based on the device causal graph and confidence labels, and aggregate different modal time series features in the device causal graph in a graph path-guided manner to obtain a feature fusion vector; apply the L1 norm constraint to the feature fusion vector to eliminate redundant features, and perform feature dimension compression through sparse principal component analysis to generate a low-dimensional, highly expressive feature tensor.
[0032] The method for obtaining the multi-strategy execution path includes: Through the causal enhancement embedding mechanism, jointly map the feature tensor and the device causal graph to the policy space to construct a joint representation vector; based on the joint representation vector, use the multi-agent curriculum learning mechanism to identify all policy paths under the current policy goal to generate a candidate policy space; the candidate policy space expresses the action dependence and causal advancement logic between devices in the form of a policy path graph; For each candidate policy path in the candidate policy space, use a graph neural network to evaluate its execution utility, which is calculated based on the execution success probability of the candidate policy path; preset an execution success probability threshold, evaluate the execution utility of all candidate policy paths, and filter out the candidate policy paths with an execution success probability greater than the preset threshold to form the final multi-policy execution path.
[0033] The method for obtaining the multi-path policy instruction set includes: Map the multi-policy execution path to a structured policy behavior tree, define a joint representation vector as the root node of the policy behavior tree, and guided by each policy path, recursively expand the state transition chain based on the root node to convert each policy path into a complete branch of the policy behavior tree; introduce a fault-tolerant distributed consensus protocol (such as PBFT, Raft), broadcast the policy execution path of the policy behavior tree to all participating devices, and when each device receives the policy execution instruction, perform consistency verification using Byzantine fault tolerance to check whether there are conflicts in the multi-policy execution path (such as conflicts in the order of events, violation of time constraints), retain the policy execution paths without conflicts to form a consistency path set, and convert it into a structured instruction set, thereby obtaining the multi-path policy instruction set; the multi-path policy instruction set includes the following fields: instruction number, policy path identifier, instruction content, execution condition, instruction confidence, and expected feedback constraint.
[0034] The method for correcting the initial group execution action plan includes: Perform action parsing according to the multi-path policy instruction set, decompose the policy into device control instructions, select the execution order based on the device causal graph, bind the device control instructions, execution order, and corresponding devices to generate the initial group execution action plan; call the data stream with confidence labels to correct the initial group execution action plan, generate correction instructions based on the dynamic correction table, and make an adaptive adjustment to the control instructions of the corresponding devices.
[0035] For example: Table 1 Dynamic Correction Table Confidence level category Correction action Example Too low (confidence < 0.3) Equipment task suspension, task migration The signal quality of a certain fan equipment is poor, and the execution is suspended Medium (0.3 ≤ confidence level ≤ 0.6) Parameter adjustment, dynamic delay The current feedback of the heater is abnormal, and the heating action is delayed for 5 seconds High (confidence level > 0.6) Enable bypass or fault-tolerant redundancy The main power supply is overloaded, switch to the standby power supply The method for realizing collaborative control between each execution node includes: An execution control module, which is used to parse the multi-path policy instruction set to generate an initial group execution action plan; correct the initial group execution action plan according to the data stream with confidence labels, generate execution feedback data and execution status data, and realize collaborative control between each execution node in combination with the group topology optimization algorithm; Make an adaptive adjustment to the control instructions for the corresponding device. After completing the correction of the action plan for the preliminary group, generate execution feedback data, and at the same time collect the execution status data during the execution of the correction instructions; Based on the execution status data and the device causal graph, construct a device-instruction bipartite graph, bind all control instructions to the corresponding devices, and use the group topology optimization algorithm to minimize the global execution cost as the optimization goal, and perform resource sharing and information synchronization between devices, so as to realize the system control between each execution node.
[0036] The method for the integrated optimization management of multimodal devices includes: The execution feedback data includes an execution success / failure flag (indicating whether it is completed according to the plan), an action deviation value (the difference between the actual execution and the original plan, such as a 3-second extension in time consumption), a feedback confidence trajectory (the confidence change trajectory during the execution process), and a fault trigger flag and reason (if it stops due to an abnormality, record the fault trigger event and the abnormality description of the sensing terminal); The execution status data is used to describe the current overall operating state of the group system, including the node active state (which nodes are online, offline, degraded), the task completion rate index (the completion ratio of the current action plan), the average confidence and volatility index (the stability of the current overall perception of the device), and the inter-node response delay matrix (the delay of each device to the control signal); Integrate the execution feedback data and the execution status data to form a multimodal tracking data set; Use a time-aware neural network to perform state evolution modeling on the integrated multimodal tracking data set to generate a state tracking model covering the evolution path of the device operating state. According to the state tracking model, construct an execution evaluation index system including execution stability, policy matching degree, and multimodal consistency, output the analysis results of device operation deviation and the evaluation results of behavior consistency, and construct a cognitive feedback path based on the analysis results of device operation deviation and the evaluation results of behavior consistency. Dynamically adjust the device cooperation mechanism based on the cognitive feedback path to perform the integrated optimization management of multimodal devices; Dynamically adjusting the device cooperation mechanism includes adjusting the task allocation ratio between devices, updating the communication topology, and adjusting the device working mode, etc.
[0037] The preset joint anomaly score threshold is set by the staff, and the average value of multiple joint anomaly scores is taken as the preset joint anomaly score threshold; Similarly, set the preset confidence threshold, the preset execution success probability threshold, and the preset correction deviation threshold.
[0038] In this embodiment, the real-time causal modeling accurately reflects the causal relationship between the internal operation mechanism of the device and external interference, which helps to deeply understand the operation principle of the device and provides a reliable basis for subsequent decision-making and control. Compared with the traditional implicit correlation analysis method, the causal relationship is highly interpretable, providing a clear logical path for fault tracing and prediction; the introduction of the self-supervised anomaly detection task does not require manual annotation, and automatically learns the intra-modal statistical features, temporal dynamic laws, cross-modal correlation relationships and spatio-temporal consistency of the normal operation of the device, solving the pain point of the traditional supervised method relying on a large number of abnormal samples and reducing the data annotation cost; at the same time, combined with the multi-modal data-triggered freezing operation of the anomaly confidence, it can effectively prevent the interference of abnormal data on the causal graph structure of the device and ensure the stability and reliability of the system; The causal enhancement embedding mechanism fully considers the causal relationship and policy objectives between devices, generates multi-policy execution paths containing causal dependencies, and significantly improves the decision coverage rate and policy adaptability in complex scenarios compared with single-policy decision-making; through the policy behavior tree, it gradually transitions from basic policies to complex collaborative policies, improving the policy exploration efficiency and avoiding the cold start problem of traditional reinforcement learning; Construct a cognitive feedback path including multi-modal state tracking and execution evaluation, and perform integrated optimization management of multi-modal devices through an adaptive feedback mechanism, enhancing the long-term stability and intelligence of the system. It can real-time monitor and evaluate the operation state of the device, and dynamically adjust the device collaboration mechanism according to the evaluation results, improving the integrated management level and operation efficiency of the device; By integrating cutting-edge technologies such as causal reasoning, self-supervised learning, and group topology optimization, a full-link intelligent framework of "perception - feature - decision - execution - feedback" is constructed, forming the collaborative advantage of interdisciplinary technologies, and realizing the precise perception, efficient decision-making, accurate execution, and intelligent optimization management of multi-modal devices.
[0039] Embodiment 2 Please refer to Figure 2 As shown, for the parts not described in detail in this embodiment, refer to the description content of Embodiment 1. A multi-modal device integrated management method based on intelligent AI is provided, including: S1. Perform real-time causal modeling on multi-modal data using embedded causal analysis; combine the self-supervised anomaly detection mechanism to output the device causal graph and the data stream with confidence labels; S2. Based on the device causal graph and the data stream with confidence labels, use a gated temporal convolutional network to extract time series features, reconstruct the feature path through a dynamic feature pooling mechanism, and combine a sparse regularization method to compress and select high-dimensional features, thereby generating a feature tensor; S3. Generate multi-strategy execution paths based on the feature tensor and the device causal graph; construct a policy behavior tree according to the multi-strategy execution paths through a multi-agent curriculum learning mechanism; use a distributed consensus mechanism to verify the consistency of the execution paths of the policy behavior tree, and output a multi-path policy instruction set; S4. Parse the multi-path policy instruction set to generate a preliminary group execution action plan; correct the preliminary group execution action plan according to the data stream with confidence labels to generate execution feedback data and execution status data, and realize the collaborative control between execution nodes by combining the group topology optimization algorithm; S5. Construct a cognitive feedback path including multi-modal state tracking and execution evaluation according to the execution feedback data and the execution status data, and perform integrated optimization management of multi-modal devices through an adaptive feedback mechanism.
[0040] Since the electronic device introduced in this embodiment is the electronic device adopted in a multi-modal device integrated management system and method based on intelligent AI in the embodiments of the present application, based on the multi-modal device integrated management system and method introduced in the embodiments of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the implementation of how this electronic device realizes the method in the embodiments of the present application will not be described in detail here. As long as those skilled in the art implement the electronic device adopted in a multi-modal device integrated management system and method based on intelligent AI in the embodiments of the present application, it belongs to the scope protected by the present application.
[0041] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula closest to the real situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.
[0042] The above is only the preferred implementation manner of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technical users in the technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A multi-modal device integrated management system based on intelligent AI, characterized in that, Including: An autonomous perception module that performs real-time causal modeling on multimodal data using embedded causal analysis; Combined with a self-supervised anomaly detection mechanism, it outputs a device causal graph and a data stream with confidence labels; A liquid feature factory module that, based on the device causal graph and the data stream with confidence labels, uses a gated temporal convolutional network to extract time series features, reconstructs the feature path through a dynamic feature pooling mechanism, combines a sparse regularization method to compress and select high-dimensional features, and then generates a feature tensor; A path decision module that generates multi-strategy execution paths based on the feature tensor and the device causal graph; through a multi-agent curriculum learning mechanism, constructs a policy behavior tree according to the multi-strategy execution paths; Uses a distributed consensus mechanism to verify the consistency of the execution paths of the policy behavior tree and outputs a multi-path policy instruction set; An execution control module for parsing the multi-path policy instruction set to generate a preliminary group execution action plan; correcting the preliminary group execution action plan according to the data stream with confidence labels to generate execution feedback data and execution status data, and realizing cooperative control between execution nodes by combining a group topology optimization algorithm; A cognitive reflection coordination module for constructing a cognitive feedback path including multimodal state tracking and execution evaluation according to the execution feedback data and the execution status data, and performing integrated optimization management of multimodal devices through an adaptive feedback mechanism.
2. The multimodal device integrated management system based on intelligent AI according to claim 1, wherein The method of the real-time causal modeling includes: Collect heterogeneous data from each sensing terminal, align the heterogeneous data to a unified sampling time axis according to the local clock synchronization signal in an interpolation manner, simultaneously use a spatial nearest neighbor interpolation method to fill in missing values, and perform noise filtering on the filled heterogeneous data, thereby completing the unified spatio-temporal standardization of the heterogeneous data and obtaining multimodal data; Preset the time window length, collect multimodal data through a sliding time window mechanism to form a multimodal data set; the multimodal data set is composed of multimodal data collected at all time points within the initial time window; construct a lag prediction model of all variables, use the multimodal data set as input, perform Granger causality test, and obtain a set of directed causal edges; Input the multimodal data set into the structure inference layer of the variational autoencoder, encode the distribution characteristics of each multimodal data, aim at minimizing the reconstruction error and the KL divergence of the prior structure, calculate and obtain the connection weights of the directed causal edges to obtain an edge weight set; use the data included in the multimodal data set as nodes, combine the set of directed causal edges and the edge weight set, and output a device causal graph.
3. The multimodal device integrated management system based on intelligent AI according to claim 2, wherein The method of the self-supervised anomaly detection mechanism includes: Receive the multimodal data collected through the sliding window mechanism and perform zero-mean unit-variance standardization to eliminate the numerical dimension difference, thereby obtaining the standardized multimodal data; introduce two pre-text tasks to learn normal behavior patterns in an unlabeled environment; the two pre-text tasks are reconstruction pre-text and prediction pre-text respectively; Reconstruct the standardized multimodal data by reconstructing the pre-text and calculate the reconstruction error. Use the reconstruction error as the loss function, and minimize the loss function through backpropagation and gradient descent to obtain the reconstruction error score, and automatically learn the normal collaboration pattern among multimodal data; Predict the multi-modal data of the next time step based on the normalized multi-modal data of the recent time steps; use the prediction error as the loss function and minimize it through backpropagation and gradient descent to obtain the prediction error score, capturing the temporal dynamic features between multi-modal data; Take 50% weights of the calculated reconstruction error score and prediction error score respectively, and obtain the joint anomaly score through weighted fusion; preset the joint anomaly score threshold, and define the confidence function, and calculate the confidence according to the joint anomaly score; mark the obtained confidence on the corresponding multimodal data, and then obtain the data stream with confidence labels; Preset the confidence threshold, and record the confidence lower than the preset confidence threshold as the abnormal confidence; for the multimodal data with abnormal confidence, trigger the freezing operation to temporarily suppress the update of the edge weight of the data in the device causal graph; for each freezing operation, dynamically calculate the freezing duration according to the current abnormal confidence.
4. An integrated management system for multimodal devices based on intelligent AI according to claim 3, characterized in that, The method for obtaining the feature tensor includes: Perform multi-scale convolution processing on the data stream with confidence labels through a gated temporal convolutional network, introduce a gating mechanism to control the information flow, and extract time series features; construct a feature fusion path based on the device causal graph and confidence labels, and aggregate the time series features of different modalities in the device causal graph in a graph path-guided manner to obtain a feature fusion vector; apply the L1 norm constraint to the feature fusion vector to eliminate redundant features, and compress the feature dimension through sparse principal component analysis, and then generate the feature tensor.
5. An integrated management system for multi-modal devices based on intelligent AI according to claim 4, characterized in that, The method for obtaining the multi-strategy execution path includes: Through the causal enhancement embedding mechanism, jointly map the feature tensor and the device causal graph to the policy space to construct a joint representation vector; based on the joint representation vector, use the multi-agent curriculum learning mechanism to identify all policy paths under the current policy goal and generate a candidate policy space; For each candidate policy path in the candidate policy space, use a graph neural network to evaluate its execution utility, and the execution utility is calculated based on the execution success probability of the candidate policy path; preset the execution success probability threshold, evaluate the execution utility of all candidate policy paths, and filter out the candidate policy paths greater than the preset execution success probability threshold to form the final multi-strategy execution path.
6. The integrated management system for multi-modal devices based on intelligent AI according to claim 5, wherein The method for obtaining the multi-path policy instruction set includes: Map the multi-strategy execution path to a structured policy behavior tree, define the joint representation vector as the root node of the policy behavior tree, and guided by each policy path, recursively expand the state transition chain on the basis of the root node to convert each policy path into a complete branch of the policy behavior tree; introduce a fault-tolerant distributed consensus protocol to broadcast the policy execution path of the policy behavior tree to all participating devices. When each device receives the policy execution instruction, use Byzantine fault tolerance for consistency verification to check whether there are conflicts in the multi-strategy execution path, retain the policy execution paths without conflicts to form a consistency path set, and convert it into a structured instruction set, and then obtain the multi-path policy instruction set.
7. An integrated management system for multimodal devices based on intelligent AI according to claim 6, characterized in that, The method for correcting the initial population execution action plan includes: Perform action parsing according to the multi-path policy instruction set, decompose the policy into device control instructions, select the execution order based on the device causality graph, bind the device control instructions, execution order, and corresponding devices, and generate a preliminary group execution action plan; Call the data stream with confidence labels to correct the preliminary group execution action plan, generate correction instructions based on the dynamic correction table, and make adaptive adjustments to the control instructions of the corresponding devices.
8. An integrated management system for multi-modal devices based on intelligent AI according to claim 7, characterized in that, The method for realizing collaborative control between each execution node includes: An execution control module, which is used to parse the multi-path policy instruction set and generate a preliminary group execution action plan; Correct the preliminary group execution action plan according to the data stream with confidence labels, generate execution feedback data and execution status data, and realize collaborative control between each execution node by combining the group topology optimization algorithm; Make adaptive adjustments to the control instructions of the corresponding devices. After completing the correction of the preliminary group execution action plan, generate execution feedback data, and at the same time collect the execution status data during the execution of the correction instructions; Based on the execution status data and the device causality graph, construct a device-instruction bipartite graph, bind all control instructions to the corresponding devices, and use the group topology optimization algorithm to minimize the global execution cost as the optimization goal to share resources and synchronize information between devices, thereby realizing system control between each execution node.
9. An integrated management system for multimodal devices based on intelligent AI according to claim 8, characterized in that, The method for integrated optimization management of the multi-modal device includes: The execution feedback data includes execution success / failure flags, action deviation values, feedback confidence trajectories, and fault trigger flags and reasons; The execution status data includes node active status, task completion rate indicators, average confidence and volatility indicators, and inter-node response delay matrices; Integrate the execution feedback data and execution status data to form a multi-modal tracking data set; Use a time-aware neural network to model the state evolution of the integrated multi-modal tracking data set and generate a state tracking model covering the evolution path of the device operating state; According to the state tracking model, construct an execution evaluation index system including execution stability, policy matching degree, and multi-modal consistency, output the device operation deviation analysis result and behavior consistency evaluation result, and construct a cognitive feedback path based on the device operation deviation analysis result and behavior consistency evaluation result; Dynamically adjust the device collaboration mechanism based on the cognitive feedback path to perform integrated optimization management of the multi-modal device.
10. A multi-modal device integrated management method based on intelligent AI, which is used to implement a multi-modal device integrated management system according to any one of claims 1 to 9, characterized in that, It includes: S1. Perform real-time causal modeling on multi-modal data using embedded causal analysis; Combine the self-supervised anomaly detection mechanism to output the device causality graph and the data stream with confidence labels; S2. Based on the device causality graph and the data stream with confidence labels, use a gated temporal convolutional network to extract time series features, reconstruct the feature path through the dynamic feature pooling mechanism, and combine the sparse regularization method to compress and select high-dimensional features, thereby generating a feature tensor; S3. Generate multi-strategy execution paths based on the feature tensor and the device causality graph; Through the multi-agent curriculum learning mechanism, construct a policy behavior tree according to the multi-strategy execution path; Use the distributed consensus mechanism to verify the consistency of the execution paths of the policy behavior tree and output the multi-path policy instruction set; S4. It is used to parse the multi-path policy instruction set to generate a preliminary group execution action plan; correct the preliminary group execution action plan according to the data stream with confidence labels to generate execution feedback data and execution status data, and realize the collaborative control between execution nodes by combining the group topology optimization algorithm; S5. It is used to construct a cognitive feedback path including multi-modal state tracking and execution evaluation according to the execution feedback data and execution status data, and perform integrated optimization management of multi-modal devices through an adaptive feedback mechanism.
Citation Information
Patent Citations
Network mode management system and management method
CN115098156A
Heterogeneous graph learning system and method for generating causal relationship of urban road network flow
CN116738725A
Unmanned aerial vehicle-based river hydrological sampling inspection method and system
CN119151387A
Commercial credit evaluation and supervision method based on multi-modal coevolution algorithm
CN119250963A
Industrial Internet of Things equipment intelligent identification method and device based on deep learning
CN119691548A
Cited By
Big data platform storage data isolation method in SaaS mode
CN120492215A
A data storage isolation method for a big data platform under a SaaS mode
CN120492215B
Intelligent network communication and edge calculation optimization method for electric vehicle
CN120935229A
Communication chip signal interference prediction and optimization system based on artificial intelligence
CN121173404A
Water-saving agent interaction control method and system based on multi-modal fusion
CN121598326A