Multi-modal industrial operation and maintenance system and method fusing hierarchical large model architecture and adaptive reinforcement learning, and application of multi-modal industrial operation and maintenance system and method
By integrating a multimodal industrial operation and maintenance system with a hierarchical large-modal model architecture and adaptive reinforcement learning, the limitations of the middle-level model architecture of the industrial operation and maintenance system, the problems of multimodal data fusion obstacles and dynamic resource allocation are solved, efficient and intelligent multi-level decision-making and resource allocation are achieved, and the intelligent level and operation and maintenance efficiency of the system are improved.
Patent Information
- Application Number
- CN202510400387.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-01
AI Technical Summary
There are limitations of layer model architecture, obstacles to multimodal data fusion and inefficiency in dynamic resource allocation in existing industrial operation and maintenance systems, resulting in high decision-making errors, large delays, failed data associations and unacceptable resource allocation to drift in response to operating conditions.
The fusion hierarchical large-model architecture is adopted, including expert layer, system layer and total engineering layer modules, combined with adaptive reinforcement learning and Bayesian collaborative network, multimodal data fusion and dynamic resource allocation are realized, and the model weights are optimized through the BERT-Whitening algorithm, LSTM feature extraction, GRU network and Bayesian network, and the inter-hierarchical API call weights are dynamically adjusted.
It improves the intelligence level of industrial operation and maintenance systems, optimizes multi-level decision response efficiency, reduces operation and maintenance costs, enhances system reliability and decision-making accuracy, and realizes real-time strategy optimization.
Smart Images

Figure CN120355134A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of industrial automation and artificial intelligence, and particularly to a multimodal industrial operation and maintenance system and method, and an application that integrate a hierarchical large model architecture and adaptive reinforcement learning. Background Art
[0002] Currently, with the rapid development of artificial intelligence technology, the intelligent upgrade of industrial operation and maintenance systems faces the following technical bottlenecks: 1. Limitations of the layer model architecture: Existing systems (such as Siemens SIMATIC PCS 7) adopt a single AI model in the cloud centralized or edge distributed manner, resulting in decision-making fragmentation between levels.
[0003] The device-side model (such as the lightweight LSTM network deployed on the PLC) is limited by computing resources (typical computing power <1 TOPS) and can only perform simple tasks such as threshold warning, with a decision error rate as high as 18.7% (based on the IEC 62443 standard test set); Although the cloud large model (such as Alibaba Cloud ET Industrial Brain) has complex decision-making capabilities, the decision-making delay reaches 800 - 1200 ms (ISO 23837 standard test results), which cannot meet the real-time control requirements specified in IEC 61131-3 (delay requirement ≤ 300 ms).
[0004] 2. Obstacles in multimodal data fusion: Industrial scenarios involve heterogeneous data such as vibration signals (0 - 20 kHz frequency domain), infrared thermal imaging (640×480 resolution), and device logs (unstructured text). Existing technologies (such as Huawei Cloud EI Industrial Intelligent Agent) lack a unified semantic interface, resulting in cross-level data association failures. Taking the fault diagnosis of a water turbine as an example, the mis-association rate between the temperature alarm data of the control cabinet (Modbus / TCP protocol) and the vibration spectrum characteristics of the subsystem (MQTT protocol) increases by 23%.
[0005] 3. Inefficiency in dynamic resource allocation: Traditional methods (such as the federated learning architecture proposed in CN114217314A with the publication number) adopt a static weight allocation strategy and cannot adapt to working condition drift.
[0006] Experiments show that when the load of a hydropower station suddenly changes (±30% of the rated power), network congestion is caused by API call conflicts at the subsystem layer, and the TCP retransmission rate surges from 0.5% to 12.3% (based on Wireshark packet capture analysis); After the device firmware is upgraded (such as PLC upgraded from v2.1 to v3.0), the data distribution shift (KL divergence>0.15) causes the decision accuracy of the traditional Bayesian network (CN113570129A) to drop by 41.2%. Summary of the invention
[0007] The purpose of the present invention is to provide a multimodal industrial operation and maintenance system, method and application that integrates a hierarchical large model architecture and adaptive reinforcement learning to improve the technical problems mentioned in the above technical background.
[0008] The present invention is achieved through the following technical solutions: A multimodal industrial operation and maintenance system that integrates hierarchical large model architecture and adaptive reinforcement learning, including expert layer module, system layer module, general engineer layer module, adaptive reinforcement learning engine and Bayesian collaborative network. The expert layer module is a lightweight DeepSeek model deployed on industrial field control equipment. M I , used to receive equipment time series data and local knowledge manual word vectors to generate preliminary operation and maintenance decisions; The system layer module is a medium-sized DeepSeek model deployed on an industrial computer. M m , receiving the decision outputs of at least two expert-level modules and corresponding system-level documents, and performing multimodal data fusion and conflict resolution; The general engineering layer module is a complete DeepSeek model deployed on the factory-level server M c , receiving the output of all system-level modules and plant-wide documents, and generating a global optimization strategy; The adaptive reinforcement learning engine dynamically adjusts the API call weights between levels based on three-dimensional reward and punishment signals; Bayesian collaborative networks are used to model the state transition probabilities between levels.
[0009] Furthermore, multimodal data fusion includes: a1. Convert the device knowledge manual text into a 128-dimensional word vector using the BERT-Whitening algorithm; b1. Perform wavelet denoising and LSTM feature extraction on sensor time series data; c1. Analyze the speech / text interaction content of the operation and maintenance personnel through the pre-trained BERT-Emotion model and output the emotion value Feeling∈[-5,+5].
[0010] Furthermore, the adaptive reinforcement learning engine includes: a2. Reward and Punishment Integration Network: The gated recurrent unit (GRU) is adopted, with the hidden layer dimension d = 64 and the activation function being Sigmoid; b2. Policy Optimization Module: The following Bellman equation model update is performed at the system layer: ; Among them, A t represents the selected action; Q( M m , A t ) represents the value of the action selected by the model M m in the current state; A t ; η represents the learning rate ( η ∈[0.05, 0.2]); R ' represents the dynamic reward; A ' represents all possible actions; γ is the discount factor (the discount factor γ = 0.99, with an exponential decay rate of 0.95 / period), used to balance long-term and short-term interests; τ represents the temperature coefficient ( τ = 0.1), used to balance the control policy selection and exploration.
[0011] Furthermore, the three-dimensional reward and punishment signal is the artificial score Rating, sentiment analysis Feeling, and model - to - model score Model; The weight update of the adaptive reinforcement learning engine satisfies the following relationship: , Among them, w i represents the weight of the i -th layer model at time step t, R ( t ) represents the reward at the current time step t; the learning rate η ∈[0.05, 0.2]; λ represents the regularization coefficient ( λ = 0.85, with a 95% confidence interval of ±0.03), used to balance policy exploration and exploitation; K represents the weight dimension involved in the calculation, and the L2 (L2 is the 2 - norm) penalty coefficient β = 1e - 4, used to prevent weight overfitting.
[0012] Furthermore, the Bayesian collaborative network defines that the conditional probability of the upper layer on the lower layer satisfies the following relationship. Taking the total layer to the system layer as an example: , Among them, c' represents all possible total worker layer models, θ c ∈ R d is the learnable parameter from the system layer to the total worker layer (dimension d = 128), T represents matrix transpose operation. The feature mapping function ϕ(•) is implemented using two layers of CNN. The first layer uses the ReLU activation function, and the second layer uses the Tanh activation function. The convolution kernel size is 3×3, and the stride is 1.
[0013] Furthermore, it also includes an efficiency evaluation module for: Ⅰ. Statistically calculate the contribution degree indicators of each expert layer module, including decision adoption rate, inference time consumption, and user satisfaction index; Ⅱ. When the contribution degree of a certain expert module is <0.2 for three consecutive evaluation cycles (each cycle is 1 week), trigger at least one of the following operations: Model parameter expansion: Upgrade the model scale according to the prompt of the hardware adaptation table; Precision quantization compression: Adopt FP8 quantization to reduce resource occupancy; Knowledge base update: Retrieve the latest documents through the Retrieval-Augmented Generation (RAG) technology and fine-tune the model.
[0014] Furthermore, the contribution degree indicators satisfy the following relational formula: , where α1 = 0.6, α2 = 0.3, α3 = 0.1, the KL divergence is calculated through variational inference, and the threshold <1e - 3.
[0015] The multi-modal industrial operation and maintenance method based on the aforementioned system includes the steps: S1. Deploy the lightweight DeepSeek- M I model on the edge device and load the local knowledge word vectors; S2. The system layer model M m receives the decisions of the expert layer and performs multi-source alignment based on the attention mechanism; S3. The total worker layer model M c generates the global policy and issues optimization instructions through the Bayesian network; S4. Update the weights between layers according to the real-time reward and punishment signals and synchronously adjust the parameters of the Bayesian network θ c .
[0016] Furthermore, in step S2, the multi-source alignment includes: S21. Apply a weighted voting mechanism to conflicting decisions, where the voting weight is proportional to the historical accuracy rate of the expert module; S22. Perform feature splicing on complementary decisions and input them into the GRU network to generate a fusion result; In step S4, the parameter update needs to satisfy: S41. Learning rate η Adjust according to the cosine annealing strategy: , where t represents the current training step, the period T = 1000, η min = 0.05 ,η max = 0.2; S42. Bayesian network parameters θ c Update through the variational EM algorithm every 50 cycles, and the convergence condition is that the change rate of ELBO < 0.1%.
[0017] The application of the multimodal industrial operation and maintenance system or method integrating the hierarchical large model architecture and adaptive reinforcement learning in hydropower stations, thermal power stations, and intelligent manufacturing production lines as described in any of the previous items.
[0018] Compared with the prior art, the present invention has the following advantages and beneficial effects: In the present invention, the proposed multimodal industrial operation and maintenance system and method integrating the hierarchical large model architecture and adaptive reinforcement learning break through the limitations of the single-layer model architecture, the obstacles of multimodal data fusion, and the inefficiency of dynamic resource allocation by establishing a hierarchical structure and a dynamic adjustment mechanism, improve the intelligent level of the system, optimize the industrial operation and maintenance efficiency, and can be applied in hydropower stations, thermal power stations, and intelligent manufacturing production lines. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is the system architecture diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0020] The present invention will be further described in detail below with reference to the embodiments, but the embodiments of the present invention are not limited thereto.
[0021] Embodiment 1 To facilitate the public's understanding of the present invention, this embodiment takes a specific multimodal industrial operation and maintenance system integrating the hierarchical large model architecture and adaptive reinforcement learning as an example, and further illustrates the solution in combination with the drawings.
[0022] A multimodal industrial operation and maintenance system integrating a hierarchical large model architecture and adaptive reinforcement learning, including an expert layer module, a system layer module, a general engineer layer module, an adaptive reinforcement learning engine, and a Bayesian collaborative network.
[0023] The expert layer module is a lightweight DeepSeek model deployed on industrial field control devices M I , which is used to receive device time-series data and local knowledge manual word vectors and generate preliminary operation and maintenance decisions. The system layer module is a medium-sized DeepSeek model deployed on an industrial control computer M m , which receives the decision outputs of at least two expert layer modules and corresponding system-level documents and performs multimodal data fusion and conflict resolution; Among them, multimodal data fusion includes: a1. Convert the device knowledge manual text into 128-dimensional word vectors through the BERT-Whitening algorithm; b1. Perform wavelet denoising and LSTM feature extraction on sensor time-series data; c1. Parse the voice / text interaction content of operation and maintenance personnel through the pre-trained BERT-Emotion model and output the emotion value Feeling∈[-5, +5].
[0024] The general engineer layer module is a complete DeepSeek model deployed on the factory-level server M c , which receives the outputs of all system layer modules and the whole-plant-level documents and generates a global optimization strategy. The adaptive reinforcement learning engine is used to dynamically adjust the API call weights between levels based on three-dimensional reward and punishment (manual scoring Rating, sentiment analysis Feeling, model-inter model scoring Model) signals, The adaptive reinforcement learning engine includes: a2. Reward and punishment fusion network: Adopt a gated recurrent unit GRU, with a hidden layer dimension d = 64 and an activation function of Sigmoid; b2. Policy optimization module: Execute the improved Bellman equation update. Taking the system layer as an example: ; Among them, A t represents the selected action, Q( M m , A t ) represents the model in the current state M m selects the action A tValue, learning rate η ∈ [0.05, 0.2], R ' represents dynamic reward, A ' represents all possible actions, discount factor γ = 0.99 (exponential decay rate 0.95 / cycle), used to balance long - term and short - term interests, temperature coefficient τ = 0.1, used to balance control strategy selection and exploration.
[0025] The weight update of the adaptive reinforcement learning engine satisfies the following relational expression: , where,, w i represents the weight of the i layer model at time step t, R ( t ) represents the reward at the current time step t, learning rate η ∈ [0.05, 0.2], regularization coefficient λ = 0.85 (95% confidence interval ±0.03), used to balance policy exploration and exploitation, K represents the weight dimension participating in the calculation, L2 (L2 is the norm of 2) penalty coefficient β = 1e - 4, used to prevent weight overfitting.
[0026] The Bayesian collaborative network is used to model the state transition probability between levels.
[0027] The Bayesian collaborative network defines the conditional probability of the upper layer to the lower layer (such as the general engineering layer to the system layer) as: , where, c' represents all possible general engineering layer models, θ c ∈ R d is the learnable parameter from the system layer to the general engineering layer (dimension d = 128), the feature mapping function ϕ(•) is implemented using two - layer CNN, the first layer uses the ReLU activation function, the second layer uses the Tanh activation function, the convolution kernel size is 3×3, and the stride is 1.
[0028] Preferably, it further includes an efficiency evaluation module for: Ⅰ. Statistically calculate the contribution degree indicators of each expert layer module, including decision adoption rate, inference time consumption, and user satisfaction index; Ⅱ. When the contribution degree of a certain expert module is <0.2 for three consecutive evaluation cycles (each cycle is 1 week), trigger at least one of the following operations: Model parameter expansion: Upgrade the model scale according to the prompt of the hardware adaptation table (7B → 14B); Precision quantization compression: Adopt FP8 quantization to reduce resource occupancy; Knowledge base update: Retrieve the latest documents through the Retrieval-Augmented Generation (RAG) technique and fine-tune the model.
[0029] The contribution degree index satisfies the following relational expression: , where α1 = 0.6, α2 = 0.3, α3 = 0.1, the KL divergence is calculated through variational inference, and the threshold is <1e-3.
[0030] The multi-modal industrial operation and maintenance method based on the above system refers to Figure 1 , including the following steps: Step 1: Deploy the lightweight DeepSeek- M I model on the edge device, and load the local knowledge word vectors and time series data.
[0031] In this step, deploy the lightweight DeepSeek model in the PLC or edge box of each control cabinet M I , as an agent expert, responsible for preliminary data processing and local decision-making. And input the knowledge manual related to this control cabinet into the model in the form of word vectors and the collected time series historical data M I .
[0032] Step 2: The system layer model M m Receives the decision of the expert layer and performs multi-source alignment based on the attention mechanism.
[0033] In this step, deploy a medium-sized DeepSeek model on an industrial control computer for each system M m , as the agent system responsible person, to integrate the decisions of each expert. And convert the knowledge manual related to this system into word vectors for the model M m to use.
[0034] The multi-source alignment includes: S21. Apply a weighted voting mechanism to conflicting decisions, and the voting weight is proportional to the historical accuracy rate of the expert module; S22. Perform feature splicing on complementary decisions and input them into the GRU network to generate a fusion result.
[0035] Step 3: The general engineer layer model M c Generates a global policy and issues optimization instructions through the Bayesian network.
[0036] In this step, deploy the full version of the DeepSeek model on the plant-level server M c , as the acting chief engineer, coordinate the decisions of each subsystem. And convert the plant-level knowledge manual into word vectors for the model M c to use.
[0037] Step 4: Update the weights between levels according to the real-time reward and punishment signals, and synchronously adjust the Bayesian network parameters θ c .
[0038] The parameter update needs to satisfy:[[]] S41. Learning rate η Adjust according to the cosine annealing strategy: , where t represents the current training step, the period T = 1000, η min = 0.05 ,η max = 0.2; S42. Bayesian network parameters θ c Update every 50 cycles through the variational EM algorithm, and the convergence condition is that the change rate of ELBO < 0.1%.
[0039] In this step, when the operation and maintenance personnel interact with the model, the system will call the APIs of this layer and the next layer (if any), combine the local knowledge manual with the feedback from the lower-layer agent (if any), and generate corresponding answers. At the same time, the system can automatically generate operation and maintenance reports at preset times when there is no manual operation.
[0040] During the operation and maintenance interaction process, adopt a hierarchical reinforcement learning mechanism that integrates Bayesian inference to dynamically optimize the model weights.
[0041] 1. Obtain multi-source reward and punishment signals.
[0042] 1.1 Manual evaluation signal: The operation and maintenance personnel give an immediate score of 1 - 10 for the generated text (Rating ∈ N + ), and establish the mapping relationship of the operation log: Rating → <device ID, decision level, timestamp>.
[0043] 1.2 Emotional semantic signal: Parse the interactive text / speech through the pre-trained BERT-Emotion model, and output the emotional value Feeling ∈ [-5, +5], and its softmax probability distribution satisfies: P(Feeling ≤ 0) < 0.2 (confidence level ≥ 95%).
[0044] Scoring signal between models: The upper-layer model scores the decision-making quality of the output of the lower layer. Model ∈ [1, 10], and semantic scoring based on LLM is adopted. The upper-layer model is given a clear role and is guided to score through a structured Prompt template.
[0045] 2. Reward fusion and weight update.
[0046] Design a reward function implemented by a gated recurrent neural network (GRU): ; where σ(•) is the Sigmoid activation function, and the output is the normalized reward value R ( t ) ∈ [0, 1], the weight matrix W r ∈ R 2 ×d , W m ∈ R 1×d (the hidden layer dimension d = 64), and the bias term b ∈ R d is initialized by Xavier.
[0047] The weight update adopts a policy gradient method with regularization: ; where, w i represents the weight of the i -th layer model at time step t, R ( t ) represents the reward at the current time step t. The learning rate η ∈ [0.05, 0.2] is dynamically adjusted by Bayesian optimization. The regularization coefficient λ = 0.85 (95% confidence interval ±0.03) is used to balance policy exploration and exploitation. K represents the weight dimension participating in the calculation, and the L2 penalty coefficient β = 1e-4 is used to prevent weight overfitting.
[0048] 3. Bayesian-reinforcement cooperation mechanism 1) Conditional probability modeling: Construct a three-layer Bayesian network, define the state transition probability, and satisfy the following relationship: ; where c' represents all possible total layer models, θ c ∈ R dIt is a learnable parameter from the system layer to the chief engineer layer (dimension d = 128). The feature mapping function ϕ(•) adopts a two-layer CNN structure. The ReLU activation function is used in the first layer, and the Tanh activation function is used in the second layer. The convolution kernel size is 3×3, and the stride is 1.
[0049] 2) Dynamic reward update: Correct the Q value based on the posterior probability, satisfying the following relational expression: ; where E is the conditional expectation, estimated by importance sampling, KL ( P prior || P posterior ) is the KL divergence of the prior distribution relative to the posterior distribution, estimated by variational method. Coefficient constraint: α1 + α2 + α3 = 1 (default α1 = 0.6, α2 = 0.3, α3 = 0.1).
[0050] 3) Cooperative decision update: Adopt an improved Bellman equation, satisfying the following relational expression: ; where, A t represents the selected action, Q( M m , A t ) represents the value of the model M m selecting the action A t at the current state. The learning rate η ∈[0.05, 0.2], R ' represents the dynamic reward, A ' represents all possible actions, and the discount factor γ = 0.99 (exponential decay rate 0.95 / cycle), used to balance long-term and short-term interests. The temperature coefficient τ = 0.1, used to balance the control strategy selection and exploration.
[0051] 4. Online learning optimization Execute per cycle: 1) Update the Bayesian network parameters through variational inference θ c (KL divergence threshold < 1e-3); 2) Use importance sampling to adjust the policy gradient, satisfying the following relational expression: ; 3) Perform soft update of the model weights, satisfying the following relational expression: , Among them, the mixing coefficient ρ = 0.85.
[0052] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Efficiency improvement brought by architecture innovation Multi-level decision optimization: Through the three-level model architecture of the expert layer ( M l ), the system layer ( M m ), and the chief engineer layer ( M c ), the coordination between device-level local decision-making and system-level global optimization is realized, avoiding the response delay problem caused by the single decision-making dimension of the traditional single-layer model.
[0053] Dynamic resource allocation efficiency: The weight update formula (including the regularization term λ) based on reinforcement learning is coordinated with the Bayesian network to solve the low efficiency problem of traditional static resource allocation.
[0054] 2. Breakthrough in the level of intelligence Multi-modal data fusion ability: Convert multi-modal inputs such as knowledge manual text, time-series data, and voice interaction into unified word vectors, breaking through the limitations of traditional single-modal analysis.
[0055] Adaptive learning mechanism: The state transition probability modeling (Dirichlet prior distribution) based on the Bayesian network is coordinated with the dynamic reward and punishment mechanism of reinforcement learning to achieve real-time policy optimization.
[0056] 3. Optimization of operation and maintenance costs and reliability Reduction of operation and maintenance costs: Through lightweight model deployment (expert layer M I ), and automated report generation (triggered at preset times), the need for manual intervention is reduced.
[0057] Enhanced system reliability: The fault tolerance mechanism of the hierarchical architecture (such as the decision fusion function of the system layer M m ) is coordinated with the anomaly detection ability of dynamic reinforcement learning.
[0058] 4. Interpretability-driven performance analysis Visualization of the inference path: The function of showing the active thinking process of the DeepSeek model can record the decision-making chain from the expert layer ( M l ) to the chief engineer layer ( M c ).
[0059] Quantitative indicators of expert effectiveness, establishing a three-level evaluation system: 1). Response quality: Calculate the user satisfaction index of the expert module based on the ratings given by operation and maintenance personnel and sentiment analysis. 2). Decision-making efficiency: Statistically analyze performance parameters such as the inference time consumption and resource occupancy rate of the expert model. 3). Knowledge coverage: Evaluate the need for document updates through the cosine similarity between the word vectors of the knowledge manual and the expert output.
[0060] The multi-modal industrial operation and maintenance system and method integrating a hierarchical large model architecture and adaptive reinforcement learning of the present invention can be applied in hydropower stations, thermal power stations, and intelligent manufacturing production lines.
[0061] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Any simple modification or equivalent change made to the above embodiments based on the technical essence of the present invention shall fall within the protection scope of the present invention.
Claims
1. A multimodal industrial operation and maintenance system integrating a hierarchical large model architecture and adaptive reinforcement learning, characterized in that: It includes an expert layer module, a system layer module, a chief engineer layer module, an adaptive reinforcement learning engine, and a Bayesian collaborative network. The expert layer module is a lightweight DeepSeek model deployed on industrial field control devices M I , which is used to receive device time series data and local knowledge manual word vectors and generate preliminary operation and maintenance decisions; The system layer module is a medium-sized DeepSeek model deployed on an industrial control computer M m , receives the decision outputs and corresponding system-level documents of at least two expert layer modules, and performs multimodal data fusion and conflict resolution; The general engineering layer module is a complete DeepSeek model deployed on the factory-level server M c , receiving the output of all system-level modules and plant-wide documents, and generating a global optimization strategy; The adaptive reinforcement learning engine dynamically adjusts the API call weights between layers based on three-dimensional reward and punishment signals. The Bayesian collaborative network is used to model the state transition probability between layers.
2. The multimodal industrial operation and maintenance system integrating a hierarchical large model architecture and adaptive reinforcement learning according to claim 1, characterized in that, Multimodal data fusion includes: a1. Converting the device knowledge manual text into 128-dimensional word vectors through the BERT-Whitening algorithm. b1. Performing wavelet denoising and LSTM feature extraction on the sensor time series data. c1. Parsing the voice / text interaction content of the operation and maintenance personnel through the pre-trained BERT-Emotion model, and outputting the emotion value Feeling ∈ [-5, +5].
3. The multimodal industrial operation and maintenance system integrating a hierarchical large model architecture and adaptive reinforcement learning according to claim 1, characterized in that, The adaptive reinforcement learning engine includes: a2. A reward and punishment fusion network: Using a gated recurrent unit GRU, with a hidden layer dimension d = 64 and an activation function of Sigmoid. b2. A policy optimization module: The system layer performs the following Bellman equation model update: ; Among them, A t represents a selection action; Q( M m , A t ) represents the value of the model M m for the selection action A t ; η represents the learning rate; R ' represents the dynamic reward; A ' represents all possible actions; γ is the discount factor, used to balance short-term and long-term interests; τ represents the temperature coefficient, used to balance the control policy selection and exploration.
4. The multimodal industrial operation and maintenance system integrating a hierarchical large model architecture and adaptive reinforcement learning according to claim 1, characterized in that: The three-dimensional reward and punishment signals are the artificial score Rating, the sentiment analysis Feeling, and the inter-model score Model. The weight update of the adaptive reinforcement learning engine satisfies the following relationship: , Among them, w i represents the weight of the i -th layer model at time step t, R ( t ) represents the reward at the current time step t; λ represents the regularization coefficient, which is used to balance policy exploration and exploitation; K represents the weight dimension involved in the calculation, and the L2 penalty coefficient β = 1e-4 is used to prevent weight overfitting.
5. The multimodal industrial operation and maintenance system integrating a hierarchical large model architecture and adaptive reinforcement learning according to claim 1, characterized in that: The Bayesian collaborative network defines that the conditional probability of the upper layer to the lower layer satisfies the following relationship: Among them, c' represents all possible total layer models, θ c ∈ R d is the learnable parameter from the system layer to the total layer, T represents the matrix transpose operation. The feature mapping function ϕ(•) is implemented using two layers of CNN. The first layer uses the ReLU activation function, and the second layer uses the Tanh activation function. The convolution kernel size is 3×3, and the stride is 1.
6. The multimodal industrial operation and maintenance system integrating a hierarchical large model architecture and adaptive reinforcement learning according to claim 1, characterized in that It also includes an efficiency evaluation module, which is used for: Ⅰ. Statistically calculating the contribution degree indicators of each expert layer module, including the decision adoption rate, the inference time consumption, and the user satisfaction index. Ⅱ. When the contribution degree of a certain expert module is < 0.2 for three consecutive evaluation periods, trigger at least one of the following operations: Model parameter expansion: Upgrade the model scale according to the prompt of the hardware adaptation table. Precision quantization compression: Adopt FP8 quantization to reduce resource occupation. Knowledge base update: Retrieve the latest documents through the retrieval-augmented generation RAG technology and fine-tune the model.
7. The multimodal industrial operation and maintenance system integrating a hierarchical large model architecture and adaptive reinforcement learning according to claim 6, characterized in that The contribution degree indicators satisfy the following relationship: , Among them, α1 = 0.6, α2 = 0.3, α3 = 0.1, the KL divergence is calculated through variational inference, and the threshold < 1e-3.
8. The multimodal industrial operation and maintenance method based on the system described in claim 1, characterized in that, It includes steps: S1. Deploy the lightweight DeepSeek- M I model on the edge device and load the local knowledge word vectors; S2. System layer model M m Receive the decisions from the expert layer and perform multi-source alignment based on the attention mechanism; S3. General Engineering Layer Model M c Generate a global policy and issue optimization instructions through the Bayesian network; S4. Update the weights between levels according to the real-time reward and punishment signals, and synchronously adjust the Bayesian network parameters θ c 。 9. The multimodal industrial operation and maintenance method according to claim 8, wherein, In step S2, the multi-source alignment includes: S21. Applying a weighted voting mechanism to conflicting decisions, and the voting weight is proportional to the historical correct rate of the expert module. S22. Performing feature splicing on complementary decisions, and inputting them into the GRU network to generate a fusion result. In step S4, the parameter update needs to satisfy: S41. Learning rate η Adjust according to the cosine annealing strategy: , Among them, t represents the current training step, and the period T = 1000; S42, Bayesian network parameters θ c Updated every 50 cycles by the variational EM algorithm, and the convergence condition is that the change rate of ELBO < 0.1%.
10. The application of the multimodal industrial operation and maintenance system or method integrating the hierarchical large model architecture and adaptive reinforcement learning according to any one of claims 1 to 9 in a hydropower station, a thermal power station, and an intelligent manufacturing production line.
Citation Information
Patent Citations
A strip steel pickling concentration prediction method and a computer readable storage medium
CN113570129A
Obstacle avoidance method and application of mobile obstacle avoidance safety system
CN114217314A
Multi-attention fusion deep residual shrinkage network soft measurement modeling method based on Bayesian optimization
CN115203954A
5G intelligent leveling method and system for civil aviation and fire rescue equipment
CN118246704A
Multi-task cooperative processing method based on artificial intelligence agent and event chain
CN118331199A
Cited By
Multi-sensor fusion aviation food truck intelligent monitoring and early warning method and device
CN120913380A
Mine production plan intelligent optimization system
CN121303480A