A human-machine collaborative intelligent control system based on AIGC
Through technologies such as multimodal biological signal completion, game theory optimization decision-making and meta-reinforcement learning, the problems of instability in input, poor decision consistency, insufficient execution feedback and algorithm deviation in the human-machine collaborative intelligent control system are solved, and efficient, interpretable and optimized intelligent control is achieved, improving the system's application capabilities in complex environments.
Patent Information
- Application Number
- CN202510423767.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The existing human-machine collaborative intelligent control system has unstable input of multimodal biological signals, poor consistency of group intelligent decision-making, lack of adaptive ability in artificial intelligence execution feedback, and deviation and fairness of algorithms, which affect the application of the system in complex environments.
Multimodal biological signal completion, game theory optimization decision-making, meta-reinforcement learning execution, adaptive federated learning and causal inference fairness monitoring technology are adopted to build an efficient, interpretable, and optimized intelligent control framework, including BI input module, CI decision-making module, AI execution feedback module, dynamic learning and cost optimization module, algorithm deviation and fairness monitoring module, user interface and interaction module.
It realizes high-complete user status data input, improves consistent decision-making capabilities across devices and across users, enhances AI execution and feedback capabilities, ensures algorithm fairness and user privacy protection, significantly enhances the system's personalized adaptability and user experience, improves computing efficiency and reduces latency.
Smart Images

Figure CN119940425B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an AIGC-based human-machine collaborative intelligent control system. Background Art
[0002] With the rapid development of artificial intelligence (AI), intelligent human-machine collaborative control systems are gaining widespread application in areas such as autonomous driving, smart homes, and medical assistance. These systems typically rely on biological intelligence to provide user input, swarm intelligence to optimize decision-making, and artificial intelligence to execute tasks, enabling multi-level human-machine interaction. However, existing human-machine collaborative systems still face numerous challenges, particularly in cross-modal data fusion, personalized learning, and decision fairness.
[0003] Current intelligent control systems face the following major challenges: limited multimodal biosignal data quality, with signal loss and noise leading to unstable input; irrational weight distribution in swarm intelligence decisions, making it difficult to ensure decision consistency across devices; a lack of dynamic learning mechanisms for AI execution feedback, making it difficult to adapt to changing user needs; algorithmic biases and fairness issues impact the user experience across different user groups; and relatively rigid system interaction methods lack personalized optimization and human-machine co-adaptation capabilities. These issues hinder the widespread application of human-machine collaborative intelligent systems in complex environments.
[0004] To address the above problems, the present invention proposes an AIGC-based human-machine collaborative intelligent control system, which adopts multimodal biosignal completion, game theory optimization decision-making, meta-reinforcement learning execution, adaptive federated learning and causal inference fairness monitoring technologies to build an efficient, explainable and optimizable intelligent control framework. Summary of the Invention
[0005] In response to the above problems, the present invention provides an AIGC-based human-machine collaborative intelligent control system to solve the problems of existing technologies in terms of unstable multimodal biosignal input, poor consistency of group intelligent decision-making, lack of adaptive ability of artificial intelligence execution feedback, and algorithm bias and fairness.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: a human-machine collaborative intelligent control system based on AIGC, comprising:
[0007] BI input module, which realizes human-computer collaborative input through multimodal biosignal perception and dynamic modeling, and provides real-time user status data support;
[0008] A CI decision module, which is used to integrate data from the BI input module, perform multi-objective optimization and federated learning scheduling, and achieve intelligent decision-making across devices and users;
[0009] An AI execution feedback module that executes instructions based on the optimized decisions of the CI decision module and provides an adaptive execution solution based on reinforcement learning and multimodal feedback;
[0010] A dynamic learning and cost optimization module, which is used to optimize model training and deployment, ensuring a balance between system computing efficiency and privacy protection;
[0011] Algorithm bias and fairness monitoring module, which is used to monitor system algorithm bias and ensure fairness;
[0012] User interface and interaction module, which are used to enhance user experience and optimize human-computer interaction.
[0013] The BI input module includes a sensing unit, a signal processing unit, a dynamic image modeling unit and a physiological signal quality assessment unit;
[0014] The sensing unit aligns multimodal data from an EEG signal acquisition device, a heart rate sensor, a voice acquisition device, and a gesture capture device through a heterogeneous data interface protocol. The multimodal data includes EEG signals, heart rate, voice, gestures, and external device data. Physiological feature vectors are derived through EEG signal analysis and heart rate signal calculation. A dual-discriminator generative adversarial network is used to complete the sensor's missing abnormal signals. The generator uses a U-Net deep neural network architecture to reconstruct and complete the missing multimodal signals to generate a complete completed signal. The dual discriminator verifies signal integrity and physiological rationality, respectively, and outputs a multimodal signal stream with a time synchronization error of less than 1ms.
[0015] The signal processing unit uses Kalman filtering and wavelet packet decomposition to eliminate baseline drift and electromyographic interference problems that occur when supplementing abnormal missing sensor signals, and extracts energy features of the α, β, and θ bands. The improved SimCLR contrastive learning framework is used to learn potential representations from cross-modal data, where positive samples are paired data of EEG signals and heart rate signals of the same user, and negative samples are randomly paired data across users. The loss function uses normalized temperature-scaled cross entropy and outputs a 128-dimensional standardized feature vector. The contrastive learning loss function is specifically shown as follows:
[0016]
[0017] in, is the contrastive learning loss function, is the embedding vector of the positive sample pair, is the temperature parameter, is the number of negative samples in the batch, is the cosine similarity function, which calculates the similarity between two vectors;
[0018] The dynamic portrait modeling unit constructs user portraits based on a semi-supervised spatiotemporal graph network, where nodes include real users and 200 virtual users generated daily. Virtual users are generated through latent variable interpolation of a variational graph autoencoder, with a KL divergence of less than 0.1. Edge weights are defined by behavioral similarity, and a pseudo-label propagation mechanism is used to fuse key events actively annotated by users, outputting dynamic short-term states and long-term behavioral preference maps. The model of virtual users generated by latent variable interpolation of the variational graph autoencoder is specifically shown in the formula:
[0019]
[0020] in, Evidence lower bounds for VGAE generation models, is the encoder network, which is used to map the input data X to the latent variable z; is the decoder network, used to reconstruct data X from the latent variable z; is the prior distribution of latent variables, for Divergence, which measures the difference between the encoder distribution and the prior distribution;
[0021] The physiological signal quality assessment unit detects the sensor contact status in real time, triggering a warning when the EEG signal electrode impedance is greater than 50kΩ. After the abnormality persists for 10 seconds, the adversarial network completion mode is activated, and a voice prompt "Sensor contact is poor, completion mode has been enabled" is pushed simultaneously; and a three-level warning of the device health status is output, including normal, warning and fault.
[0022] The CI decision module includes a data fusion unit, a consensus unit, a multi-objective collaborative optimization unit, a federated learning scheduling unit and a conflict resolution engine unit;
[0023] The data fusion unit uses an attention mechanism to dynamically weight the integrated data, which comes from the physiological feature vectors of the BI module and the external device data, eliminating redundant information and generating a spatiotemporally aligned fusion feature tensor;
[0024] The consensus unit dynamically assigns decision weights based on Shapley values in game theory and calculates the historical accuracy of each agent, which includes an EEG signal acquisition device, a heart rate sensor, a voice acquisition device, and a gesture capture device. It generates an initial set of command candidates through weighted voting, introduces a conflict detection mechanism to ensure cross-device decision consistency, and outputs a prioritized command sequence.
[0025] The multi-objective collaborative optimization unit uses an improved NSGA-III algorithm to perform multi-objective optimization based on the priority-ranked instruction sequence generated by the consensus unit, comprehensively considering energy consumption, response delay, and user satisfaction to solve the optimal Pareto front solution set. The Pareto front solution set represents the optimal instruction solution that strikes a balance between different optimization objectives, ensuring that the final decision meets both system performance requirements and user experience. The optimized instruction set is converted into a natural language description through the T5 model, and a human-understandable optimized instruction set is output for system execution and adjustment.
[0026] The federated learning scheduling unit adopts a hierarchical federated architecture, sets personalized training goals based on an optimized instruction set, compresses model parameters through differentiated knowledge distillation, and dynamically adjusts the communication threshold for model synchronization based on the model update amplitude, computing resource usage, and network bandwidth. A fairness regularization term is added during edge aggregation to adjust the balance between the personalized model and the global model, and finally outputs the updated global model. The loss function is shown in the formula:
[0027]
[0028] in, is the federal fairness constraint loss, Loss of main mission, is the fairness regularization term, is the fairness weight coefficient;
[0029] The conflict resolution engine unit takes the global optimization model and optimized instruction set output by the federated learning scheduling unit as input, matches historical similar scenarios based on case reasoning, uses Sentence-BERT to encode the instruction context, and retrieves the top three historical solutions for user reference; if the match fails, the reinforcement learning unit is called to generate a new strategy to ensure the consistency and executableness of the final instructions, and outputs the final execution instructions after conflict resolution.
[0030] The AI execution feedback module includes a reinforcement learning unit, an execution unit, a feedback unit, and an interpretation unit;
[0031] The reinforcement learning unit pre-trains a cross-scenario policy network based on a meta-reinforcement learning framework, inputs the instruction set and environment data optimized by the CI decision module, and uses the Unity simulator to generate diverse control scenario data. Through model-independent meta-learning, it quickly adapts to new tasks within 10 gradient updates, dynamically balances user manual intervention and autonomous decision-making, and outputs a real-time control strategy, as shown in the formula:
[0032]
[0033] in, are the updated model parameters, are the initial model parameters, β is the meta-learning rate, is the task loss function, is the parameter change of one gradient update;
[0034] The execution unit uses the predictive execution model LSTM timing prediction, combined with resource-aware scheduling for optimized execution, inputs the Pareto optimization instruction set generated by the CI module, preloads context-related actions, and dynamically adjusts the model complexity based on the device computing power to output physical control signals, as shown in the formula:
[0035]
[0036] in, For predictive execution loss function, For the actual execution results, is the model prediction result, is the L1 regularization weight, are model parameters;
[0037] The feedback unit uses multimodal generative feedback based on the physical control signals of the actuator to provide a personalized interactive experience. Stable Diffusion is used to generate guidance animations and voice prompts, and the CLIP model is used for user state matching, analyzing physiological signals, interaction history, and environmental parameters. The feedback mechanism adopts a real-time and offline two-layer architecture: the real-time layer calculates key feature weights and provides immediate feedback; the offline layer constructs a decision logic tree, analyzes long-term interaction patterns, and provides users with interpretable interaction content.
[0038] The explanation unit adopts a dynamic explainability framework to analyze the explainability of model decisions based on the physical control signals of the execution unit, the user interaction data of the feedback unit and the system's historical decision logs. The online layer locates key decision feature areas in the input data through gradient-weighted class activation mapping and provides real-time explanations. The offline layer combines causal reasoning methods to construct a causal graph model, which supports users to trace back any historical decision path, track the system's decision-making process in different scenarios, and output visual reports and fairness audit logs.
[0039] The dynamic learning and cost optimization module combines the learning unit, the model compression unit, and the adaptive unit;
[0040] The joint learning unit adopts a hierarchical federated architecture, deploys a lightweight student model on the terminal device, and aggregates the gradient of the teacher model on the edge server for update, ensuring a balance between privacy protection and computational efficiency. The unit uses differential knowledge distillation technology to compress the model and reduce the amount of communication data. At the same time, it uses the Paillier homomorphic encryption algorithm to protect the privacy security of gradient transmission, ensuring that the accuracy of the global model does not decay by more than 2% within the monthly update cycle. It outputs a federated model version adapted to multiple devices. The loss function is shown in the formula:
[0041]
[0042] in, is the knowledge distillation loss function, is the temperature parameter, is the probability distribution output by the student model, is the probability distribution output by the teacher model, is the distillation loss weight;
[0043] The model compression unit dynamically adjusts the model structure based on a resource-aware mechanism, inputs real-time device status data, and generates lightweight sub-models through knowledge distillation technology. It uses elastic scaling to dynamically adjust the channel width of MobileNetV3 to ensure that the compressed model size is less than or equal to 5MB, while also increasing inference speed by 3 times. The final output is a model library adapted to different device resources.
[0044] The adaptive unit designs a multi-objective optimization strategy selector, which dynamically matches the optimal model compression strategy through a random forest classifier according to the task type and device resource status, and outputs real-time configuration instructions, reducing system response delay by 40%.
[0045] The algorithm bias and fairness monitoring module includes a bias detection unit, a privacy protection unit, a bias correction unit, and a fairness unit;
[0046] The bias detection unit monitors the decision data flow of the CI module in real time, uses a causal inference model to perform counterfactual fairness analysis, and identifies potential discrimination patterns through structural equation fit. It inputs the instruction set and user attribute labels generated by the CI module, calculates the bias risk score, and triggers a fairness alarm when the score exceeds 0.6. The model is validated based on the UCI Adult dataset, as shown in the formula:
[0047]
[0048] in, For sensitive attributes, is the counterfactual decision result when the sensitive attribute is intervened by k, It is a non-sensitive feature;
[0049] The privacy protection unit builds a triple protection chain, specifically: data collection end: injects differential privacy noise to prevent user identity leakage; transmission end: uses Paillier homomorphic encryption to protect gradient transmission; storage end: combines blockchain technology to record audit logs to ensure that data cannot be tampered with; input is raw biometric signals and user attributes, and output is a desensitized training dataset for subsequent fairness optimization;
[0050] The bias correction unit adopts a dual-channel adversarial fairness constraint, generates counterfactual samples through the GAN generator, and introduces fairness loss while optimizing the main task loss in the discriminator to reduce system bias. During federated learning aggregation, fairness regularization terms are added to ensure fairness when optimizing the global model. The input is biased model parameters, and the output is an unbiased model, as shown in the formula:
[0051]
[0052] in, is an unbiased model, G is the generator, D is the discriminator, is the fairness constraint weight, Loss of fairness;
[0053] The fairness unit continuously monitors system fairness and provides feedback and corrections to potential problems by tracking four core indicators: group fairness, ensuring fairness between different groups; individual fairness, ensuring that similar users receive similar decisions; equal opportunity, ensuring that different groups receive the same positive decision-making opportunities; causal fairness, analyzing potential biases in system decisions based on causal reasoning; using a dynamic dashboard to display real-time fairness monitoring results; and inputting historical decision logs and user feedback to output compliance reports.
[0054] The user interface and interaction module includes a visualization unit, a user feedback unit and a collaborative learning unit;
[0055] The visualization unit dynamically presents system decisions and feedback through a multimodal fusion interface, integrates StableDiffusion to generate personalized AR guidance animations, and combines the CLIP model to match user physiological characteristics with the knowledge base to display key decision-making basis in real time. The input is the layered interpretation results of the AI module and the user's real-time status, and the output is multimodal interactive content, which can improve user operation efficiency by 55%.
[0056] The user feedback unit is designed with a two-way collaborative learning mechanism to optimize the interactive experience through reinforcement learning and prediction of user preferences. After the user manually corrects the data, the PPO algorithm is used to reversely optimize the reinforcement learning strategy. Based on LSTM, the user preference migration trend is predicted, and touch / voice commands and physiological signals are input to dynamically adjust the user preference model.
[0057] The collaborative learning unit builds a human-machine co-adaptive decision-making framework, identifies the causal effects of user correction behavior through causal reinforcement learning, and uses a dynamic reward function to balance system goals and user preferences. The input is historical interaction logs and real-time feedback, and the output is a causal strategy map. The causal reward function and dynamic reward function are specifically shown as follows:
[0058]
[0059] in, is the causal reward function, Actions automatically performed by the system, is the action manually corrected by the user, sim is the action cosine similarity, and N is the number of historical interactions;
[0060]
[0061] in, is the weight matrix, is the hidden state at the previous moment, is the current input feature, and b is the bias term.
[0062] Compared with the prior art, the present invention has the following beneficial effects:
[0063] The present invention ensures high-integrity user status data input through multimodal fusion. Real-time signal completion and noise reduction make the data more continuous and reliable, avoiding decision lags and misjudgments caused by loss or delay of sensor data, and improving the accuracy and timeliness of the system's perception of user status.
[0064] This paper constructs a swarm intelligent decision-making module to improve the ability to make consistent decisions across devices and users. It uses an attention mechanism to dynamically weight and fuse data from various sensor devices and external systems, filtering out redundant information and aligning spatiotemporal features. It assigns decision weights to each device based on a consensus algorithm based on the Shapley value of game theory, and introduces a conflict detection mechanism to ensure the consistency of decision results under multi-device input.
[0065] This invention greatly enhances the execution and feedback capabilities of AI by introducing reinforcement learning and adaptive feedback mechanisms; the reinforcement learning unit adopts the meta-reinforcement learning framework to pre-train the policy network, enabling it to quickly migrate in different industry scenarios, and only requires about 10 gradient updates to adapt to new tasks and environmental changes; the execution unit uses a predictive model to predict future actions, and dynamically adjusts the model complexity in combination with resource-aware scheduling to adapt to the computing power of the terminal device, thereby ensuring the real-time execution of instructions.
[0066] This invention strengthens algorithmic fairness and user privacy protection. The bias detection unit performs causal inference analysis on the data stream output by the CI decision module, detects potential discrimination using a counterfactual method, and triggers a fairness alarm when the calculated bias risk score exceeds the threshold. The bias correction unit introduces adversarial fairness constraint training, generates counterfactual samples through GAN, minimizes task loss and fairness loss simultaneously during model training, reduces algorithmic bias, and outputs a model that eliminates unfairness.
[0067] This invention has made innovative designs in the user interface and interaction mode, significantly enhancing the system's personalized adaptation capabilities and user experience. The visualization unit uses a multimodal fusion interface to dynamically present the system decision-making process and feedback information, integrates the Stable Diffusion model to generate personalized AR animation guidance in real time, and uses the CLIP model to match the user's physiological characteristics with the knowledge base, intuitively displaying the key basis for AI decision-making. This rich visualization and explanatory content enables users to more clearly understand the system behavior.
[0068] This invention achieves higher computing efficiency, lower cost and latency by optimizing the system architecture and algorithms in many aspects. The joint learning unit adopts a hierarchical federated learning architecture, deploys lightweight student models on terminal devices, and aggregates teacher model gradients on edge servers, thereby reducing centralized communication bandwidth occupancy. While ensuring user privacy, it compresses model parameters through differentiated knowledge distillation, reduces model size and transmission volume, and ensures that even after long-term distributed training, the accuracy decay of the global model does not exceed 2% within a monthly cycle. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. It is understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0070] Figure 1 It is a system architecture diagram of the present invention;
[0071] Figure 2This is the BI input module architecture diagram of the present invention;
[0072] Figure 3 It is the CI decision module architecture diagram of the present invention;
[0073] Figure 4 This is the AI execution feedback module architecture diagram of the present invention;
[0074] Figure 5 This is a diagram of the dynamic learning and cost optimization module architecture of the present invention;
[0075] Figure 6 This is the architecture diagram of the algorithm bias and fairness monitoring module of the present invention;
[0076] Figure 7 It is a diagram of the user interface and interaction module architecture of the present invention. DETAILED DESCRIPTION
[0077] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention for which protection is sought, but is merely for selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0078] Please refer to Figure 1-7 Schematic diagram of a human-machine collaborative intelligent control system based on AIGC provided by an embodiment of the present invention, including:
[0079] BI input module: The BI input module realizes human-machine collaborative input through multimodal bio-signal perception and dynamic modeling, and provides real-time user status data support;
[0080] The CI decision module is used to integrate data from the BI input module, perform multi-objective optimization and federated learning scheduling, and achieve intelligent decision-making across devices and users;
[0081] The AI execution feedback module executes instructions based on the optimized decisions of the CI decision module and provides an adaptive execution solution based on reinforcement learning and multimodal feedback;
[0082] Dynamic learning and cost optimization module: This module is used to optimize model training and deployment, ensuring a balance between system computing efficiency and privacy protection.
[0083] Algorithm bias and fairness monitoring module, which is used to monitor system algorithm bias and ensure fairness;
[0084] User interface and interaction module: The user interface and interaction module are used to enhance user experience and optimize human-computer interaction.
[0085] The BI input module includes a sensing unit, a signal processing unit, a dynamic image modeling unit, and a biosignal quality assessment unit;
[0086] The sensing unit aligns multimodal data from EEG signal acquisition equipment, heart rate sensors, voice acquisition devices and gesture capture devices through a heterogeneous data interface protocol, where the multimodal data includes EEG signals, heart rate, voice and gestures; a dual discriminator is used to generate an adversarial network to complete the missing signals when the sensor is abnormal. The missing signals during abnormalities include but are not limited to: EEG signal loss due to EEG electrode detachment, heart rate data interruption due to heart rate sensor disconnection, voice signal loss due to voice acquisition device failure, and gesture data loss due to occlusion or loss of focus of gesture capture device; the generator uses a U-Net deep neural network architecture to reconstruct and complete the missing multimodal signals to generate a complete completed signal. The generator belongs to the signal generation model in the generative adversarial network; the dual discriminator verifies the signal integrity and physiological rationality respectively, where: the first discriminator is used to check the time continuity, consistency and modal matching degree of the signal to ensure the integrity of the completed signal; the second discriminator is used to evaluate the physiological rationality of the completed signal, that is, whether it conforms to the physiological parameter distribution and the normal pattern; the output multimodal signal stream has a time synchronization error of less than 1ms;
[0087] The signal processing unit uses Kalman filtering and wavelet packet decomposition to eliminate baseline drift and electromyographic interference problems that occur when supplementing abnormal missing sensor signals, and extracts energy features in the α, β, and θ bands. The improved SimCLR contrastive learning framework is used to learn potential representations from cross-modal data, where positive samples are paired EEG and heart rate signals of the same user, and negative samples are randomly paired data across users. The loss function uses normalized temperature-scaled cross entropy and outputs a 128-dimensional standardized feature vector. The contrastive learning loss function is shown in the formula:
[0088]
[0089] in, is the contrastive learning loss function, is the embedding vector of the positive sample pair, is the temperature parameter, is the number of negative samples in the batch, is the cosine similarity function, which calculates the similarity between two vectors;
[0090] The dynamic portrait modeling unit constructs user portraits based on a semi-supervised spatiotemporal graph network. The nodes include real users and 200 virtual users generated daily. Virtual users are generated through latent variable interpolation of a variational graph autoencoder with a KL divergence of less than 0.1. Edge weights are defined by behavioral similarity, and a pseudo-label propagation mechanism is used to integrate key events actively annotated by users to output dynamic short-term states and long-term behavioral preference maps. The model of virtual users generated by latent variable interpolation of the variational graph autoencoder is shown in the following formula:
[0091]
[0092] in, Evidence lower bounds for VGAE generation models, is the encoder network, which is used to map the input data X to the latent variable z; is the decoder network, used to reconstruct data X from the latent variable z; is the prior distribution of latent variables, for Divergence, which measures the difference between the encoder distribution and the prior distribution;
[0093] It should be noted that in this process, the nodes represent user instances, including real user nodes and virtual user nodes generated by variational graph autoencoders; real user nodes are directly derived from the behavioral data of actual users, while virtual user nodes are generated through the VGAE model, and the latent variable interpolation method is used to ensure that the generated virtual users are similar in behavior to real users; 200 virtual user nodes are generated every day. During the generation process, the KL divergence is controlled to be less than 0.1 to ensure that the distribution of latent variables is close to the prior distribution.
[0094] These nodes, as the basic elements in the spatiotemporal graph network, carry the behavioral characteristics of each user. The edges between nodes are defined by the behavioral similarity between users, reflecting the similarity of different users in certain behaviors. Behavioral similarity is calculated by analyzing the user's interaction history, preferences, and other relevant features, thereby establishing a connection between each pair of similar users in the graph.
[0095] The physiological signal quality assessment unit detects the sensor contact status in real time. When the EEG signal electrode impedance is greater than 50kΩ, a warning is triggered. After the abnormality persists for 10 seconds, the adversarial network completion mode is activated, and a voice prompt "Sensor contact is poor, completion mode has been enabled" is pushed simultaneously; a three-level warning of the device health status is output, including normal, warning and fault.
[0096] The CI decision module includes a data fusion unit, a consensus unit, a multi-objective collaborative optimization unit, a federated learning scheduling unit, and a conflict resolution engine unit;
[0097] The data fusion unit uses the attention mechanism to dynamically weight the integrated data, integrating the physiological feature vectors from the BI module and the external device data, eliminating redundant information, and generating a spatiotemporally aligned fusion feature tensor;
[0098] The consensus unit dynamically assigns decision weights based on the Shapley value in game theory and calculates the historical accuracy of each agent, which includes an EEG signal acquisition device, a heart rate sensor, a voice acquisition device, and a gesture capture device. It generates an initial set of command candidates through weighted voting, introduces a conflict detection mechanism to ensure decision consistency across devices, and outputs a prioritized command sequence.
[0099] Based on the prioritized instruction sequence generated by the consensus unit, the multi-objective collaborative optimization unit uses an improved NSGA-III algorithm for multi-objective optimization. Taking into account energy consumption, response latency, and user satisfaction, the multi-objective collaborative optimization unit solves for the optimal Pareto frontier solution set. The Pareto frontier solution set represents the optimal instruction plan that strikes a balance between different optimization objectives, ensuring that the final decision meets both system performance requirements and user experience. The optimized instruction set is converted into a natural language description through the T5 model, outputting a human-understandable optimized instruction set for system execution and adjustment.
[0100] The federated learning scheduling unit adopts a hierarchical federated architecture, sets personalized training goals based on an optimized instruction set, compresses model parameters through differentiated knowledge distillation, and dynamically adjusts the communication threshold for model synchronization based on the model update amplitude, computing resource usage, and network bandwidth. A fairness regularization term is added during edge aggregation to adjust the balance between the personalized model and the global model, and finally outputs the updated global model. The loss function is shown in the formula:
[0101]
[0102] in, is the federal fairness constraint loss, Loss of main mission, is the fairness regularization term, is the fairness weight coefficient;
[0103] The conflict resolution engine unit takes the global optimization model and optimized instruction set output by the federated learning scheduling unit as input, matches similar historical scenarios based on case reasoning, uses Sentence-BERT to encode the instruction context, and retrieves the top three historical solutions for user reference; if the match fails, the reinforcement learning unit is called to generate a new strategy to ensure the consistency and executableness of the final instructions, and outputs the final execution instructions after conflict resolution.
[0104] The AI execution feedback module includes a reinforcement learning unit, an execution unit, a feedback unit, and an explanation unit;
[0105] The reinforcement learning unit pre-trains a cross-scenario policy network based on a meta-reinforcement learning framework, inputs the instruction set and environment data optimized by the CI decision module, and uses the Unity simulator to generate diverse control scenario data. Through model-independent meta-learning, it quickly adapts to new tasks within 10 gradient updates, dynamically balances user manual intervention and autonomous decision-making, and outputs a real-time control policy, as shown in the formula:
[0106]
[0107] in, are the updated model parameters, are the initial model parameters, β is the meta-learning rate, is the task loss function, is the parameter change of one gradient update;
[0108] The execution unit uses the predictive execution model LSTM timing prediction, combined with resource-aware scheduling for optimized execution. It inputs the Pareto optimization instruction set generated by the CI module, preloads context-related actions, and dynamically adjusts the model complexity based on the device computing power to output physical control signals. The specific formula is as follows:
[0109]
[0110] in, For predictive execution loss function, For the actual execution results, is the model prediction result, is the L1 regularization weight, are model parameters;
[0111] The feedback unit uses multimodal generative feedback based on the physical control signals from the actuator to provide a personalized interactive experience. Stable Diffusion generates guidance animations and voice prompts, and the CLIP model is used to match user status, analyzing physiological signals, interaction history, and environmental parameters. The feedback mechanism adopts a real-time and offline two-layer architecture: the real-time layer calculates key feature weights and provides immediate feedback; the offline layer constructs a decision logic tree, analyzes long-term interaction patterns, and provides users with explainable interaction content.
[0112] The explanation unit adopts a dynamic explainability framework to analyze the explainability of model decisions based on the physical control signals of the execution unit, the user interaction data of the feedback unit, and the system's historical decision logs. The online layer locates the key decision feature areas in the input data through gradient-weighted class activation mapping and provides real-time explanations. The offline layer combines causal reasoning methods to construct a causal graph model, which supports users to trace back any historical decision path, track the system's decision-making process in different scenarios, and output visual reports and fairness audit logs.
[0113] Dynamic learning and cost optimization module combines learning unit, model compression unit, and adaptive unit;
[0114] The federated learning unit adopts a hierarchical federated architecture, deploys a lightweight student model on the terminal device, and aggregates the gradient of the teacher model on the edge server for update, ensuring a balance between privacy protection and computational efficiency. The unit uses differential knowledge distillation technology to compress the model and reduce the amount of communication data. At the same time, it uses the Paillier homomorphic encryption algorithm to protect the privacy security of gradient transmission, ensuring that the accuracy of the global model does not decay by more than 2% within the monthly update cycle. It outputs a federated model version that is adapted to multiple devices. The loss function is shown in the formula:
[0115]
[0116] in, is the knowledge distillation loss function, is the temperature parameter, is the probability distribution output by the student model, is the probability distribution output by the teacher model, is the distillation loss weight;
[0117] The model compression unit dynamically adjusts the model structure based on a resource-aware mechanism, inputs real-time device status data, and generates lightweight sub-models through knowledge distillation technology. It uses elastic scaling to dynamically adjust the channel width of MobileNetV3 to ensure that the compressed model size is less than or equal to 5MB, while also increasing inference speed by 3 times. The final output is a model library adapted to different device resources.
[0118] The adaptive unit designs a multi-objective optimization strategy selector. Based on the task type and device resource status, it dynamically matches the optimal model compression strategy through a random forest classifier and outputs real-time configuration instructions, reducing system response delay by 40%.
[0119] The algorithm bias and fairness monitoring module includes a bias detection unit, a privacy protection unit, a bias correction unit, and a fairness unit;
[0120] The bias detection unit monitors the decision data flow of the CI module in real time, uses a causal inference model to perform counterfactual fairness analysis, and identifies potential discrimination patterns through structural equation fit. It inputs the instruction set and user attribute labels generated by the CI module, calculates the bias risk score, and triggers a fairness alarm when the score exceeds 0.6. This model is validated based on the UCI Adult dataset and is shown in the following formula:
[0121]
[0122] in, For sensitive attributes, is the counterfactual decision result when the sensitive attribute is intervened by k, It is a non-sensitive feature;
[0123] The privacy protection unit builds a triple protection chain, specifically: data collection: injecting differential privacy noise to prevent user identity leakage; transmission: using Paillier homomorphic encryption to protect gradient transmission; storage: combining blockchain technology to record audit logs to ensure data cannot be tampered with; input is raw biometric signals and user attributes, and output is a desensitized training dataset for subsequent fairness optimization;
[0124] The bias correction unit uses a dual-channel adversarial fairness constraint to generate counterfactual samples through the GAN generator. While optimizing the main task loss in the discriminator, it also introduces fairness loss to reduce system bias. During federated learning aggregation, a fairness regularization term is added to ensure fairness when optimizing the global model. The input is biased model parameters, and the output is an unbiased model, as shown in the formula:
[0125]
[0126] in, is an unbiased model, G is the generator, D is the discriminator, is the fairness constraint weight, Loss of fairness;
[0127] It should be noted that the federated learning scheduling unit is responsible for coordinating and managing the local model training process of each terminal device. It ensures that each device performs adaptive training according to its resource status by setting personalized training goals, optimizing knowledge distillation, and dynamically adjusting the communication threshold for model synchronization. After the local model training is completed, these devices will transmit their model updates to the central server, and the central server will perform the aggregation process, that is, merging the local model updates into a global model through weighted averaging and other methods. This process is called federated learning aggregation.
[0128] The fairness unit continuously monitors system fairness and provides feedback and corrections to potential problems by tracking four core indicators: group fairness, ensuring fairness between different groups; individual fairness, ensuring that similar users receive similar decisions; equal opportunity, ensuring that different groups receive the same positive decision-making opportunities; causal fairness, analyzing potential biases in system decisions based on causal reasoning; using a dynamic dashboard to display real-time fairness monitoring results; and inputting historical decision logs and user feedback to output compliance reports.
[0129] The user interface and interaction module includes a visualization unit, a user feedback unit, and a collaborative learning unit;
[0130] The visualization unit dynamically presents system decisions and feedback through a multimodal fusion interface, integrates StableDiffusion to generate personalized AR guidance animations, and combines the CLIP model to match user physiological characteristics with the knowledge base to display key decision-making basis in real time. The input is the layered interpretation results of the AI module and the user's real-time status, and the output is multimodal interactive content, which improves user operation efficiency by 55%.
[0131] The user feedback unit is designed with a two-way collaborative learning mechanism to optimize the interactive experience through reinforcement learning and prediction of user preferences. After the user manually corrects the data, the PPO algorithm is used to reversely optimize the reinforcement learning strategy. Based on LSTM, the user preference migration trend is predicted, and touch / voice commands and physiological signals are input to dynamically adjust the user preference model.
[0132] The collaborative learning unit builds a human-machine co-adaptive decision-making framework, identifies the causal effects of user correction behavior through causal reinforcement learning, and uses a dynamic reward function to balance system goals and user preferences. The input is historical interaction logs and real-time feedback, and the output is a causal strategy map. The causal reward function and dynamic reward function are shown in the formula:
[0133]
[0134] in, is the causal reward function, Actions automatically performed by the system, is the action manually corrected by the user, sim is the action cosine similarity, and N is the number of historical interactions;
[0135]
[0136] in, is the weight matrix, is the hidden state at the previous moment, is the current input feature, and b is the bias term.
[0137] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A human-machine collaborative intelligent control system based on AIGC, characterized in that: include: BI input module, which realizes human-computer collaborative input through multimodal biosignal perception and dynamic modeling, and provides real-time user status data support; A CI decision module, which is used to integrate data from the BI input module, perform multi-objective optimization and federated learning scheduling, and achieve intelligent decision-making across devices and users; An AI execution feedback module that executes instructions based on the optimized decisions of the CI decision module and provides an adaptive execution solution based on reinforcement learning and multimodal feedback; A dynamic learning and cost optimization module, which is used to optimize model training and deployment, ensuring a balance between system computing efficiency and privacy protection; Algorithm bias and fairness monitoring module, which is used to monitor system algorithm bias and ensure fairness; User interface and interaction module, which are used to enhance user experience and optimize human-computer interaction; The BI input module includes a sensing unit, a signal processing unit, a dynamic image modeling unit and a physiological signal quality assessment unit; The sensing unit aligns multimodal data from an EEG signal acquisition device, a heart rate sensor, a voice acquisition device, and a gesture capture device through a heterogeneous data interface protocol, wherein the multimodal data includes EEG signals, heart rate, voice, gestures, and external device data; The physiological feature vector is obtained by analyzing the EEG signal and calculating the heart rate signal; A dual-discriminator generative adversarial network is used to complete the sensor's abnormal missing signals; The generator uses the U-Net deep neural network architecture to reconstruct and complete missing multimodal signals and generate complete complement signals. The dual discriminators verify signal integrity and physiological plausibility respectively, and output a multimodal signal stream with a time synchronization error of less than 1ms. The signal processing unit uses Kalman filtering and wavelet packet decomposition to eliminate baseline drift and electromyographic interference problems that occur when the sensor abnormally lacks signals, and extracts energy characteristics of α, β, and θ bands; The improved SimCLR contrastive learning framework is used to learn latent representations from cross-modal data, where positive samples are paired data of EEG signals and heart rate signals of the same user, and negative samples are randomly paired data across users. The loss function uses normalized temperature-scaled cross entropy and outputs a 128-dimensional standardized feature vector. The contrastive learning loss function is specifically shown as follows: ,in, is the contrastive learning loss function, is the embedding vector of the positive sample pair, is the temperature parameter, is the number of negative samples in the batch, is the cosine similarity function, which calculates the similarity between two vectors; The dynamic portrait modeling unit constructs user portraits based on a semi-supervised spatiotemporal graph network, where nodes include real users and 200 virtual users generated daily. Virtual users are generated through latent variable interpolation of a variational graph autoencoder, with a KL divergence of less than 0.
1. Edge weights are defined by behavioral similarity, and a pseudo-label propagation mechanism is used to fuse user-actively labeled events to output dynamic short-term states and long-term behavioral preference maps. The model of virtual users generated by latent variable interpolation of the variational graph autoencoder is specifically shown in the formula: ,in, Evidence lower bounds for VGAE generation models, is the encoder network, which is used to map the input data X to the latent variable z; is the decoder network, used to reconstruct data X from the latent variable z; is the prior distribution of latent variables, for Divergence, which measures the difference between the encoder distribution and the prior distribution; The physiological signal quality assessment unit detects the sensor contact status in real time and triggers a warning when the EEG signal electrode impedance is greater than 50kΩ. After the abnormality persists for 10 seconds, the adversarial network completion mode is activated, and a voice prompt "Sensor contact is poor, completion mode has been enabled" is pushed simultaneously; and a three-level warning of the device health status is output, including normal, warning and fault.
2. The AIGC-based human-machine collaborative intelligent control system according to claim 1, characterized in that: The CI decision module includes a data fusion unit, a consensus unit, a multi-objective collaborative optimization unit, a federated learning scheduling unit and a conflict resolution engine unit; The data fusion unit uses an attention mechanism to dynamically weight the integrated data, which comes from the physiological feature vectors of the BI module and the external device data, eliminating redundant information and generating a spatiotemporally aligned fusion feature tensor; The consensus unit dynamically assigns decision weights based on Shapley values in game theory and calculates the historical accuracy of each agent, which includes an EEG signal acquisition device, a heart rate sensor, a voice acquisition device, and a gesture capture device. It generates an initial set of command candidates through weighted voting, introduces a conflict detection mechanism to ensure cross-device decision consistency, and outputs a prioritized command sequence. The multi-objective collaborative optimization unit uses an improved NSGA-III algorithm to perform multi-objective optimization based on the priority-ranked instruction sequence generated by the consensus unit, comprehensively considering energy consumption, response delay, and user satisfaction to solve the optimal Pareto front solution set. The Pareto front solution set represents the optimal instruction solution that strikes a balance between different optimization objectives, ensuring that the final decision meets both system performance requirements and user experience. The optimized instruction set is converted into a natural language description through the T5 model, and a human-understandable optimized instruction set is output for system execution and adjustment. The federated learning scheduling unit adopts a hierarchical federated architecture, sets personalized training goals based on an optimized instruction set, compresses model parameters through differentiated knowledge distillation, and dynamically adjusts the communication threshold for model synchronization based on the model update amplitude, computing resource usage, and network bandwidth. A fairness regularization term is added during edge aggregation to adjust the balance between the personalized model and the global model, and finally outputs the updated global model. The loss function is shown in the formula: ,in, is the federal fairness constraint loss, Loss of main mission, is the fairness regularization term, is the fairness weight coefficient; The conflict resolution engine unit takes the global optimization model and optimization instruction set output by the federated learning scheduling unit as input, matches historical similar scenarios based on case-based reasoning, uses Sentence-BERT to encode the instruction context, and retrieves the top three historical solutions for user reference; If the matching fails, the reinforcement learning unit is called to generate a new strategy to ensure the consistency and executability of the final instructions, and output the final execution instructions after the conflict is resolved.
3. The AIGC-based human-machine collaborative intelligent control system according to claim 1, characterized in that: The AI execution feedback module includes a reinforcement learning unit, an execution unit, a feedback unit, and an interpretation unit; The reinforcement learning unit pre-trains a cross-scenario policy network based on a meta-reinforcement learning framework, inputs the instruction set and environment data optimized by the CI decision module, and uses the Unity simulator to generate diverse control scenario data. Through model-independent meta-learning, it quickly adapts to new tasks within 10 gradient updates, dynamically balances user manual intervention and autonomous decision-making, and outputs a real-time control strategy, as shown in the formula: ,in, are the updated model parameters, are the initial model parameters, β is the meta-learning rate, is the task loss function, is the parameter change of one gradient update; The execution unit uses the predictive execution model LSTM timing prediction, combined with resource-aware scheduling for optimized execution, inputs the Pareto optimization instruction set generated by the CI module, preloads context-related actions, and dynamically adjusts the model complexity based on the device computing power to output physical control signals, as shown in the formula: ,in, For predictive execution loss function, For the actual execution results, is the model prediction result, is the L1 regularization weight, are model parameters; The feedback unit uses multimodal generative feedback based on the physical control signals of the actuator to provide a personalized interactive experience. Stable Diffusion is used to generate guidance animations and voice prompts, and the CLIP model is used for user state matching, analyzing physiological signals, interaction history, and environmental parameters. The feedback mechanism adopts a real-time and offline two-layer architecture: the real-time layer calculates key feature weights and provides immediate feedback; the offline layer constructs a decision logic tree, analyzes long-term interaction patterns, and provides users with interpretable interaction content. The explanation unit adopts a dynamic explainability framework to analyze the explainability of model decisions based on the physical control signals of the execution unit, the user interaction data of the feedback unit and the system's historical decision logs. The online layer locates key decision feature areas in the input data through gradient-weighted class activation mapping and provides real-time explanations. The offline layer combines causal reasoning methods to construct a causal graph model, which supports users to trace back any historical decision path, track the system's decision-making process in different scenarios, and output visual reports and fairness audit logs.
4. The AIGC-based human-machine collaborative intelligent control system according to claim 1, characterized in that: The dynamic learning and cost optimization module combines the learning unit, the model compression unit, and the adaptive unit; The joint learning unit adopts a hierarchical federated architecture, deploys a lightweight student model on the terminal device, and aggregates the gradient of the teacher model on the edge server for update, ensuring a balance between privacy protection and computational efficiency. The unit uses differential knowledge distillation technology to compress the model and reduce the amount of communication data. At the same time, it uses the Paillier homomorphic encryption algorithm to protect the privacy security of gradient transmission, ensuring that the accuracy of the global model does not decay by more than 2% within the monthly update cycle. It outputs a federated model version adapted to multiple devices. The loss function is shown in the formula: ,in, is the knowledge distillation loss function, is the temperature parameter, is the probability distribution output by the student model, is the probability distribution output by the teacher model, is the distillation loss weight; The model compression unit dynamically adjusts the model structure based on a resource-aware mechanism, inputs real-time device status data, and generates lightweight sub-models through knowledge distillation technology. It uses elastic scaling to dynamically adjust the channel width of MobileNetV3 to ensure that the compressed model size is less than or equal to 5MB, while also increasing inference speed by 3 times. The final output is a model library adapted to different device resources. The adaptive unit designs a multi-objective optimization strategy selector, which dynamically matches the optimal model compression strategy through a random forest classifier according to the task type and device resource status, and outputs real-time configuration instructions, reducing system response delay by 40%.
5. The AIGC-based human-machine collaborative intelligent control system according to claim 1, characterized in that: The algorithm bias and fairness monitoring module includes a bias detection unit, a privacy protection unit, a bias correction unit, and a fairness unit; The bias detection unit monitors the decision data stream of the CI module in real time, uses a causal inference model to perform counterfactual fairness analysis, and identifies potential discrimination patterns through structural equation fit. It inputs the instruction set and user attribute labels generated by the CI module, calculates the bias risk score, and triggers a fairness alarm when the score exceeds 0.
6. This model is validated based on the UCI Adult dataset and is shown in the following formula: ,in, For sensitive attributes, is the counterfactual decision result when the sensitive attribute is intervened by k, It is a non-sensitive feature; The privacy protection unit builds a triple protection chain, specifically: data collection end: injects differential privacy noise to prevent user identity leakage; transmission end: uses Paillier homomorphic encryption to protect gradient transmission; storage end: combines blockchain technology to record audit logs to ensure that data cannot be tampered with; input is raw biometric signals and user attributes, and output is a desensitized training dataset for subsequent fairness optimization; The bias correction unit adopts a dual-channel adversarial fairness constraint, generates counterfactual samples through the GAN generator, and introduces fairness loss while optimizing the main task loss in the discriminator to reduce system bias. During federated learning aggregation, fairness regularization terms are added to ensure fairness when optimizing the global model. The input is biased model parameters, and the output is an unbiased model, as shown in the formula: ,in, is an unbiased model, G is the generator, D is the discriminator, is the fairness constraint weight, Loss of fairness; The fairness unit continuously monitors system fairness and provides feedback and corrections to potential problems by tracking four core indicators: group fairness, ensuring fairness between different groups; individual fairness, ensuring that similar users receive similar decisions; equal opportunity, ensuring that different groups receive the same positive decision-making opportunities; causal fairness, analyzing potential biases in system decisions based on causal reasoning; using a dynamic dashboard to display real-time fairness monitoring results; and inputting historical decision logs and user feedback to output compliance reports.
6. The AIGC-based human-machine collaborative intelligent control system according to claim 1, characterized in that: The user interface and interaction module includes a visualization unit, a user feedback unit and a collaborative learning unit; The visualization unit dynamically presents system decisions and feedback through a multimodal fusion interface, integrates StableDiffusion to generate personalized AR guidance animations, and combines the CLIP model to match user physiological characteristics with the knowledge base to display key decision-making basis in real time. The input is the layered interpretation results of the AI module and the user's real-time status, and the output is multimodal interactive content, which can improve user operation efficiency by 55%. The user feedback unit is designed with a two-way collaborative learning mechanism to optimize the interactive experience through reinforcement learning and prediction of user preferences. After the user manually corrects the data, the PPO algorithm is used to reversely optimize the reinforcement learning strategy. Based on LSTM, the user preference migration trend is predicted, and touch / voice commands and physiological signals are input to dynamically adjust the user preference model. The collaborative learning unit builds a human-machine co-adaptive decision-making framework, identifies the causal effects of user correction behavior through causal reinforcement learning, and uses a dynamic reward function to balance system goals and user preferences. The input is historical interaction logs and real-time feedback, and the output is a causal strategy map. The causal reward function and dynamic reward function are specifically shown as follows: ,in, is the causal reward function, Actions automatically performed by the system, is the action manually corrected by the user, sim is the action cosine similarity, and N is the number of historical interactions; ,in, is the weight matrix, is the hidden state at the previous moment, is the current input feature, and b is the bias term.
Citation Information
Patent Citations
Dynamic fair federal learning method and device based on reinforcement learning
CN117273119A
Human-machine cooperation intelligent control method and system based on AIGC and storage medium
CN119024723A