Man-machine cooperation intelligent control system based on AIGC
By adopting multimodal biological signal completion, game theory optimization decision-making, meta-reinforcement learning execution, adaptive federated learning and causal inference fairness monitoring in the human-machine collaborative intelligent control system, the system's instability in the input of multimodal biological signal, poor consistency of group intelligent decision-making, lack of adaptive ability in the artificial intelligence execution feedback, and deviation and fairness of the algorithm, achieving efficient, interpretable, and optimized intelligent control effects.
Patent Information
- Application Number
- CN202510423767.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The existing human-machine collaborative intelligent control system has unstable input of multimodal biological signals, poor consistency of group intelligent decision-making, lack of adaptive ability to perform feedback on artificial intelligence, and algorithms have problems with deviations and fairness.
Using technologies such as multimodal biological signal completion, game theory optimization decision-making, meta-reinforcement learning execution, adaptive federated learning and causal inference fairness monitoring, we will build an efficient, interpretable and optimized intelligent control framework.
Through multimodal fusion, ensure the stability and integrity of user status data input, improve decision-making consistency across devices and across users, enhance the execution and feedback capabilities of AI, ensure algorithm fairness and user privacy protection, and improve the system's personalized adaptability and user experience.
Smart Images

Figure CN119940425A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an AIGC-based human-machine collaborative intelligent control system. Background Art
[0002] With the rapid development of artificial intelligence technology, human-machine collaborative intelligent control systems have been widely used in the fields of autonomous driving, smart homes, medical assistance, etc. Such systems usually rely on biological intelligence to provide user input, group intelligence to optimize decisions, and combine artificial intelligence to perform tasks to achieve multi-level human-machine interaction. However, existing human-machine collaborative systems still face many challenges, especially in cross-modal data fusion, personalized learning, and decision fairness.
[0003] The current intelligent control systems have the following main problems: the quality of multimodal biosignal data is limited, and signal loss and noise influence lead to unstable input; the weight distribution of group intelligent decision-making is unreasonable, and it is difficult to ensure the consistency of decision-making across devices; artificial intelligence execution feedback lacks a dynamic learning mechanism and is difficult to adapt to changes in user needs; the algorithm has bias and fairness issues, which affect the user experience of different user groups; the system interaction mode is relatively fixed, lacking personalized optimization and human-machine co-adaptation capabilities. These problems restrict the widespread application of human-machine collaborative intelligent systems in complex environments.
[0004] In response to the above problems, the present invention proposes a human-machine collaborative intelligent control system based on AIGC, which adopts multimodal biological signal completion, game theory optimization decision-making, meta-reinforcement learning execution, adaptive federated learning and causal inference fairness monitoring to build an efficient, explainable and optimizable intelligent control framework. Summary of the invention
[0005] In response to the above problems, the present invention provides a human-machine collaborative intelligent control system based on AIGC to solve the problems of the prior art in terms of unstable multimodal biological signal input, poor consistency of group intelligent decision-making, lack of adaptive ability of artificial intelligence execution feedback, and bias and fairness of the algorithm.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a human-machine collaborative intelligent control system based on AIGC, comprising:
[0007] BI input module, which realizes human-machine collaborative input through multimodal bio-signal perception and dynamic modeling, and provides real-time user status data support;
[0008] A CI decision module, which is used to integrate the data of the BI input module, perform multi-objective optimization and federated learning scheduling, and realize intelligent decision-making across devices and users;
[0009] An AI execution feedback module, which executes instructions according to the optimized decisions of the CI decision module and provides an adaptive execution solution based on reinforcement learning and multimodal feedback;
[0010] A dynamic learning and cost optimization module, which is used to optimize model training and deployment to ensure a balance between system computing efficiency and privacy protection;
[0011] An algorithm deviation and fairness monitoring module, which is used to monitor system algorithm deviation and ensure fairness;
[0012] User interface and interaction module, the user interface and interaction module are used to enhance user experience and optimize human-computer interaction.
[0013] The BI input module includes a sensing unit, a signal processing unit, a dynamic image modeling unit and a physiological signal quality assessment unit;
[0014] The sensing unit aligns multimodal data from an EEG signal acquisition device, a heart rate sensor, a voice acquisition device, and a gesture capture device through a heterogeneous data interface protocol, wherein the multimodal data includes EEG signals, heart rate, voice, gestures, and external device data; a physiological feature vector is obtained by analyzing EEG signals and calculating heart rate signals; a dual discriminator is used to generate an adversarial network to complete the abnormal missing signal of the sensor; the generator uses a U-Net deep neural network architecture to reconstruct and complete the missing multimodal signal to generate a complete completed signal; the dual discriminator verifies the signal integrity and physiological rationality respectively, and outputs a multimodal signal stream with a time synchronization error of less than 1ms;
[0015] The signal processing unit uses Kalman filtering and wavelet packet decomposition to eliminate the baseline drift and electromyographic interference problems that occur when the sensor abnormal missing signal is completed, and extracts the energy features of the α, β, and θ bands; the potential representation is learned from the cross-modal data through the improved SimCLR contrast learning framework, where the positive sample is the paired data of the EEG signal and the heart rate signal of the same user, and the negative sample is the random paired data across users. The loss function uses the normalized temperature scale cross entropy and outputs a 128-dimensional standardized feature vector; the contrast learning loss function is specifically shown in the formula:
[0016]
[0017] in, is the contrastive learning loss function, is the embedding vector of the positive sample pair, is the temperature parameter, is the number of negative samples in the batch, is the cosine similarity function, which calculates the similarity between two vectors;
[0018] The dynamic portrait modeling unit constructs user portraits based on a semi-supervised spatiotemporal graph network, where nodes include real users and 200 virtual users generated daily; virtual users are generated by latent variable interpolation of a variational graph autoencoder, with a KL divergence of less than 0.1; edge weights are defined by behavioral similarity, and a pseudo-label propagation mechanism is used to fuse key events actively annotated by users, outputting dynamic short-term states and long-term behavioral preference maps; the model of virtual users generated by latent variable interpolation of the variational graph autoencoder is specifically shown in the formula:
[0019]
[0020] in, Evidence lower bounds for VGAE generation models, is the encoder network, which is used to map the input data X to the latent variable z; is the decoder network, used to reconstruct data X from latent variables z; is the prior distribution of latent variables, for Divergence, which measures the difference between the encoder distribution and the prior distribution;
[0021] The physiological signal quality assessment unit detects the sensor contact status in real time, triggers a warning when the EEG signal electrode impedance is greater than 50kΩ, starts the adversarial network completion mode after the abnormality lasts for 10 seconds, and simultaneously pushes a voice prompt "Sensor contact is poor, completion mode has been enabled"; outputs three levels of warning of the device health status, including normal, warning and fault.
[0022] The CI decision module includes a data fusion unit, a consensus unit, a multi-objective collaborative optimization unit, a federated learning scheduling unit and a conflict resolution engine unit;
[0023] The data fusion unit uses an attention mechanism to dynamically weight the integrated data, the integrated data comes from the physiological feature vector of the BI module and the external device data, eliminates redundant information, and generates a spatiotemporally aligned fused feature tensor;
[0024] The consensus unit dynamically allocates decision weights based on the Shapley value in game theory, calculates the historical accuracy of each agent, and each agent includes an EEG signal acquisition device, a heart rate sensor, a voice acquisition device, and a gesture capture device; generates an initial instruction candidate set through weighted voting, introduces a conflict detection mechanism to ensure cross-device decision consistency, and outputs a sequence of instructions sorted by priority;
[0025] The multi-objective collaborative optimization unit uses an improved NSGA-III algorithm to perform multi-objective optimization based on the priority-ranked instruction sequence generated by the consensus unit, comprehensively considers energy consumption, response delay and user satisfaction, and solves the optimal Pareto frontier solution set, which represents the optimal instruction scheme that strikes a balance between different optimization objectives, ensuring that the final decision meets both system performance requirements and user experience; the optimized instruction set is converted into a natural language description through the T5 model, and an optimized instruction set that can be understood by humans is output for system execution adjustment;
[0026] The federated learning scheduling unit adopts a hierarchical federated architecture, sets personalized training goals based on the optimized instruction set, compresses model parameters through differentiated knowledge distillation, and dynamically adjusts the communication threshold of model synchronization based on the model update amplitude, computing resource occupancy and network bandwidth. It adds fairness regularization terms during edge aggregation, adjusts the balance between the personalized model and the global model, and finally outputs the updated global model. The loss function is specifically shown in the formula:
[0027]
[0028] in, For the federal fairness constraint loss, Loss of main task, is the fairness regularization term, is the fairness weight coefficient;
[0029] The conflict resolution engine unit takes the global optimization model and optimization instruction set output by the federated learning scheduling unit as input, matches historical similar scenarios based on case reasoning, uses Sentence-BERT to encode the instruction context, and retrieves the top three historical solutions for user reference; if the match fails, the reinforcement learning unit is called to generate a new strategy to ensure the consistency and executability of the final instructions, and outputs the final execution instructions after conflict resolution.
[0030] The AI execution feedback module includes a reinforcement learning unit, an execution unit, a feedback unit, and an interpretation unit;
[0031] The reinforcement learning unit pre-trains the cross-scenario policy network based on the meta-reinforcement learning framework, inputs the instruction set and environmental data optimized by the CI decision module, and uses the Unity simulator to generate diverse control scenario data; through model-independent meta-learning, it quickly adapts to new tasks within 10 gradient updates, dynamically balances user manual intervention and autonomous decision-making, and outputs real-time control strategies, as shown in the formula:
[0032]
[0033] in, are the updated model parameters, is the initial model parameter, β is the meta-learning rate, is the task loss function, is the parameter change of one gradient update;
[0034] The execution unit uses the predictive execution model LSTM timing prediction, combined with resource-aware scheduling for optimized execution, inputs the Pareto optimization instruction set generated by the CI module, preloads context-related actions, and dynamically adjusts the model complexity according to the device computing power to output physical control signals, as shown in the formula:
[0035]
[0036] in, is the loss function for predictive execution, For the actual execution result, For the model prediction results, is the L1 regularization weight, are model parameters;
[0037] The feedback unit adopts multimodal generative feedback based on the physical control signal of the execution unit to provide personalized interactive experience; Stable Diffusion is used to generate guidance animation and voice prompts, and the CLIP model is used for user state matching to analyze physiological signals, interaction history and environmental parameters; the feedback mechanism adopts a real-time-offline two-layer architecture: the real-time layer calculates the key feature weights and provides instant feedback; the offline layer constructs a decision logic tree, analyzes long-term interaction patterns, and provides explainable interaction content for users;
[0038] The explanation unit adopts a dynamic explainability framework to analyze the explainability of model decisions based on the physical control signals of the execution unit, the user interaction data of the feedback unit and the system's historical decision logs. The online layer locates the key decision feature areas in the input data through gradient weighted class activation mapping to provide real-time explanations. The offline layer combines causal reasoning methods to build a causal graph model, which supports users to trace back any historical decision path, track the system's decision-making process in different scenarios, and output visual reports and fairness audit logs.
[0039] The dynamic learning and cost optimization module combines a learning unit, a model compression unit, and an adaptive unit;
[0040] The joint learning unit adopts a hierarchical federated architecture, deploys a lightweight student model on the terminal device, and aggregates the gradient of the teacher model on the edge server for update to ensure a balance between privacy protection and computational efficiency. The unit uses differential knowledge distillation technology to compress the model and reduce the amount of communication data. At the same time, the Paillier homomorphic encryption algorithm is used to protect the privacy security of gradient transmission, ensuring that the accuracy of the global model does not decay by more than 2% within the monthly update cycle, and outputs a federated model version adapted to multiple devices. The loss function is shown in the formula:
[0041]
[0042] in, is the knowledge distillation loss function, is the temperature parameter, is the probability distribution output by the student model, is the probability distribution output by the teacher model, is the distillation loss weight;
[0043] The model compression unit dynamically adjusts the model structure based on the resource perception mechanism, inputs the real-time status data of the device, and generates a lightweight sub-model through knowledge distillation technology; the elastic scaling method is used to dynamically adjust the channel width of MobileNetV3 to ensure that the compressed model size is less than or equal to 5MB, while the inference speed is increased by 3 times, and finally outputs a model library adapted to different device resources;
[0044] The adaptive unit designs a multi-objective optimization strategy selector, which dynamically matches the best model compression strategy through a random forest classifier according to the task type and device resource status, and outputs real-time configuration instructions, thereby reducing the system response delay by 40%.
[0045] The algorithm bias and fairness monitoring module includes a bias detection unit, a privacy protection unit, a bias correction unit, and a fairness unit;
[0046] The bias detection unit monitors the decision data flow of the CI module in real time, uses a causal inference model to perform counterfactual fairness analysis, and identifies potential discrimination patterns through structural equation fit; inputs the instruction set and user attribute labels generated by the CI module, calculates the bias risk score, and triggers a fairness alarm when the score exceeds 0.6; the model is verified based on the UCIAdult dataset, as shown in the formula:
[0047]
[0048] in, For sensitive attributes, is the counterfactual decision result when the sensitive attribute is intervened to k, It is a non-sensitive feature;
[0049] The privacy protection unit constructs a triple protection chain, specifically: data collection end: inject differential privacy noise to avoid user identity leakage; transmission end: use Paillier homomorphic encryption to protect gradient transmission; storage end: combine blockchain technology to record audit logs to ensure that data cannot be tampered with; input is original biological signals and user attributes, and output is a desensitized training data set for subsequent fairness optimization;
[0050] The bias correction unit adopts dual-channel adversarial fairness constraints, generates counterfactual samples through the GAN generator, and introduces fairness loss to reduce system bias while optimizing the main task loss in the discriminator; when aggregating in federated learning, fairness is ensured when optimizing the global model by adding fairness regularization terms; the input is biased model parameters, and the output is an unbiased model, as shown in the formula:
[0051]
[0052] in, is an unbiased model, G is the generator, D is the discriminator, is the fairness constraint weight, Loss of fairness;
[0053] The fairness unit continuously monitors the fairness of the system and provides feedback and corrections to potential problems by tracking four core indicators: group fairness, ensuring fairness between different groups; individual fairness, ensuring that similar users receive similar decisions; equal opportunity, ensuring that different groups receive the same positive decision opportunities; causal fairness, analyzing potential biases in system decisions based on causal reasoning; using a dynamic dashboard to display real-time fairness monitoring results; inputting historical decision logs and user feedback, and outputting compliance reports.
[0054] The user interface and interaction module includes a visualization unit, a user feedback unit and a collaborative learning unit;
[0055] The visualization unit dynamically presents system decisions and feedback through a multimodal fusion interface, integrates StableDiffusion to generate personalized AR guidance animations, and combines the CLIP model to match user physiological characteristics and knowledge bases to display key decision-making basis in real time; the input is the hierarchical interpretation results of the AI module and the real-time status of the user, and the output is multimodal interactive content to improve user operation efficiency by 55%;
[0056] The user feedback unit designs a two-way collaborative learning mechanism to optimize the interactive experience through reinforcement learning + prediction of user preferences; after the user manually corrects the data, the PPO algorithm is used to reversely optimize the reinforcement learning strategy; based on LSTM, the user preference migration trend is predicted, touch / voice commands and physiological signals are input, and the user preference model is dynamically adjusted;
[0057] The collaborative learning unit constructs a human-machine co-adaptive decision-making framework, identifies the causal effect of user correction behavior through causal reinforcement learning, and uses a dynamic reward function to balance system goals and user preferences. The input is historical interaction logs and real-time feedback, and the output is a causal strategy map. The causal reward function and the dynamic reward function are specifically shown as follows:
[0058]
[0059] in, is the causal reward function, Actions automatically performed by the system. is the action manually corrected by the user, sim is the action cosine similarity, and N is the number of historical interactions;
[0060]
[0061] in, is the weight matrix, is the hidden state at the previous moment, is the current input feature, and b is the bias term.
[0062] Compared with the prior art, the present invention has the following beneficial effects:
[0063] The present invention ensures high-integrity user status data input through multimodal fusion. Real-time signal completion and noise reduction make the data more continuous and reliable, avoiding decision lags and misjudgments caused by loss or delay of sensor data, and improving the accuracy and timeliness of the system's perception of user status.
[0064] The present invention constructs a swarm intelligent decision-making module, improves the consistent decision-making ability across devices and users, adopts the attention mechanism to dynamically weighted fuse the data from various sensor devices and external systems, filters out redundant information and aligns spatiotemporal features, allocates the decision weight of each device based on the consensus algorithm of the Shapley value of game theory, and introduces a conflict detection mechanism to ensure the consistency of decision results under multi-device input.
[0065] The present invention greatly enhances the execution and feedback capabilities of AI by introducing reinforcement learning and adaptive feedback mechanisms; the reinforcement learning unit adopts a meta-reinforcement learning framework to pre-train the policy network, enabling it to quickly migrate in different industry scenarios and adapt to new tasks and environmental changes with only about 10 gradient updates; the execution unit uses a predictive model to predict future actions, and dynamically adjusts the model complexity in combination with resource-aware scheduling to adapt to the computing power of the terminal device, thereby ensuring the real-time execution of instructions.
[0066] The present invention strengthens algorithmic fairness and user privacy protection. The bias detection unit performs causal inference analysis on the data stream output by the CI decision module, detects potential discrimination using a counterfactual method, and triggers a fairness alarm when the calculated bias risk score exceeds a threshold. The bias correction unit introduces adversarial fairness constraint training, generates counterfactual samples through GAN, minimizes task loss and fairness loss during model training, reduces algorithmic bias, and outputs a model that eliminates unfairness.
[0067] The present invention has made innovative designs in the user interface and interaction mode, which significantly enhances the personalized adaptation capability and user experience of the system. The visualization unit adopts a multimodal fusion interface to dynamically present the system decision-making process and feedback information, integrates the Stable Diffusion model to generate personalized AR animation guidance in real time, and uses the CLIP model to match the user's physiological characteristics with the knowledge base, intuitively displaying the key basis of AI decision-making. This rich visualization and explanatory content enables users to understand the system behavior more clearly.
[0068] The present invention achieves higher computing efficiency, lower cost and latency by optimizing the system architecture and algorithms in many aspects. The joint learning unit adopts a hierarchical federated learning architecture, deploys lightweight student models on terminal devices, and aggregates teacher model gradients on edge servers, thereby reducing centralized communication bandwidth occupancy. While ensuring user privacy, it compresses model parameters through differentiated knowledge distillation, reduces model size and transmission volume, and ensures that even after long-term distributed training, the accuracy decay of the global model does not exceed 2% within a monthly cycle. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It is understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0070] Figure 1 is a system architecture diagram of the present invention;
[0071] Figure 2is a diagram of the BI input module architecture of the present invention;
[0072] Figure 3 It is the CI decision module architecture diagram of the present invention;
[0073] Figure 4 It is the architecture diagram of the AI execution feedback module of the present invention;
[0074] Figure 5 is a diagram of the architecture of the dynamic learning and cost optimization module of the present invention;
[0075] Figure 6 It is the architecture diagram of the algorithm deviation and fairness monitoring module of the present invention;
[0076] Figure 7 It is a user interface and interaction module architecture diagram of the present invention. DETAILED DESCRIPTION
[0077] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention claimed for protection, but is only for selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.
[0078] Please refer to Figure 1-7 Schematic diagram of a human-machine collaborative intelligent control system based on AIGC provided by an embodiment of the present invention, including:
[0079] BI input module: The BI input module realizes human-machine collaborative input through multimodal bio-signal perception and dynamic modeling, and provides real-time user status data support;
[0080] CI decision module: The CI decision module is used to integrate the data of the BI input module, perform multi-objective optimization and federated learning scheduling, and realize intelligent decision-making across devices and users;
[0081] AI execution feedback module: The AI execution feedback module executes instructions according to the optimized decisions of the CI decision module and provides an adaptive execution solution based on reinforcement learning and multimodal feedback;
[0082] Dynamic learning and cost optimization module: The dynamic learning and cost optimization module is used to optimize model training and deployment to ensure a balance between system computing efficiency and privacy protection;
[0083] Algorithm bias and fairness monitoring module: The algorithm bias and fairness monitoring module is used to monitor system algorithm bias and ensure fairness;
[0084] User interface and interaction module: The user interface and interaction module are used to enhance user experience and optimize human-computer interaction.
[0085] The BI input module includes a sensing unit, a signal processing unit, a dynamic image modeling unit, and a biosignal quality assessment unit;
[0086] The sensing unit aligns multimodal data from EEG signal acquisition equipment, heart rate sensor, voice acquisition device and gesture capture device through a heterogeneous data interface protocol, wherein the multimodal data includes EEG signal, heart rate, voice and gesture; a dual discriminator is used to generate an adversarial network to complete the missing signal when the sensor is abnormal, and the missing signal during the abnormality includes but is not limited to: EEG signal loss due to EEG electrode detachment, heart rate data interruption due to heart rate sensor disconnection, voice signal loss due to voice acquisition device failure, and gesture data loss due to occlusion or loss of focus of gesture capture device; the generator uses the U-Net deep neural network architecture to reconstruct and complete the missing multimodal signal to generate a complete completed signal, and the generator belongs to the signal generation model in the generative adversarial network; the dual discriminator verifies the signal integrity and physiological rationality respectively, wherein: the first discriminator is used to check the time continuity, consistency and modal matching degree of the signal to ensure the integrity of the completed signal; the second discriminator is used to evaluate the physiological rationality of the completed signal, that is, whether it conforms to the physiological parameter distribution and the normal mode; the output time synchronization error is less than 1ms of the multimodal signal stream;
[0087] The signal processing unit uses Kalman filtering and wavelet packet decomposition to eliminate the baseline drift and electromyographic interference problems that occur when filling in the abnormal missing signals of the sensor, and extracts the energy features of the α, β, and θ bands; the potential representation is learned from the cross-modal data through the improved SimCLR contrastive learning framework, where the positive samples are paired data of the EEG signal and heart rate signal of the same user, and the negative samples are randomly paired data across users. The loss function uses normalized temperature scaled cross entropy and outputs a 128-dimensional standardized feature vector; the contrastive learning loss function is specifically shown in the formula:
[0088]
[0089] in, is the contrastive learning loss function, is the embedding vector of the positive sample pair, is the temperature parameter, is the number of negative samples in the batch, is the cosine similarity function, which calculates the similarity between two vectors;
[0090] The dynamic portrait modeling unit builds user portraits based on a semi-supervised spatiotemporal graph network, where nodes include real users and 200 virtual users generated daily. Virtual users are generated by latent variable interpolation of a variational graph autoencoder, with a KL divergence of less than 0.1. Edge weights are defined by behavioral similarity, and a pseudo-label propagation mechanism is used to fuse key events actively annotated by users, outputting dynamic short-term states and long-term behavioral preference maps. The model of virtual users generated by latent variable interpolation of the variational graph autoencoder is specifically shown in the formula:
[0091]
[0092] in, Evidence lower bounds for VGAE generation models, is the encoder network, which is used to map the input data X to the latent variable z; is the decoder network, used to reconstruct data X from latent variables z; is the prior distribution of latent variables, for Divergence, which measures the difference between the encoder distribution and the prior distribution;
[0093] It should be noted that in this process, the nodes represent user instances, including real user nodes and virtual user nodes generated by variational graph autoencoders; the real user nodes are directly derived from the behavioral data of actual users, while the virtual user nodes are generated by the VGAE model, and the latent variable interpolation method is used to ensure that the generated virtual users are similar to real users in behavior; 200 virtual user nodes are generated every day, and during the generation process, the KL divergence is controlled to be less than 0.1 to ensure that the distribution of latent variables is close to the prior distribution.
[0094] As the basic elements in the spatiotemporal graph network, these nodes carry the behavioral characteristics of each user. The edges between nodes are defined by the behavioral similarity between users, reflecting the similarity of different users in certain behaviors. The behavioral similarity is calculated by analyzing the user's interaction history, preferences, and other relevant features, thereby establishing a connection between each pair of similar users in the graph.
[0095] The physiological signal quality assessment unit detects the sensor contact status in real time. When the EEG signal electrode impedance is greater than 50kΩ, a warning is triggered. After the abnormality lasts for 10 seconds, the adversarial network completion mode is started, and a voice prompt "Sensor contact is poor, completion mode has been enabled" is pushed simultaneously; three-level warnings of the output device health status include normal, warning and fault.
[0096] The CI decision-making module includes a data fusion unit, a consensus unit, a multi-objective collaborative optimization unit, a federated learning scheduling unit, and a conflict resolution engine unit;
[0097] The data fusion unit uses the attention mechanism to dynamically weight the integrated data, integrates the physiological feature vectors from the BI module and the external device data, eliminates redundant information, and generates a spatiotemporally aligned fused feature tensor;
[0098] The consensus unit dynamically allocates decision weights based on the Shapley value in game theory and calculates the historical accuracy of each agent, which includes an EEG signal acquisition device, a heart rate sensor, a voice acquisition device, and a gesture capture device. It generates an initial command candidate set through weighted voting, introduces a conflict detection mechanism to ensure cross-device decision consistency, and outputs a priority-ordered command sequence.
[0099] The multi-objective collaborative optimization unit uses the improved NSGA-III algorithm to perform multi-objective optimization based on the priority-ranked instruction sequence generated by the consensus unit, comprehensively considering energy consumption, response delay and user satisfaction, and solving the optimal Pareto frontier solution set. The Pareto frontier solution set represents the optimal instruction scheme that strikes a balance between different optimization goals, ensuring that the final decision meets both system performance requirements and user experience. The optimized instruction set is converted into a natural language description through the T5 model, and an optimized instruction set that can be understood by humans is output for system execution adjustment.
[0100] The federated learning scheduling unit adopts a hierarchical federated architecture, sets personalized training goals based on the optimized instruction set, compresses model parameters through differentiated knowledge distillation, and dynamically adjusts the communication threshold of model synchronization based on the model update amplitude, computing resource usage, and network bandwidth. It adds fairness regularization terms during edge aggregation, adjusts the balance between the personalized model and the global model, and finally outputs the updated global model. The loss function is shown in the formula:
[0101]
[0102] in, For the federal fairness constraint loss, Loss of main task, is the fairness regularization term, is the fairness weight coefficient;
[0103] The conflict resolution engine unit takes the global optimization model and optimized instruction set output by the federated learning scheduling unit as input, matches historical similar scenarios based on case reasoning, uses Sentence-BERT to encode the instruction context, and retrieves the top three historical solutions for user reference; if the match fails, the reinforcement learning unit is called to generate a new strategy to ensure the consistency and executableness of the final instructions, and outputs the final execution instructions after conflict resolution.
[0104] The AI execution feedback module includes a reinforcement learning unit, an execution unit, a feedback unit, and an explanation unit;
[0105] The reinforcement learning unit pre-trains the cross-scenario policy network based on the meta-reinforcement learning framework, inputs the instruction set and environmental data optimized by the CI decision module, and uses the Unity simulator to generate diverse control scenario data; through model-independent meta-learning, it quickly adapts to new tasks within 10 gradient updates, dynamically balances user manual intervention and autonomous decision-making, and outputs real-time control strategies, as shown in the formula:
[0106]
[0107] in, are the updated model parameters, is the initial model parameter, β is the meta-learning rate, is the task loss function, is the parameter change of one gradient update;
[0108] The execution unit uses the predictive execution model LSTM timing prediction, combined with resource-aware scheduling for optimized execution, inputs the Pareto optimization instruction set generated by the CI module, preloads context-related actions, and dynamically adjusts the model complexity according to the device computing power to output physical control signals, as shown in the formula:
[0109]
[0110] in, is the loss function for predictive execution, For the actual execution result, For the model prediction results, is the L1 regularization weight, are model parameters;
[0111] The feedback unit uses multimodal generative feedback based on the physical control signal of the execution unit to provide a personalized interactive experience. Stable Diffusion is used to generate guidance animations and voice prompts. The CLIP model is used for user state matching to analyze physiological signals, interaction history, and environmental parameters. The feedback mechanism adopts a real-time-offline two-layer architecture: the real-time layer calculates the weights of key features and provides instant feedback; the offline layer constructs a decision logic tree, analyzes long-term interaction patterns, and provides users with explainable interaction content.
[0112] The explanation unit adopts a dynamic explainability framework to analyze the explainability of model decisions based on the physical control signals of the execution unit, the user interaction data of the feedback unit, and the system's historical decision logs. The online layer locates the key decision feature areas in the input data through gradient-weighted class activation mapping to provide real-time explanations. The offline layer combines causal reasoning methods to build a causal graph model, which supports users to trace back any historical decision path, track the system's decision-making process in different scenarios, and output visual reports and fairness audit logs.
[0113] Dynamic learning and cost optimization module combines learning unit, model compression unit and adaptive unit;
[0114] The joint learning unit adopts a hierarchical federated architecture, deploys a lightweight student model on the terminal device, and aggregates the gradient of the teacher model on the edge server for update to ensure a balance between privacy protection and computational efficiency. The unit uses differential knowledge distillation technology to compress the model and reduce the amount of communication data. At the same time, the Paillier homomorphic encryption algorithm is used to protect the privacy security of gradient transmission, ensuring that the accuracy of the global model does not decay by more than 2% within the monthly update cycle, and outputs a federated model version that is suitable for multiple devices. The loss function is shown in the formula:
[0115]
[0116] in, is the knowledge distillation loss function, is the temperature parameter, is the probability distribution output by the student model, is the probability distribution output by the teacher model, is the distillation loss weight;
[0117] The model compression unit dynamically adjusts the model structure based on the resource perception mechanism, inputs the real-time status data of the device, and generates a lightweight sub-model through knowledge distillation technology. It uses an elastic scaling method to dynamically adjust the channel width of MobileNetV3 to ensure that the size of the compressed model is less than or equal to 5MB, while increasing the inference speed by 3 times, and finally outputs a model library that adapts to different device resources.
[0118] The adaptive unit designs a multi-objective optimization strategy selector, which dynamically matches the best model compression strategy through a random forest classifier according to the task type and device resource status, and outputs real-time configuration instructions, reducing the system response delay by 40%.
[0119] The algorithm bias and fairness monitoring module includes bias detection unit, privacy protection unit, bias correction unit, and fairness unit;
[0120] The bias detection unit monitors the decision data flow of the CI module in real time, uses a causal inference model to perform counterfactual fairness analysis, and identifies potential discrimination patterns through structural equation fit. It inputs the instruction set and user attribute labels generated by the CI module, calculates the bias risk score, and triggers a fairness alarm when the score exceeds 0.6. The model is verified based on the UCI Adult dataset, as shown in the formula:
[0121]
[0122] in, For sensitive attributes, is the counterfactual decision result when the sensitive attribute is intervened to k, It is a non-sensitive feature;
[0123] The privacy protection unit builds a triple protection chain, specifically: data collection end: inject differential privacy noise to avoid user identity leakage; transmission end: use Paillier homomorphic encryption to protect gradient transmission; storage end: combine blockchain technology to record audit logs to ensure that data cannot be tampered with; input is the original biological signal and user attributes, and the output is the desensitized training data set for subsequent fairness optimization;
[0124] The bias correction unit adopts dual-channel adversarial fairness constraints, generates counterfactual samples through the GAN generator, and introduces fairness loss to reduce system bias while optimizing the main task loss in the discriminator. When aggregating in federated learning, fairness is added to ensure fairness when optimizing the global model. The input is biased model parameters, and the output is an unbiased model, as shown in the formula:
[0125]
[0126] in, is an unbiased model, G is the generator, D is the discriminator, is the fairness constraint weight, Loss of fairness;
[0127] It should be noted that the federated learning scheduling unit is responsible for coordinating and managing the local model training process of each terminal device. It ensures that each device performs adaptive training according to its resource status by setting personalized training goals, optimizing knowledge distillation, and dynamically adjusting the communication threshold of model synchronization. After the local model training is completed, these devices transmit their model updates to the central server, and the central server performs the aggregation process, that is, merging the local model updates into a global model through weighted averaging and other methods. This process is federated learning aggregation.
[0128] The fairness unit continuously monitors system fairness and provides feedback and corrections to potential problems by tracking four core indicators: group fairness, ensuring fairness between different groups; individual fairness, ensuring that similar users receive similar decisions; equal opportunity, ensuring that different groups receive the same positive decision opportunities; causal fairness, analyzing potential biases in system decisions based on causal reasoning; using dynamic dashboards to display real-time fairness monitoring results; inputting historical decision logs and user feedback, and outputting compliance reports.
[0129] The user interface and interaction module includes a visualization unit, a user feedback unit, and a collaborative learning unit;
[0130] The visualization unit dynamically presents system decisions and feedback through a multimodal fusion interface, integrates StableDiffusion to generate personalized AR guidance animations, and combines the CLIP model to match user physiological characteristics and knowledge bases to display key decision-making basis in real time. The input is the hierarchical interpretation results of the AI module and the real-time status of the user, and the output is multimodal interactive content to improve user operation efficiency by 55%;
[0131] The user feedback unit is designed with a two-way collaborative learning mechanism to optimize the interactive experience through reinforcement learning + prediction of user preferences. After the user manually corrects the data, the PPO algorithm is used to reversely optimize the reinforcement learning strategy. Based on LSTM, the user preference migration trend is predicted, and touch / voice commands and physiological signals are input to dynamically adjust the user preference model.
[0132] The collaborative learning unit builds a human-machine co-adaptive decision-making framework, identifies the causal effect of user correction behavior through causal reinforcement learning, and uses a dynamic reward function to balance system goals and user preferences. The input is historical interaction logs and real-time feedback, and the output is a causal strategy map. The causal reward function and dynamic reward function are shown in the formula:
[0133]
[0134] in, is the causal reward function, Actions automatically performed by the system. is the action manually corrected by the user, sim is the action cosine similarity, and N is the number of historical interactions;
[0135]
[0136] in, is the weight matrix, is the hidden state at the previous moment, is the current input feature, and b is the bias term.
[0137] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention has various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A human-machine collaborative intelligent control system based on AIGC, characterized in that: include: BI input module, which realizes human-machine collaborative input through multimodal bio-signal perception and dynamic modeling, and provides real-time user status data support; A CI decision module, which is used to integrate the data of the BI input module, perform multi-objective optimization and federated learning scheduling, and realize intelligent decision-making across devices and users; An AI execution feedback module, which executes instructions according to the optimized decisions of the CI decision module and provides an adaptive execution solution based on reinforcement learning and multimodal feedback; A dynamic learning and cost optimization module, which is used to optimize model training and deployment to ensure a balance between system computing efficiency and privacy protection; An algorithm deviation and fairness monitoring module, which is used to monitor system algorithm deviation and ensure fairness; User interface and interaction module, the user interface and interaction module are used to enhance user experience and optimize human-computer interaction.
2. According to claim 1, the human-machine collaborative intelligent control system based on AIGC is characterized by: The BI input module includes a sensing unit, a signal processing unit, a dynamic image modeling unit and a physiological signal quality assessment unit; The sensing unit aligns multimodal data from an EEG signal acquisition device, a heart rate sensor, a voice acquisition device, and a gesture capture device through a heterogeneous data interface protocol, wherein the multimodal data includes EEG signals, heart rate, voice, gestures, and external device data; The physiological feature vector is obtained by analyzing the EEG signal and calculating the heart rate signal; A dual discriminator is used to generate an adversarial network to complete the abnormal missing signals of the sensor; The generator uses the U-Net deep neural network architecture to reconstruct and complete the missing multimodal signals and generate complete completed signals; the dual discriminators verify the signal integrity and physiological rationality respectively, and output a multimodal signal stream with a time synchronization error of less than 1ms; The signal processing unit uses Kalman filtering and wavelet packet decomposition to eliminate the baseline drift and electromyographic interference problems that occur when the sensor abnormally lacks signals, and extracts the energy characteristics of the α, β, and θ bands; The potential representation is learned from cross-modal data through the improved SimCLR contrastive learning framework, where the positive samples are paired data of the EEG signal and heart rate signal of the same user, and the negative samples are randomly paired data across users. The loss function uses normalized temperature scaled cross entropy and outputs a 128-dimensional standardized feature vector. The contrastive learning loss function is specifically shown in the formula: ,in, is the contrastive learning loss function, is the embedding vector of the positive sample pair, is the temperature parameter, is the number of negative samples in the batch, is the cosine similarity function, which calculates the similarity between two vectors; The dynamic portrait modeling unit constructs user portraits based on a semi-supervised spatiotemporal graph network, where nodes include real users and 200 virtual users generated daily; virtual users are generated by latent variable interpolation of a variational graph autoencoder, with a KL divergence of less than 0.1; edge weights are defined by behavioral similarity, and pseudo-label propagation mechanisms are used to fuse user-actively annotated events, outputting dynamic short-term states and long-term behavioral preference maps; the model of virtual users generated by latent variable interpolation of the variational graph autoencoder is specifically shown in the formula: ,in, Evidence lower bounds for VGAE generation models, is the encoder network, which is used to map the input data X to the latent variable z; is the decoder network, used to reconstruct data X from latent variables z; is the prior distribution of latent variables, for Divergence, which measures the difference between the encoder distribution and the prior distribution; The physiological signal quality assessment unit detects the sensor contact status in real time, triggers a warning when the EEG signal electrode impedance is greater than 50kΩ, starts the adversarial network completion mode after the abnormality lasts for 10 seconds, and simultaneously pushes a voice prompt "Sensor contact is poor, completion mode has been enabled"; outputs three levels of warning of the device health status, including normal, warning and fault.
3. The human-machine collaborative intelligent control system based on AIGC according to claim 1 is characterized in that: The CI decision module includes a data fusion unit, a consensus unit, a multi-objective collaborative optimization unit, a federated learning scheduling unit and a conflict resolution engine unit; The data fusion unit uses an attention mechanism to dynamically weight the integrated data, the integrated data comes from the physiological feature vector of the BI module and the external device data, eliminates redundant information, and generates a spatiotemporally aligned fused feature tensor; The consensus unit dynamically allocates decision weights based on the Shapley value in game theory, calculates the historical accuracy of each agent, and each agent includes an EEG signal acquisition device, a heart rate sensor, a voice acquisition device, and a gesture capture device; generates an initial instruction candidate set through weighted voting, introduces a conflict detection mechanism to ensure cross-device decision consistency, and outputs a sequence of instructions sorted by priority; The multi-objective collaborative optimization unit uses an improved NSGA-III algorithm to perform multi-objective optimization based on the priority-ranked instruction sequence generated by the consensus unit, comprehensively considers energy consumption, response delay and user satisfaction, and solves the optimal Pareto frontier solution set, which represents the optimal instruction scheme that strikes a balance between different optimization objectives, ensuring that the final decision meets both system performance requirements and user experience; the optimized instruction set is converted into a natural language description through the T5 model, and an optimized instruction set that can be understood by humans is output for system execution adjustment; The federated learning scheduling unit adopts a hierarchical federated architecture, sets personalized training goals based on the optimized instruction set, compresses model parameters through differentiated knowledge distillation, and dynamically adjusts the communication threshold of model synchronization based on the model update amplitude, computing resource occupancy and network bandwidth. It adds fairness regularization terms during edge aggregation, adjusts the balance between the personalized model and the global model, and finally outputs the updated global model. The loss function is specifically shown in the formula: ,in, For the federal fairness constraint loss, Loss of main task, is the fairness regularization term, is the fairness weight coefficient; The conflict resolution engine unit takes the global optimization model and optimization instruction set output by the federated learning scheduling unit as input, matches historical similar scenarios based on case reasoning, uses Sentence-BERT to encode the instruction context, and retrieves the top three historical solutions for user reference; If the match fails, the reinforcement learning unit is called to generate a new strategy to ensure the consistency and executability of the final instructions, and output the final execution instructions after the conflict is resolved.
4. The human-machine collaborative intelligent control system based on AIGC according to claim 1 is characterized in that: The AI execution feedback module includes a reinforcement learning unit, an execution unit, a feedback unit, and an interpretation unit; The reinforcement learning unit pre-trains the cross-scenario policy network based on the meta-reinforcement learning framework, inputs the instruction set and environmental data optimized by the CI decision module, and uses the Unity simulator to generate diverse control scenario data; through model-independent meta-learning, it quickly adapts to new tasks within 10 gradient updates, dynamically balances user manual intervention and autonomous decision-making, and outputs real-time control strategies, as shown in the formula: ,in, are the updated model parameters, is the initial model parameter, β is the meta-learning rate, is the task loss function, is the parameter change of one gradient update; The execution unit uses the predictive execution model LSTM timing prediction, combined with resource-aware scheduling for optimized execution, inputs the Pareto optimization instruction set generated by the CI module, preloads context-related actions, and dynamically adjusts the model complexity according to the device computing power to output physical control signals, as shown in the formula: ,in, is the loss function for predictive execution, For the actual execution result, For the model prediction results, is the L1 regularization weight, are model parameters; The feedback unit adopts multimodal generative feedback based on the physical control signal of the execution unit to provide personalized interactive experience; Stable Diffusion is used to generate guidance animation and voice prompts, and the CLIP model is used for user state matching to analyze physiological signals, interaction history and environmental parameters; the feedback mechanism adopts a real-time-offline two-layer architecture: the real-time layer calculates the key feature weights and provides instant feedback; the offline layer constructs a decision logic tree, analyzes long-term interaction patterns, and provides explainable interaction content for users; The explanation unit adopts a dynamic explainability framework to analyze the explainability of model decisions based on the physical control signals of the execution unit, the user interaction data of the feedback unit and the system's historical decision logs. The online layer locates the key decision feature areas in the input data through gradient weighted class activation mapping to provide real-time explanations. The offline layer combines causal reasoning methods to build a causal graph model, which supports users to trace back any historical decision path, track the system's decision-making process in different scenarios, and output visual reports and fairness audit logs.
5. The human-machine collaborative intelligent control system based on AIGC according to claim 1 is characterized in that: The dynamic learning and cost optimization module combines a learning unit, a model compression unit, and an adaptive unit; The joint learning unit adopts a hierarchical federated architecture, deploys a lightweight student model on the terminal device, and aggregates the gradient of the teacher model on the edge server for update to ensure a balance between privacy protection and computational efficiency. The unit uses differential knowledge distillation technology to compress the model and reduce the amount of communication data. At the same time, the Paillier homomorphic encryption algorithm is used to protect the privacy security of gradient transmission, ensuring that the accuracy of the global model does not decay by more than 2% within the monthly update cycle, and outputs a federated model version adapted to multiple devices. The loss function is shown in the formula: ,in, is the knowledge distillation loss function, is the temperature parameter, is the probability distribution output by the student model, is the probability distribution output by the teacher model, is the distillation loss weight; The model compression unit dynamically adjusts the model structure based on the resource perception mechanism, inputs the real-time status data of the device, and generates a lightweight sub-model through knowledge distillation technology; the elastic scaling method is used to dynamically adjust the channel width of MobileNetV3 to ensure that the compressed model size is less than or equal to 5MB, while the inference speed is increased by 3 times, and finally outputs a model library adapted to different device resources; The adaptive unit designs a multi-objective optimization strategy selector, dynamically matches the best model compression strategy through a random forest classifier according to the task type and device resource status, and outputs real-time configuration instructions, thereby reducing the system response delay by 40%.
6. The human-machine collaborative intelligent control system based on AIGC according to claim 1 is characterized in that: The algorithm bias and fairness monitoring module includes a bias detection unit, a privacy protection unit, a bias correction unit, and a fairness unit; The bias detection unit monitors the decision data flow of the CI module in real time, uses a causal inference model to perform counterfactual fairness analysis, and identifies potential discrimination patterns through structural equation fit; inputs the instruction set and user attribute labels generated by the CI module, calculates the bias risk score, and triggers a fairness alarm when the score exceeds 0.6; the model is verified based on the UCI Adult dataset, as shown in the formula: ,in, For sensitive attributes, is the counterfactual decision result when the sensitive attribute is intervened to k, It is a non-sensitive feature; The privacy protection unit constructs a triple protection chain, specifically: data collection end: inject differential privacy noise to avoid user identity leakage; transmission end: use Paillier homomorphic encryption to protect gradient transmission; storage end: combine blockchain technology to record audit logs to ensure that data cannot be tampered with; input is original biological signals and user attributes, and output is a desensitized training data set for subsequent fairness optimization; The bias correction unit adopts dual-channel adversarial fairness constraints, generates counterfactual samples through the GAN generator, and introduces fairness loss to reduce system bias while optimizing the main task loss in the discriminator; when aggregating in federated learning, fairness is ensured when optimizing the global model by adding fairness regularization terms; the input is biased model parameters, and the output is an unbiased model, as shown in the formula: ,in, is an unbiased model, G is the generator, D is the discriminator, is the fairness constraint weight, Loss of fairness; The fairness unit continuously monitors the fairness of the system and provides feedback and corrections to potential problems by tracking four core indicators: group fairness, ensuring fairness between different groups; individual fairness, ensuring that similar users receive similar decisions; equal opportunity, ensuring that different groups receive the same positive decision opportunities; causal fairness, analyzing potential biases in system decisions based on causal reasoning; using a dynamic dashboard to display real-time fairness monitoring results; inputting historical decision logs and user feedback, and outputting compliance reports.
7. The human-machine collaborative intelligent control system based on AIGC according to claim 1 is characterized in that: The user interface and interaction module includes a visualization unit, a user feedback unit and a collaborative learning unit; The visualization unit dynamically presents system decisions and feedback through a multimodal fusion interface, integrates StableDiffusion to generate personalized AR guidance animations, and combines the CLIP model to match user physiological characteristics and knowledge bases to display key decision-making basis in real time; the input is the hierarchical interpretation results of the AI module and the real-time status of the user, and the output is multimodal interactive content to improve user operation efficiency by 55%; The user feedback unit designs a two-way collaborative learning mechanism to optimize the interactive experience through reinforcement learning + prediction of user preferences; after the user manually corrects the data, the PPO algorithm is used to reversely optimize the reinforcement learning strategy; based on LSTM, the user preference migration trend is predicted, touch / voice commands and physiological signals are input, and the user preference model is dynamically adjusted; The collaborative learning unit constructs a human-machine co-adaptive decision-making framework, identifies the causal effect of user correction behavior through causal reinforcement learning, and uses a dynamic reward function to balance system goals and user preferences. The input is historical interaction logs and real-time feedback, and the output is a causal strategy map. The causal reward function and the dynamic reward function are specifically shown as follows: ,in, is the causal reward function, Actions automatically performed by the system. is the action manually corrected by the user, sim is the action cosine similarity, and N is the number of historical interactions; ,in, is the weight matrix, is the hidden state at the previous moment, is the current input feature, and b is the bias term.
Citation Information
Patent Citations
Artificial intelligence management system based on cross-platform data interaction
CN117271981A
Dynamic fair federal learning method and device based on reinforcement learning
CN117273119A
Human-machine cooperation intelligent control method and system based on AIGC and storage medium
CN119024723A
Digital twinning-based multi-modal anomaly detection system
CN119128778A
Intelligent measurement and control device and industrial design application method
CN119150673A
Cited By
Personalized content recommendation and ROI (Region of Interest) improvement method and system based on multi-dimensional user portraits
CN120125294A
Knowledge distillation-based lightweight model transfer learning method
CN120181188A
Cloud networking monitoring video lightweight processing method and system based on U-ACE
CN120388322A
A lightweight processing method and system for cloud-based network surveillance video based on U-ACE
CN120388322B
Unmanned aerial vehicle electric power inspection safety supervision and emergency response system
CN120491665A