Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

67 results about "Strategic learning" patented technology

Three-dimensional perception robot operation knowledge distillation method based on monocular image

The invention relates to the field of robot operation and three-dimensional perception, in particular to a monocular image-based three-dimensional perception robot operation knowledge distillation method, which comprises the following steps of: establishing a strategy learning framework comprising a student model and a teacher model; constructing a strategy prediction model in the strategy learning framework; training the strategy prediction model, wherein a three-level knowledge distillation mechanism is adopted in the training process to complete knowledge migration between the teacher model and the student model; combining the three distillation losses with strategy optimization losses to form a total loss function, and performing end-to-end training on the student model until convergence to obtain a deployable monocular strategy model; the deployable monocular strategy model only retains a student model, and generates a robot operation instruction. The method has the beneficial effects that the robot under monocular RGB input has three-dimensional perception and high-precision operation capabilities while the reasoning efficiency is kept, the task success rate and generalization performance are remarkably improved, and the effectiveness and robustness of the method are verified.
Owner:ZHEJIANG UNIV OF TECH

Complex network disintegration method based on evolution deep reinforcement learning

The invention discloses a complex network disintegration method based on evolution deep reinforcement learning. According to the method, an encoder-decoder model fusing a graph convolutional neural network and a deep Q network is constructed, and is used for efficiently extracting importance features of nodes in a complex network and realizing dynamic decision-making of a node disassembling sequence according to the importance features. In order to optimize model parameters and improve search capability, an evolutionary algorithm is introduced to perform global exploration on the model parameters, and the problem that a directional optimization strategy is easy to fall into local optimum is avoided. Meanwhile, deep mining is carried out on an evolution result in combination with a reinforcement learning strategy, the overall optimization process is accelerated, and advantage complementation of parameter evolution and strategy learning is achieved. Experimental results show that the method significantly improves the efficiency and precision of network disassembly while maintaining the robustness of the model, and has good practical value and wide application prospects.
Owner:NANJING UNIV OF SCI & TECH +2

Multi-modal fusion virtual simulation experiment teaching system

The invention provides a multi-modal fusion virtual simulation experiment teaching system, and relates to the technical field of virtual simulation, and the system comprises an information collection module which is used for obtaining the interaction information of a learner; the information fusion module is used for constructing a multi-modal input data set; the instruction analysis module is used for analyzing and generating an operation instruction sequence; the teaching intervention strategy generation module is used for generating a self-adaptive teaching intervention strategy; the teaching intervention strategy execution module is used for executing a self-adaptive teaching intervention strategy; and the learning effect evaluation module is used for quantifying the learning effect evaluation score. The technical problems that in traditional virtual simulation experiment teaching, multi-modal operation behaviors of learners are difficult to accurately capture in real time, personalized teaching intervention is lacked and learning effect evaluation is incomplete can be solved, and virtual experiment simulation and intelligent teaching intervention driven by multi-modal information fusion are achieved. And the experimental interaction experience and the learning effect evaluation precision are improved.
Owner:CHINA UNIV OF GEOSCIENCES (BEIJING)

Digital twin multi-agent reinforcement learning intelligent decision-making system with secure memory playback mechanism

The invention discloses a digital twinning multi-agent reinforcement learning intelligent decision-making system and method with a secure memory playback mechanism, and the system comprises a digital twinning module which is used for constructing a virtual model and synchronizing the virtual model with a physical entity in real time; the multi-agent reinforcement learning module is used for carrying out strategy learning based on a constrained Markov decision process and balancing performance and safety through a Lagrange multiplier; the safe memory playback module is used for weighting and playing back the experience samples according to the risk and the timeliness so as to improve the learning safety; the reversible grey influence network module is used for causal modeling and reasoning and enhancing decision interpretability; the double-loop self-constraint control module ensures that a control action is always in a physical safety boundary through a barrier function and safety projection; and the convergence and stability criterion module is used for verifying strategy security convergence and system asymptotic stability. According to the method, the problems of strategy border crossing, virtual-real mismatching and the like in the high-risk manufacturing process are solved, and multi-target optimal control under the safety constraint is realized.
Owner:CHONGQING UNIV +1

Incremental reinforcement learning path planning method based on dynamic reward remodeling

The invention provides an incremental reinforcement learning path planning method based on dynamic reward remodeling, which is used for solving the problems of training efficiency and composite optimization in path planning. According to the method, a rapid exploration random tree and an artificial potential field method are combined with reinforcement learning, a dynamic reward function is designed, real-time feedback is provided for an intelligent agent, and strategy learning is accelerated. The system specifically comprises a reward remodeling module which provides process rewards through a fast exploration random tree and an artificial potential field method; the sub-target course learning module is used for gradually increasing task complexity and optimizing sample utilization; the progressive reward adjustment module dynamically adjusts the reward weight and solves the reward coupling problem; and the strategy fine tuning module dynamically adjusts the learning rate and supports strategy optimization. Experimental results show that in static and dynamic path planning tasks, compared with a traditional method, the method has the advantages that the training efficiency and the composite optimization performance are remarkably improved, the reward coupling problem is effectively solved, and the intelligent agent can efficiently learn and converge to a global optimal strategy.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Multi-AUV safe path planning method based on deep reinforcement learning

The invention discloses a multi-AUV safe path planning method based on deep reinforcement learning. According to the method, a deep reinforcement learning algorithm model based on an MATD3 method is adopted, a policy network and a value network are updated by adopting security constraints, the security of policy learning is improved, the expected reward revenue is maximized under the condition that the expected security cost constraints are met, a frequent minimum and maximum optimization process is avoided by adopting first-order penalty optimization, and the security of policy learning is improved. Meanwhile, safety correction based on a safety layer is added in the training process to guarantee safety in the early stage of training, and exploratory and safety balance is brought to strategy optimization of reinforcement learning through safety constraint and safety correction. According to the multi-AUV path planning method, the time cooperation constraint and the space cooperation constraint of multi-AUV path planning are comprehensively considered, the centralized training and decentralized decision-making architecture is applied to multi-AUV path planning, the path planning method capable of ensuring cooperation safety is provided for a multi-AUV system, and the safety and the reliability of path planning are improved.
Owner:HARBIN ENG UNIV

Communication lightweight multi-micro energy network cooperative control method for large model assisted reward recovery

The invention belongs to the technical field of multi-micro-energy-network cooperative control, and particularly relates to a communication lightweight multi-micro-energy-network cooperative control method for reward recovery assisted by a large model, and the method comprises the following steps: S1, constructing a mathematical model of distributed energy in a micro-energy network, and enabling the mathematical model to be used for describing the dynamic behaviors of energy conversion, storage and constraint; s2, modeling a micro energy network cooperative control problem as a partially observable Markov decision process POMDP; s3, solving the POMDP of the step S2 by using an optimized MAPPO method to obtain a cooperative control scheme; and issuing to each micro-energy network terminal for execution. The method can break through the limitation of the communication bottleneck on the premise of ensuring the learning quality of the control strategy.
Owner:CHONGQING UNIV

Federated backdoor defense method based on decoupling contrast learning

The invention discloses a federated backdoor defense method based on decoupling contrast learning, and the method comprises the steps: training a backdoor model based on a backdoor sample, and immediately stopping training after the backdoor model converges on the backdoor sample; respectively extracting a penultimate layer vector of the backdoor model and the local model from a sample pair held by the malicious client as a backdoor feature and a clean feature; comparing and learning the separated back door features and the clean features, and learning the clean features for the local model by using a sample weighting strategy to train the local model to obtain a trained local model; and sending local model parameters of the trained local model to a global server, and generating model parameters of a new global model based on the local model parameters through an aggregation function. The method aims at reducing information dependence between backdoor features and clean features through comparative learning, so that local model learning is free of backdoor representation, and the robustness of a global model is improved.
Owner:BEIJING ELECTRONICS SCI & TECH INST

End-to-end control system and method fusing multi-modal perception and strategy collaborative optimization mechanism

The invention specifically discloses an end-to-end control system and method fusing multi-modal sensing and a strategy collaborative optimization mechanism. The system comprises a multi-modal sensing module; the key frame selection module and the time sequence modeling module are used for dynamically screening key frames from the historical sensing frames and carrying out time sequence modeling and feature coding by adopting a time sequence Transform model to obtain time sequence state features with semantic continuity; the strategy optimization module is used for constructing an Actor-Critic strategy network based on a maximum entropy reinforcement learning algorithm; the system is composed of a multi-mode sensing module, a time sequence modeling module, a strategy learning module and an end-to-end control module, a sensing-decision-control integrated closed-loop structure is formed, the system takes multi-source sensor data in a real traffic scene as input, and the multi-source sensor data in the real traffic scene are subjected to uniform feature coding and time sequence modeling, so that the real traffic scene is obtained. The method has the advantages of being short in information processing link, high in response speed, high in adaptive capacity and the like.
Owner:TONGHUI INTELLIGENT MACHINERY (SHANGHAI) TECHNOLOGY CO LTD

Physical guidance-based graph multi-agent reinforcement learning drainage basin water resource distributed allocation method and system

The invention relates to the crossing field of water management and artificial intelligence. The method disclosed by the invention solves the problems of difficulty in coordination of multi-subject benefit conflicts, insufficient utilization of a basin spatial topological structure, separation of physical constraints and decision-making processes, limitation of adaptive ability and the like in the prior art. The method is characterized by comprising the steps of abstracting a drainage basin into a directed graph structure; generating a node embedding vector by using a physically guided graph attention network, and explicitly introducing a physical constraint factor in an attention mechanism; configuring an intelligent agent for each sub-basin to perform distributed strategy learning; cooperative training is carried out by adopting a multi-agent depth deterministic strategy gradient algorithm, fair distribution is guided by a global reward function in a Nash product form, and physical constraints such as water balance and the like are ensured in combination with a local reward function containing a nonlinear physical penalty term. According to the method, distributed, physically consistent and self-adaptive optimal configuration of drainage basin water resources is realized, and the method is mainly used for improving fairness, efficiency and feasibility of water resource distribution.
Owner:李博

Quadruped robot robust adaptive multi-skill learning method based on key frame guidance

The invention discloses a robust self-adaptive multi-skill learning method for a quadruped robot based on key frame guidance. The robust self-adaptive multi-skill learning method comprises the following steps: determining a target key frame set of multi-skill learning; in a flat terrain environment, probability sampling is carried out on the key frames according to skill task values to generate skill instructions, the quadruped robot is controlled to execute actions, and training tracks are collected and stored in a strategy learning experience playback pool and a self-simulation learning experience playback pool respectively. Training the strategy network and the value function network through strategy learning circulation; a self-mimicking learning discriminator network is trained through a self-mimicking learning loop to extract a high quality trajectory. And after the flat terrain training is completed, entering a complex terrain environment, and carrying out adaptive training under the guidance of the trained strategy network and the self-simulation learning discriminator network so as to realize the migration of the skill primitives to the complex environment. According to the method, multi-skill learning and terrain adaptive capacity can be effectively considered, and the robustness, flexibility and multi-task execution performance of the quadruped robot are improved.
Owner:FUDAN UNIVERSITY

Autonomous learning and optimization control construction method based on plant-level cloud ICS platform

The invention provides an autonomous learning and optimization control construction method based on a plant-level cloud ICS platform, and the method comprises the steps: collecting the data of each subsystem of a power plant in real time through a plant-level cloud ICS platform architecture, and storing the data in a private cloud platform database; carrying out preprocessing and feature extraction on the collected data; constructing and training an autonomous learning model by using the data after preprocessing and feature extraction, and completing the learning of a power plant operation rule and an optimization control strategy; analyzing and predicting power plant data collected in real time through the trained autonomous learning model, generating optimization control strategies for different subsystems, and applying the optimization control strategies to the subsystems of the power plant; and performing performance evaluation on the applied optimization control strategy, collecting feedback information, and adjusting and optimizing the autonomous learning model according to the feedback information. Through the construction of the autonomous learning model and the generation of the optimization control strategy, the intelligent management of the operation of the power plant is realized, and the operation efficiency of the power plant is improved.
Owner:SOUTHWEST ELECTRIC POWER DESIGN INST OF CHINA POWER ENG CONSULTING GROUP CORP

Decentralized dynamic policy learning and implementation system

A decentralized dynamic policy learning and implementation system automatically processes incoming report and generates real-time data based on pre-configured rules and uses a graph-based transformer to generate new policies. The generation of the new polices is based on new alerts or notifications indicating an error with existing policies. An intelligent auto-router detects new security issues and re-routes information to generate the new policy and initiate analysis via a digital twin based on decisioning data points. The digital twin system qualifies the new policy to determine effectiveness in overcoming the security issue. An extended reality environment automatically generates an extended reality environment based on a graph node structure of the policy.
Owner:BANK OF AMERICA CORP

Policy learning method with privacy protection in mobile edge computing for intelligent agent

A policy learning method with privacy protection in mobile edge computing for an intelligent agent is provided, relating to the technical field of mobile communication. The method includes: establishing an edge-collaborative computing offloading model, where the edge-collaborative computing offloading model includes a service caching model, a task offloading model, and a system cost model; establishing an optimization problem for task offloading, service caching, computing resource allocation and transmission power control based on the edge-collaborative computing offloading model for minimizing task processing costs; abstracting the optimization problem to a partially observable Markov decision process; and autonomously learning a task offloading strategy, a service caching strategy, a computing resource allocation strategy, and a transmission power control strategy by using a federated learning-based multi-agent deep reinforcement learning algorithm based on the Markov decision process.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Information processing device, information processing method, and program

This provides a technology that enables policy learning such that constraint violations are less than or equal to the acceptable probability. [Solution] An information processing device for performing constrained reinforcement learning, comprising: an input unit for inputting the state and actions of an agent; a reward calculation unit for calculating a reward based on the state and actions of the agent; a determination unit for determining whether or not constraint conditions are satisfied based on the state and actions of the agent; and a learning unit for learning a policy to maximize the reward under the condition that the probability of violating the constraint conditions is less than or equal to a predetermined allowable probability.
Owner:OMRON CORP

Unmanned logistics customs clearance cooperation method based on multi-agent continuous learning

The invention discloses an unmanned logistics customs clearance cooperation method based on multi-agent continuous learning, and the method comprises the following steps: S1, carrying out scene modeling and constraint expression, abstracting a multi-agent system, formalizing constraint and compliance judgment conditions, and building a task state machine consistent with a customs clearance business process; s2, a multi-agent collaborative decision-making mechanism is used for generating an alternative path, an alternative intersection point and an alternative resource allocation scheme for execution; s3, self-game scene generation and strategy learning are carried out, a cooperation-confrontation mixed training mechanism is constructed, and a multi-target learning criterion consistent with a customs target is introduced; s4, continuous learning and rule increment adaptation are carried out; s5, cross-agent knowledge distillation and shared memory are carried out, and group consistency and mobility are improved; the scheme has the advantages that customs clearance throughput and timeliness are improved, and queuing and congestion are reduced; manual rule changing and manual scheduling costs are reduced; 'no chain breakage, no retention and no violation 'are guaranteed; and large-scale popularization and application are effectively supported.
Owner:NANJING SAMPLE TECHNOLOGY CO LTD +6

LLM-based domain-specific pipelined task-oriented dialogue system

PendingCN122112156ASave labor and material costsReduce hallucination problemsDigital data information retrievalNatural language data processingNatural language understandingDialog system
The application discloses a specific field pipeline task type dialogue system based on an LLM, which comprises an inquiry subsystem and an answering subsystem, wherein: the inquiry subsystem parses and guides a user to supplement key information of a question according to user input task text, the answering subsystem summarizes key information of a dialogue task, and performs vector matching through a built-in vertical field local vector database to generate a final answer to the dialogue task. The application uses a large language model to realize natural language understanding, dialogue state tracking and natural language generation modules in a task type dialogue system pipeline, simultaneously realizes a rule matching-based strategy learning module, understands and replies to user input text, and combines the content of a local knowledge base, so that the whole process is more interpretable, the illusion problem of the large language model is reduced, only a small amount of sample prompt learning is needed, a large amount of data training or fine tuning is not needed, and the workload during field migration is reduced.
Owner:SHANGHAI JIAOTONG UNIV

A game confrontation method and device for an unmanned cluster based on dynamic meta-learning

The application provides a game confrontation method and device for an unmanned cluster based on dynamic meta-learning, and belongs to the technical field of unmanned cluster decision control. The method comprises the following steps: acquiring a game situation index of the unmanned cluster in a strategy learning process; acquiring a fuzzy input variable corresponding to the game situation index; inputting the fuzzy input variable into a fuzzy controller to obtain an adjustment amount of a hyperparameter of an initial strategy generator; updating the hyperparameter based on the adjustment amount, and updating the initial strategy generator based on the updated hyperparameter to obtain a target strategy generator; and generating a game confrontation strategy of the unmanned cluster based on the target strategy generator. The application uses a fuzzy controller as a controller of a strategy reinforcement learning process, dynamically adjusts the hyperparameter of the strategy generator by online analysis of real-time game situation indexes, injects an online meta-learning capability based on experience into the unmanned cluster, and solves the problems of strategy oscillation, convergence to a suboptimal solution, and even complete inability to converge caused by environmental non-stationarity.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Ramp merge strategy learning method and device, electronic equipment and storage medium

The present disclosure provides a ramp merging strategy learning method and device, electronic equipment and storage medium, relating to the technical field of autonomous driving. Among them, the strategy learning method comprises: configuring cumulative expected return in the training environment, and the cumulative expected return is used to represent the cumulative value of the expected return of continuously executing the merging strategy to complete the ramp merging process; constructing a strategy evaluation model and a strategy improvement model based on the cumulative expected return, the strategy evaluation model is used to evaluate the first expected return obtained by following the merging strategy from the specified driving state, and the strategy improvement model is used to evaluate the second expected return obtained by adopting the specified driving action; alternately iteratively learning the strategy evaluation model and the strategy improvement model to obtain a distributed driving action; and performing collision avoidance correction operation on the distributed driving action based on a safety shield model to generate a corrected target merging strategy. Through the technical scheme of the present disclosure, the safety of the output strategy can be improved when applied to the scene of merging onto the highway from the ramp.
Owner:JINGDONG KUNPENG (JIANGSU) TECH CO LTD

Experiment teaching effect evaluation method based on reinforcement learning

The invention relates to the technical field of teaching evaluation, in particular to an experimental teaching effect evaluation method based on reinforcement learning, and the method comprises the steps: obtaining multi-dimensional impact factors, such as a teaching environment, student characteristics and experimental characteristics, which affect the experimental teaching effect, carrying out the clustering and grouping of experimental courses, and forming course groups with similar characteristics; determining an initial weight of an evaluation parameter in combination with an analytic hierarchy process, performing preliminary evaluation on the group, and dynamically screening a target evaluation parameter suitable for the group according to an evaluation result; constructing a membership function of the target parameters, and quantifying the membership degree of each parameter in the evaluation level; and finally fusing the parameter weight and the membership degree, introducing a reinforcement learning agent to carry out dynamic optimization evaluation, automatically adjusting an evaluation model through a strategy learning and reward feedback mechanism, and outputting a final evaluation result of the teaching effect. According to the method, intelligentization and adaptive optimization of the evaluation process are realized, and the accuracy, objectivity and practicability of experimental teaching evaluation are improved.
Owner:NANJING ZHONGCHUANG BOOKSTORE CO LTD

Unmanned aerial vehicle cluster autonomous collaborative countering method

The invention provides a method for an unmanned aerial vehicle cluster to autonomously collaboratively counteract the unmanned aerial vehicle cluster, which is applied to the technical field of unmanned aerial vehicle counteracting. Inputting the environment information into a neural network model for task allocation to obtain a task allocation result output by the neural network model; the task allocation result is input into a multi-agent reinforcement learning model for cooperative control strategy learning, and a target cooperative control strategy is obtained; inputting the target cooperative control strategy into a trained large language model for knowledge enhancement processing to obtain an action control instruction output by the large language model; and converting the action control instruction into an unmanned aerial vehicle flight control instruction, controlling each unmanned aerial vehicle in the unmanned aerial vehicle cluster based on the unmanned aerial vehicle flight control instruction, and countering the target object. According to the invention, the decision-making efficiency, generalization ability and safety of the anti-unmanned aerial vehicle system in a dynamic environment can be improved, and the new anti-unmanned aerial vehicle ability of'counteracting none by none 'and'making groups by groups' is formed.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Multi-agent simulation system for electricity market bidding decision

The invention discloses a multi-agent simulation system for electricity market bidding decision making. The multi-agent simulation system comprises an environment simulation module and a multi-agent module. The environment simulation module comprises a market output unit and a market settlement unit and is used for executing a market output settlement method, performing financial settlement and generating a reward signal and a market environment state; the multi-agent module comprises a plurality of agents, and each agent is composed of an agent strategy network and an agent action decoder; the multi-agent module is used for receiving the reward signal and the market environment state information fed back by the environment simulation module, and carrying out autonomous strategy learning and bidding decision making; the agent action decoder comprises an original action receiving and analyzing unit, a normalization and scale mapping unit, a constraint projection and structure integration unit and a compliance quotation output unit. According to the method, the convergence speed of multi-agent reinforcement learning in power market simulation and the economical efficiency of the strategy are remarkably improved.
Owner:SOUTH CHINA UNIV OF TECH +1

Clustered adversarial method and system based on expert knowledge assisted deep reinforcement learning

The application provides a cluster confrontation method and system based on expert knowledge assisted deep reinforcement learning. By introducing an expert knowledge base and imitation learning technology, combining deep reinforcement learning, the efficiency of the collaborative decision of the agent is optimized, thereby improving the initial strategy learning speed of the system and the overall combat effectiveness. The method aims to provide an effective initial strategy acquisition mechanism, accelerate the early strategy learning of the agent by using the expert knowledge base, reduce the training time, and optimize the strategy of the agent in a complex dynamic environment through a multi-agent deep reinforcement learning algorithm to improve the collaborative combat capability. The scheme can greatly shorten the time required for initial strategy learning, obtain a more optimized strategy by combining deep reinforcement learning, ensure the efficiency of strategy optimization, and further ensure the real-time performance in large-scale cluster confrontation.
Owner:TONGJI UNIV

A Reinforcement Learning-Based Method for UAV Swarm Formation Decision-Making and Encirclement

This invention relates to a reinforcement learning-based method for UAV swarm formation decision-making and encirclement, belonging to the field of multi-agent formation control and decision-making technology. This invention establishes a consensus negotiation update equation and a multi-agent dynamic model, obtaining the desired positions of each agent and control commands for the multi-agent dynamic model. It achieves the desired formation configuration given by the decision system in a distributed manner without a central node, solving the multi-agent formation control problem. Secondly, it can intuitively manipulate the multi-agent formation to adapt to dynamically changing tasks, solving the multi-agent formation manipulation problem. Based on actual adversarial tasks, it realizes online autonomous decision-making for the learnable and evolutionary multi-agent system, solving the multi-agent autonomous formation decision-making problem. Finally, it utilizes control model mechanisms to enhance the interpretability and security of the learned decisions, considering the interpretability and security issues in the formation strategy learning process.
Owner:BEIHANG UNIV

Multi-modal adaptive knowledge graph construction and intelligent fusion system

The invention discloses a multi-modal adaptive knowledge graph construction and intelligent fusion system, and particularly relates to the field of knowledge graph construction and fusion, and the system comprises a data collection module, a data preprocessing module, a knowledge graph construction module, a cross-modal matching module and a verification application module. The data acquisition module is used for acquiring public resource data sources of texts, images and audios; the data preprocessing module is used for performing feature alignment and strategy learning on the acquired multi-modal data, and extracting, enhancing, clustering and associating cross-media data feature information; the knowledge graph construction module is used for modeling a multi-modal entity and a cross-modal relationship thereof, and constructing a multi-modal knowledge graph; and the cross-modal matching module is used for matching cross-modal information between entities. According to the method, the rapid and high-generalization construction capability of the multi-modal domain knowledge graph is realized, the capabilities of sensing, discovering and early warning of abnormal event multi-source data are established, and technical support is provided for multi-modal public resource data fusion application of a computing power center.
Owner:CHENGDU UNIV OF INFORMATION TECH

Equipment cooperative control system for coal mine energy scheduling

The invention discloses an equipment cooperative control system for coal mine energy scheduling, which relates to the technical field of coal mine equipment cooperative control, and comprises an intelligent agent module, a decision module, a power supply module and a power supply module, an improved Q-learning algorithm is adopted to carry out strategy learning on each agent, a game theory mechanism is utilized to process a resource competition problem, a cooperative scheduling decision and a solving module are generated, a multi-objective function is constructed according to the cooperative scheduling decision, an NSGA-II algorithm is adopted to solve a Pareto optimal solution set, and a multi-objective function is constructed according to the Pareto optimal solution set. And generating a frequency conversion speed regulation instruction, a load balancing strategy and a start-stop time sequence control signal. According to the coal mine equipment cooperative control method, the coal mine equipment generates the cooperative scheduling decision, the cooperative efficiency between the coal mine equipment is improved, and energy consumption and mechanical loss are reduced.
Owner:SHANDONG SANHEKOU MINE CO LTD

Robot strategy learning method and related device

The invention discloses a robot strategy learning method and a related device, and the method comprises the steps: constructing a training data set based on a sharing control framework of teleoperation; the training data set is utilized, a course training method for gradually weakening visual input disturbance is adopted, the robot strategy is trained and learned, the trained and learned robot strategy is obtained, the method and the related device can improve the data collection efficiency and naturalness, and meanwhile the generalization ability and the task adaptability of the robot strategy are improved.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Auxiliary teleoperation method based on vision-language-action model

The invention belongs to the technical field of teleoperation and robot control, and particularly relates to an auxiliary teleoperation method based on a vision-language-action model. Based on a few-sample strong generalization auxiliary teleoperation framework, the key technology of the method is divided into two core stages, namely a data preprocessing stage and a strategy learning and reasoning stage, and the method comprises the following steps: S1, injecting random noise into a supervised track, and constructing intention disturbance distribution; s2, extracting a key frame of a supervised track, and constructing an intention representation of geometric perception; and S3, encoding the processed trajectory as potential embedding, and providing conditions for a vision-language-action model controller. According to the method, the visual information, the language instruction and the action strategy are fused, rapid adaptation of the teleoperation task is achieved, good generalization ability is achieved among different operators, cross-operator migration and robust control are supported, and accurate cross-operator intention recognition and strategy execution are achieved.
Owner:JILIN UNIVERSITY

Flexible job shop scheduling method based on hybrid graph attention network and reward shaping

The invention discloses a flexible job shop scheduling method based on a hybrid graph attention network and reward shaping, and belongs to the field of intelligent manufacturing and scheduling optimization. The method comprises the following steps: firstly, constructing a heterogeneous graph model, and integrating processes, machine nodes and multi-type adjacency relationships; a heterogeneous graph neural network (HGNN) is adopted to aggregate local features, a double Transform structure is combined to enhance a global context, and state representation fusing local and global information is generated; designing a scheduling strategy network based on an Actor-Critic framework, generating an optimal decision through Actor, and estimating a state value through Critic; a reward shaping mechanism is introduced, and a multi-dimensional reward function containing the maximum completion time improvement amount and the machine load balance degree is constructed; and a near-end strategy optimization algorithm is adopted to jointly train the network, so that adaptive scheduling optimization is realized. According to the method, the completion time can be effectively shortened, the machine load balance is improved, and the strategy learning stability and generalization ability are enhanced.
Owner:FUZHOU UNIV

Multi-sensor management and control method based on multi-agent deep reinforcement learning

PendingCN121787506AImprove tracking accuracyAchieve stable characterizationBiological modelsSimulationMultiple sensor
The invention discloses a multi-sensor management and control method based on multi-agent deep reinforcement learning, and the method comprises the steps: firstly building a state equation and a motion model of a target, and an observation model and a motion model of a sensor node; then constructing a multi-agent strategy learning module; the multi-agent strategy learning module adopts an MAPPO algorithm; then, each intelligent agent processes measurement information in the vision field range of the intelligent agent to obtain a multi-dimensional matrix state diagram and a search utility diagram; and splicing the two information with the state information of the agents, inputting the spliced information to a multi-agent strategy learning module, outputting and executing the control action of each agent at the next moment, and repeating the steps to realize the global estimation of the target in the monitoring area. According to the method, while the bottleneck of centralized calculation is effectively avoided, the estimation capability of target information outside a visual field is expanded, and deep fusion of GM-PHD probability information and a reinforcement learning reward mechanism is realized, so that the global estimation and tracking performance of a system on multiple targets is remarkably improved.
Owner:HANGZHOU DIANZI UNIV