Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

96 results about "Strategic learning" patented technology

Intelligent storage robot group collaborative scheduling method based on deep reinforcement learning

The invention provides an intelligent storage robot group cooperative scheduling method based on deep reinforcement learning, and relates to the technical field of intelligent storage, and the method comprises the steps: constructing a group perception module through a hierarchical attention mechanism, and generating a dynamic cooperative perception matrix; establishing a deep reinforcement learning model for strategy learning; and designing a multi-level reward function and optimizing a training process through an adaptive weight adjustment mechanism. According to the invention, the cooperative efficiency of warehouse robot group scheduling is improved, the task conflict rate is reduced, and the adaptability of the system to a complex dynamic environment is enhanced.
Owner:QINSILK COM

Three-dimensional perception robot operation knowledge distillation method based on monocular image

The invention relates to the field of robot operation and three-dimensional perception, in particular to a monocular image-based three-dimensional perception robot operation knowledge distillation method, which comprises the following steps of: establishing a strategy learning framework comprising a student model and a teacher model; constructing a strategy prediction model in the strategy learning framework; training the strategy prediction model, wherein a three-level knowledge distillation mechanism is adopted in the training process to complete knowledge migration between the teacher model and the student model; combining the three distillation losses with strategy optimization losses to form a total loss function, and performing end-to-end training on the student model until convergence to obtain a deployable monocular strategy model; the deployable monocular strategy model only retains a student model, and generates a robot operation instruction. The method has the beneficial effects that the robot under monocular RGB input has three-dimensional perception and high-precision operation capabilities while the reasoning efficiency is kept, the task success rate and generalization performance are remarkably improved, and the effectiveness and robustness of the method are verified.
Owner:ZHEJIANG UNIV OF TECH

Bridge group multi-target maintenance decision-making method fusing evolutionary algorithm and artificial intelligence

The invention provides a bridge group multi-target maintenance decision-making method fusing an evolutionary algorithm and artificial intelligence, and relates to the technical field of civil engineering and artificial intelligence crossing. The method comprises the steps of defining bridge group maintenance cost and structure failure risks, representing preferences of decision makers for different decision targets by weight combinations, and constructing a multi-target maintenance decision optimization model; encoding the weight combination into an individual of a multi-objective evolutionary algorithm, and randomly generating an initial population; aiming at each generation of weight combination, constructing a bridge group Markov decision-making environment; learning an optimal maintenance strategy by adopting an A2C training reinforcement learning agent; the optimal maintenance strategy is evaluated, and an evaluation result is used as individual fitness to be fed back to the multi-objective evolutionary algorithm; using a multi-objective evolutionary algorithm to perform evolutionary search on the multi-objective weight combination; and through a closed-loop feedback mechanism, outputting a Pareto optimal solution set containing an optimal maintenance strategy under various weight combinations, thereby realizing collaborative optimization of weight optimization and strategy learning.
Owner:UNIV OF SCI & TECH BEIJING

Multi-agent reinforcement learning method

The invention discloses a multi-agent reinforcement learning method. The method comprises the following steps: step 1, system initialization and edge device modeling; 2, generating a data source and constructing a task package; step 3, edge equipment trajectory planning and movement acquisition; 4, local task execution and unloading strategy decision making; 5, communication link modeling and bandwidth resource allocation are carried out; step 6, scheduling optimization driven by an information age and depreciation mechanism; step 7, multi-agent strategy optimization based on Actor-Critic is carried out; and step 8, reward function design and reinforcement learning process. The technical problems that an existing centralized scheduling or heuristic algorithm cannot perform efficient learning and scheduling in a resource heterogeneous and data dynamic environment, and cannot still have good learning ability and generalization ability under the condition of lack of global information are solved; the method is suitable for system modeling and optimization work of efficient strategy learning and collaborative decision making of multiple agents in a complete cooperation task.
Owner:HARBIN INST OF TECH +1

Complex network disintegration method based on evolution deep reinforcement learning

The invention discloses a complex network disintegration method based on evolution deep reinforcement learning. According to the method, an encoder-decoder model fusing a graph convolutional neural network and a deep Q network is constructed, and is used for efficiently extracting importance features of nodes in a complex network and realizing dynamic decision-making of a node disassembling sequence according to the importance features. In order to optimize model parameters and improve search capability, an evolutionary algorithm is introduced to perform global exploration on the model parameters, and the problem that a directional optimization strategy is easy to fall into local optimum is avoided. Meanwhile, deep mining is carried out on an evolution result in combination with a reinforcement learning strategy, the overall optimization process is accelerated, and advantage complementation of parameter evolution and strategy learning is achieved. Experimental results show that the method significantly improves the efficiency and precision of network disassembly while maintaining the robustness of the model, and has good practical value and wide application prospects.
Owner:NANJING UNIV OF SCI & TECH +2

Multi-modal fusion virtual simulation experiment teaching system

The invention provides a multi-modal fusion virtual simulation experiment teaching system, and relates to the technical field of virtual simulation, and the system comprises an information collection module which is used for obtaining the interaction information of a learner; the information fusion module is used for constructing a multi-modal input data set; the instruction analysis module is used for analyzing and generating an operation instruction sequence; the teaching intervention strategy generation module is used for generating a self-adaptive teaching intervention strategy; the teaching intervention strategy execution module is used for executing a self-adaptive teaching intervention strategy; and the learning effect evaluation module is used for quantifying the learning effect evaluation score. The technical problems that in traditional virtual simulation experiment teaching, multi-modal operation behaviors of learners are difficult to accurately capture in real time, personalized teaching intervention is lacked and learning effect evaluation is incomplete can be solved, and virtual experiment simulation and intelligent teaching intervention driven by multi-modal information fusion are achieved. And the experimental interaction experience and the learning effect evaluation precision are improved.
Owner:CHINA UNIV OF GEOSCIENCES (BEIJING)

Digital twin multi-agent reinforcement learning intelligent decision-making system with secure memory playback mechanism

The invention discloses a digital twinning multi-agent reinforcement learning intelligent decision-making system and method with a secure memory playback mechanism, and the system comprises a digital twinning module which is used for constructing a virtual model and synchronizing the virtual model with a physical entity in real time; the multi-agent reinforcement learning module is used for carrying out strategy learning based on a constrained Markov decision process and balancing performance and safety through a Lagrange multiplier; the safe memory playback module is used for weighting and playing back the experience samples according to the risk and the timeliness so as to improve the learning safety; the reversible grey influence network module is used for causal modeling and reasoning and enhancing decision interpretability; the double-loop self-constraint control module ensures that a control action is always in a physical safety boundary through a barrier function and safety projection; and the convergence and stability criterion module is used for verifying strategy security convergence and system asymptotic stability. According to the method, the problems of strategy border crossing, virtual-real mismatching and the like in the high-risk manufacturing process are solved, and multi-target optimal control under the safety constraint is realized.
Owner:CHONGQING UNIV +1

Incremental reinforcement learning path planning method based on dynamic reward remodeling

The invention provides an incremental reinforcement learning path planning method based on dynamic reward remodeling, which is used for solving the problems of training efficiency and composite optimization in path planning. According to the method, a rapid exploration random tree and an artificial potential field method are combined with reinforcement learning, a dynamic reward function is designed, real-time feedback is provided for an intelligent agent, and strategy learning is accelerated. The system specifically comprises a reward remodeling module which provides process rewards through a fast exploration random tree and an artificial potential field method; the sub-target course learning module is used for gradually increasing task complexity and optimizing sample utilization; the progressive reward adjustment module dynamically adjusts the reward weight and solves the reward coupling problem; and the strategy fine tuning module dynamically adjusts the learning rate and supports strategy optimization. Experimental results show that in static and dynamic path planning tasks, compared with a traditional method, the method has the advantages that the training efficiency and the composite optimization performance are remarkably improved, the reward coupling problem is effectively solved, and the intelligent agent can efficiently learn and converge to a global optimal strategy.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Multi-mechanism deep integrated multi-objective optimization framework and optimization method and application thereof

InactiveCN120725087ANeural learning methodsEvolutionary learningAlgorithm
The invention relates to the technical field of multi-objective optimization, in particular to a multi-mechanism deep integrated multi-objective optimization framework and an optimization method and application thereof, comprising a reinforcement learning module, a genetic algorithm module, a bidirectional evolutionary learning coupling module, a Pareto experience pool module, a preference management module and a graph neural network module, the bidirectional evolutionary learning coupling module is used for realizing bidirectional information interaction between the reinforcement learning module and the genetic algorithm module, and the reinforcement learning module generates candidate solutions and feeds the candidate solutions back to the genetic algorithm module; the genetic algorithm module screens the elite solution with the crowding distance larger than a set threshold value in the first Pareto layer and updates the strategy of the feedback reinforcement learning module. According to the method, the strategy learning ability of reinforcement learning and the global population search ability of the genetic algorithm are fused, and the overall optimization efficiency and the approximation ability to the complex Pareto frontier are remarkably improved.
Owner:BEIJING ZHONGDIAN JINGYI TECH CO LTD

Cluster confrontation method and system based on expert knowledge assisted deep reinforcement learning

The invention provides a cluster confrontation method and system based on expert knowledge-assisted deep reinforcement learning, and the method and system improve the initial strategy learning speed and overall combat effectiveness of the system by introducing an expert knowledge base and an imitation learning technology and combining deep reinforcement learning to optimize the collaborative decision-making efficiency of an intelligent agent. The method aims at providing an effective initial strategy acquisition mechanism, utilizing an expert knowledge base to accelerate the early strategy learning of the agents, reducing the training time, optimizing the strategies of the agents in a complex dynamic environment through a multi-agent deep reinforcement learning algorithm, and improving the cooperative combat ability. According to the scheme, the time required for initial strategy learning can be greatly shortened, a more optimized strategy is obtained in combination with deep reinforcement learning, the high efficiency of strategy tuning is guaranteed, and then the real-time guarantee in large-scale cluster confrontation is guaranteed.
Owner:TONGJI UNIV

Manufacturing industry data intelligent analysis method and system based on deep reinforcement learning

The invention relates to the technical field of data processing, and discloses a manufacturing industry data intelligent analysis method and system based on deep reinforcement learning. The method comprises the following steps: carrying out time sequence processing on manufacturing industry equipment state, order characteristics, inventory level and quality index data to obtain a four-dimensional production data matrix, carrying out strategy learning through an LSTM-Actor-Critic algorithm to obtain a manufacturing decision strategy network, carrying out classification processing according to a production cycle to obtain a hierarchical data set, constructing an intelligent experience playback buffer area, and carrying out intelligent experience playback. And carrying out collaborative optimization on order scheduling, inventory replenishment and equipment task allocation decisions to obtain a manufacturing industry data intelligent analysis result. The technical problem that an existing manufacturing industry data analysis method lacks adaptive learning ability and cannot process multi-domain collaborative decision optimization is solved.
Owner:TIANJIN HONGHUANG TECH CO LTD

Metareinforcement learning-driven adaptive task unloading mechanism in edge computing environment

The invention discloses an adaptive task unloading mechanism driven by meta-reinforcement learning in an edge computing environment. According to the mechanism, a system architecture composed of a user equipment layer and an edge server layer is constructed, and task analysis, state perception and strategy optimization processes are combined to realize unloading scheduling optimization of a multi-task dependent structure. Task unloading is modeled as a Markov decision process, a double-layer training mechanism is adopted, local strategy training is realized by utilizing a near-end strategy optimization algorithm, and the generalization ability of the system is improved in combination with cross-task meta-strategy learning. According to the method, the sequence is fused into the sequence neural network structure and the multi-head attention mechanism, and the accuracy and efficiency of unloading strategy generation are improved while the task dependency relationship is modeled. The mechanism has good task adaptability and delay optimization performance in a dynamic edge computing environment, and is suitable for various mobile computing scenes.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

Deep reinforcement learning method and system for building energy control

The invention is suitable for the technical field of deep reinforcement learning, and provides a deep reinforcement learning method and system for building energy control, and the method comprises the steps: constructing an intelligent agent based on a target building environment; collecting common experience data when the intelligent agent interacts with the target building environment; constructing a prompt text; inputting the prompt text into the large language model to obtain an action range of the control action; performing correction processing on the common empirical data according to the action range to obtain empirical data based on LLM correction; and according to the common experience pool and the LLM experience pool, training the intelligent agent by minimizing a loss function. According to the method, the action range of the control action in different environment states is analyzed by utilizing the large language model, so that low-quality empirical data in random exploration of the intelligent agent is guided and corrected, the intelligent agent can obtain high-quality experience, and the learning efficiency and the overall performance of the building energy control strategy are improved.
Owner:UESTC (SHENZHEN) ADVANCED RES INST +1

Multi-AUV safe path planning method based on deep reinforcement learning

The invention discloses a multi-AUV safe path planning method based on deep reinforcement learning. According to the method, a deep reinforcement learning algorithm model based on an MATD3 method is adopted, a policy network and a value network are updated by adopting security constraints, the security of policy learning is improved, the expected reward revenue is maximized under the condition that the expected security cost constraints are met, a frequent minimum and maximum optimization process is avoided by adopting first-order penalty optimization, and the security of policy learning is improved. Meanwhile, safety correction based on a safety layer is added in the training process to guarantee safety in the early stage of training, and exploratory and safety balance is brought to strategy optimization of reinforcement learning through safety constraint and safety correction. According to the multi-AUV path planning method, the time cooperation constraint and the space cooperation constraint of multi-AUV path planning are comprehensively considered, the centralized training and decentralized decision-making architecture is applied to multi-AUV path planning, the path planning method capable of ensuring cooperation safety is provided for a multi-AUV system, and the safety and the reliability of path planning are improved.
Owner:HARBIN ENG UNIV

Communication lightweight multi-micro energy network cooperative control method for large model assisted reward recovery

The invention belongs to the technical field of multi-micro-energy-network cooperative control, and particularly relates to a communication lightweight multi-micro-energy-network cooperative control method for reward recovery assisted by a large model, and the method comprises the following steps: S1, constructing a mathematical model of distributed energy in a micro-energy network, and enabling the mathematical model to be used for describing the dynamic behaviors of energy conversion, storage and constraint; s2, modeling a micro energy network cooperative control problem as a partially observable Markov decision process POMDP; s3, solving the POMDP of the step S2 by using an optimized MAPPO method to obtain a cooperative control scheme; and issuing to each micro-energy network terminal for execution. The method can break through the limitation of the communication bottleneck on the premise of ensuring the learning quality of the control strategy.
Owner:CHONGQING UNIV

Task type dialogue strategy learning method and system based on reinforcement learning

The invention belongs to the technical field of intelligent task-based dialogue, and particularly relates to a task-based dialogue strategy learning method and system based on reinforcement learning, which adopts a soft actor commentator algorithm to be combined with behavior cloning and maximum entropy inverse reinforcement learning to relieve a cold start problem. According to the method, maximum entropy inverse reinforcement learning is utilized, a reward value is calculated, and a user target in a dialogue is accurately deduced according to maximum log likelihood estimation of a human dialogue and a simulated dialogue; a potential reward function is extracted from a successful track through inverse reinforcement learning, manual reward design is replaced, accumulated reward maximization is pursued during strategy optimization, diversity and exploratory performance of the strategy are encouraged, and the dialogue strategy is prevented from being converged to a single mode too early.
Owner:QUFU NORMAL UNIV

Federated backdoor defense method based on decoupling contrast learning

The invention discloses a federated backdoor defense method based on decoupling contrast learning, and the method comprises the steps: training a backdoor model based on a backdoor sample, and immediately stopping training after the backdoor model converges on the backdoor sample; respectively extracting a penultimate layer vector of the backdoor model and the local model from a sample pair held by the malicious client as a backdoor feature and a clean feature; comparing and learning the separated back door features and the clean features, and learning the clean features for the local model by using a sample weighting strategy to train the local model to obtain a trained local model; and sending local model parameters of the trained local model to a global server, and generating model parameters of a new global model based on the local model parameters through an aggregation function. The method aims at reducing information dependence between backdoor features and clean features through comparative learning, so that local model learning is free of backdoor representation, and the robustness of a global model is improved.
Owner:BEIJING ELECTRONICS SCI & TECH INST

End-to-end control system and method fusing multi-modal perception and strategy collaborative optimization mechanism

The invention specifically discloses an end-to-end control system and method fusing multi-modal sensing and a strategy collaborative optimization mechanism. The system comprises a multi-modal sensing module; the key frame selection module and the time sequence modeling module are used for dynamically screening key frames from the historical sensing frames and carrying out time sequence modeling and feature coding by adopting a time sequence Transform model to obtain time sequence state features with semantic continuity; the strategy optimization module is used for constructing an Actor-Critic strategy network based on a maximum entropy reinforcement learning algorithm; the system is composed of a multi-mode sensing module, a time sequence modeling module, a strategy learning module and an end-to-end control module, a sensing-decision-control integrated closed-loop structure is formed, the system takes multi-source sensor data in a real traffic scene as input, and the multi-source sensor data in the real traffic scene are subjected to uniform feature coding and time sequence modeling, so that the real traffic scene is obtained. The method has the advantages of being short in information processing link, high in response speed, high in adaptive capacity and the like.
Owner:TONGHUI INTELLIGENT MACHINERY (SHANGHAI) TECHNOLOGY CO LTD

Physical guidance-based graph multi-agent reinforcement learning drainage basin water resource distributed allocation method and system

The invention relates to the crossing field of water management and artificial intelligence. The method disclosed by the invention solves the problems of difficulty in coordination of multi-subject benefit conflicts, insufficient utilization of a basin spatial topological structure, separation of physical constraints and decision-making processes, limitation of adaptive ability and the like in the prior art. The method is characterized by comprising the steps of abstracting a drainage basin into a directed graph structure; generating a node embedding vector by using a physically guided graph attention network, and explicitly introducing a physical constraint factor in an attention mechanism; configuring an intelligent agent for each sub-basin to perform distributed strategy learning; cooperative training is carried out by adopting a multi-agent depth deterministic strategy gradient algorithm, fair distribution is guided by a global reward function in a Nash product form, and physical constraints such as water balance and the like are ensured in combination with a local reward function containing a nonlinear physical penalty term. According to the method, distributed, physically consistent and self-adaptive optimal configuration of drainage basin water resources is realized, and the method is mainly used for improving fairness, efficiency and feasibility of water resource distribution.
Owner:李博

Quadruped robot robust adaptive multi-skill learning method based on key frame guidance

The invention discloses a robust self-adaptive multi-skill learning method for a quadruped robot based on key frame guidance. The robust self-adaptive multi-skill learning method comprises the following steps: determining a target key frame set of multi-skill learning; in a flat terrain environment, probability sampling is carried out on the key frames according to skill task values to generate skill instructions, the quadruped robot is controlled to execute actions, and training tracks are collected and stored in a strategy learning experience playback pool and a self-simulation learning experience playback pool respectively. Training the strategy network and the value function network through strategy learning circulation; a self-mimicking learning discriminator network is trained through a self-mimicking learning loop to extract a high quality trajectory. And after the flat terrain training is completed, entering a complex terrain environment, and carrying out adaptive training under the guidance of the trained strategy network and the self-simulation learning discriminator network so as to realize the migration of the skill primitives to the complex environment. According to the method, multi-skill learning and terrain adaptive capacity can be effectively considered, and the robustness, flexibility and multi-task execution performance of the quadruped robot are improved.
Owner:FUDAN UNIVERSITY

Autonomous learning and optimization control construction method based on plant-level cloud ICS platform

The invention provides an autonomous learning and optimization control construction method based on a plant-level cloud ICS platform, and the method comprises the steps: collecting the data of each subsystem of a power plant in real time through a plant-level cloud ICS platform architecture, and storing the data in a private cloud platform database; carrying out preprocessing and feature extraction on the collected data; constructing and training an autonomous learning model by using the data after preprocessing and feature extraction, and completing the learning of a power plant operation rule and an optimization control strategy; analyzing and predicting power plant data collected in real time through the trained autonomous learning model, generating optimization control strategies for different subsystems, and applying the optimization control strategies to the subsystems of the power plant; and performing performance evaluation on the applied optimization control strategy, collecting feedback information, and adjusting and optimizing the autonomous learning model according to the feedback information. Through the construction of the autonomous learning model and the generation of the optimization control strategy, the intelligent management of the operation of the power plant is realized, and the operation efficiency of the power plant is improved.
Owner:SOUTHWEST ELECTRIC POWER DESIGN INST OF CHINA POWER ENG CONSULTING GROUP CORP

Policy learning apparatus and method based on skill diffusion, and operation apparatus based on diffused skill-based policy

The present invention relates to a skill diffusion-based policy learning device and method, as well as a diffused skill-based policy-based operation device. A policy learning apparatus comprises an encoder configured to be trained to obtain a domain-invariant skill embedding based on at least one state and at least one action, and obtain a domain-variant skill embedding based on a domain parameter corresponding to the at least one state and the at least one action and the domain-invariant skill embedding and a decoder configured to be trained to obtain a skill as an output value using the domain-invariant skill embedding and the domain-variant skill embedding as inputs.
Owner:RES & BUSINESS FOUND SUNGKYUNKWAN UNIV

Decentralized dynamic policy learning and implementation system

A decentralized dynamic policy learning and implementation system automatically processes incoming report and generates real-time data based on pre-configured rules and uses a graph-based transformer to generate new policies. The generation of the new polices is based on new alerts or notifications indicating an error with existing policies. An intelligent auto-router detects new security issues and re-routes information to generate the new policy and initiate analysis via a digital twin based on decisioning data points. The digital twin system qualifies the new policy to determine effectiveness in overcoming the security issue. An extended reality environment automatically generates an extended reality environment based on a graph node structure of the policy.
Owner:BANK OF AMERICA CORP

Policy learning method with privacy protection in mobile edge computing for intelligent agent

A policy learning method with privacy protection in mobile edge computing for an intelligent agent is provided, relating to the technical field of mobile communication. The method includes: establishing an edge-collaborative computing offloading model, where the edge-collaborative computing offloading model includes a service caching model, a task offloading model, and a system cost model; establishing an optimization problem for task offloading, service caching, computing resource allocation and transmission power control based on the edge-collaborative computing offloading model for minimizing task processing costs; abstracting the optimization problem to a partially observable Markov decision process; and autonomously learning a task offloading strategy, a service caching strategy, a computing resource allocation strategy, and a transmission power control strategy by using a federated learning-based multi-agent deep reinforcement learning algorithm based on the Markov decision process.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Information processing device, information processing method, and program

This provides a technology that enables policy learning such that constraint violations are less than or equal to the acceptable probability. [Solution] An information processing device for performing constrained reinforcement learning, comprising: an input unit for inputting the state and actions of an agent; a reward calculation unit for calculating a reward based on the state and actions of the agent; a determination unit for determining whether or not constraint conditions are satisfied based on the state and actions of the agent; and a learning unit for learning a policy to maximize the reward under the condition that the probability of violating the constraint conditions is less than or equal to a predetermined allowable probability.
Owner:OMRON CORP

Unmanned logistics customs clearance cooperation method based on multi-agent continuous learning

The invention discloses an unmanned logistics customs clearance cooperation method based on multi-agent continuous learning, and the method comprises the following steps: S1, carrying out scene modeling and constraint expression, abstracting a multi-agent system, formalizing constraint and compliance judgment conditions, and building a task state machine consistent with a customs clearance business process; s2, a multi-agent collaborative decision-making mechanism is used for generating an alternative path, an alternative intersection point and an alternative resource allocation scheme for execution; s3, self-game scene generation and strategy learning are carried out, a cooperation-confrontation mixed training mechanism is constructed, and a multi-target learning criterion consistent with a customs target is introduced; s4, continuous learning and rule increment adaptation are carried out; s5, cross-agent knowledge distillation and shared memory are carried out, and group consistency and mobility are improved; the scheme has the advantages that customs clearance throughput and timeliness are improved, and queuing and congestion are reduced; manual rule changing and manual scheduling costs are reduced; 'no chain breakage, no retention and no violation 'are guaranteed; and large-scale popularization and application are effectively supported.
Owner:NANJING SAMPLE TECHNOLOGY CO LTD +6

LLM-based domain-specific pipelined task-oriented dialogue system

PendingCN122112156ASave labor and material costsReduce hallucination problemsDigital data information retrievalNatural language data processingNatural language understandingDialog system
The application discloses a specific field pipeline task type dialogue system based on an LLM, which comprises an inquiry subsystem and an answering subsystem, wherein: the inquiry subsystem parses and guides a user to supplement key information of a question according to user input task text, the answering subsystem summarizes key information of a dialogue task, and performs vector matching through a built-in vertical field local vector database to generate a final answer to the dialogue task. The application uses a large language model to realize natural language understanding, dialogue state tracking and natural language generation modules in a task type dialogue system pipeline, simultaneously realizes a rule matching-based strategy learning module, understands and replies to user input text, and combines the content of a local knowledge base, so that the whole process is more interpretable, the illusion problem of the large language model is reduced, only a small amount of sample prompt learning is needed, a large amount of data training or fine tuning is not needed, and the workload during field migration is reduced.
Owner:SHANGHAI JIAOTONG UNIV

Deep reinforcement learning method and system for building energy control

This application is applicable to the field of deep reinforcement learning technology and provides a deep reinforcement learning method and system for building energy control. The method includes: constructing an intelligent agent based on a target building environment; collecting common experience data when the intelligent agent interacts with the target building environment; constructing prompt text; inputting the prompt text into a large language model to obtain the action range of the control action; correcting the common experience data based on the action range to obtain experience data corrected based on the LLM; and training the intelligent agent by minimizing the loss function based on the common experience pool and the LLM experience pool. This application uses a large language model to analyze the action range of the control action under different environmental conditions, thereby guiding and correcting the low-quality experience data in the agent's random exploration, enabling the intelligent agent to obtain high-quality experience and improving the efficiency and overall performance of building energy control strategy learning.
Owner:UESTC (SHENZHEN) ADVANCED RES INST +1

A game confrontation method and device for an unmanned cluster based on dynamic meta-learning

The application provides a game confrontation method and device for an unmanned cluster based on dynamic meta-learning, and belongs to the technical field of unmanned cluster decision control. The method comprises the following steps: acquiring a game situation index of the unmanned cluster in a strategy learning process; acquiring a fuzzy input variable corresponding to the game situation index; inputting the fuzzy input variable into a fuzzy controller to obtain an adjustment amount of a hyperparameter of an initial strategy generator; updating the hyperparameter based on the adjustment amount, and updating the initial strategy generator based on the updated hyperparameter to obtain a target strategy generator; and generating a game confrontation strategy of the unmanned cluster based on the target strategy generator. The application uses a fuzzy controller as a controller of a strategy reinforcement learning process, dynamically adjusts the hyperparameter of the strategy generator by online analysis of real-time game situation indexes, injects an online meta-learning capability based on experience into the unmanned cluster, and solves the problems of strategy oscillation, convergence to a suboptimal solution, and even complete inability to converge caused by environmental non-stationarity.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI