Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

485 results about "Value network" patented technology

A value network is a business analysis perspective that describes social and technical resources within and between businesses. The nodes in a value network represent people (or roles). The nodes are connected by interactions that represent tangible and intangible deliverables. These deliverables take the form of knowledge or other intangibles and/or financial value. Value networks exhibit interdependence. They account for the overall worth of products and services. Companies have both internal and external value networks.

Collaborative knowledge fusion reinforcement learning method for sparse reward environment

The invention discloses a sparse reward environment-oriented collaborative knowledge fusion reinforcement learning method, and relates to the field of collaborative knowledge fusion reinforcement learning methods. By constructing a lightweight collaborative knowledge fusion model and a dynamic reward remodeling mechanism, the problems of low intelligent agent exploration efficiency and difficulty in strategy convergence in a sparse reward environment are solved. The method comprises the following steps: constructing a reinforcement learning framework comprising a policy network and a value network; designing an action space mutation supervision mechanism and a lightweight collaborative knowledge fusion model, and generating a smooth substitution action when a strategy is detected to be unstable; and a reward function is designed in combination with the task target and the dynamic constraint, and reward remodeling is realized by activating rewards through sub-target potential energy difference and knowledge fusion. According to the method, effective intermediate feedback can be provided for the agents in the sparse reward environment, the exploration efficiency is improved, the convergence time is shortened, the stability and cross-scene migration ability of the strategy are enhanced, and an effective solution is provided for sparse reward scenes such as robot control and multi-agent game.
Owner:CHANGCHUN UNIV OF TECH

Industrial production line multi-equipment dynamic collaborative scheduling method and system based on reinforcement learning

The invention relates to the technical field of industrial production lines, and discloses an industrial production line multi-device dynamic collaborative scheduling method based on reinforcement learning, comprising the following steps: S1, modeling a three-dimensional state space; s2, hierarchical reinforcement learning architecture; and S3, edge-cloud cooperative execution. According to the industrial production line multi-device dynamic collaborative scheduling method and system based on reinforcement learning, device states, task constraints and resource occupation are integrated into a structured matrix through three-dimensional state space modeling, and a global decision-making layer captures production time sequence dependence by using a bidirectional long-short-term memory network; modeling equipment space association and process constraints through a graph attention network, and generating a global strategy including task allocation, capacity adjustment and resource pre-allocation; and after the edge layer detects the dynamic event, the cloud platform generates a candidate scheme through Monte Carlo tree search, and realizes dynamic event response and multi-target collaborative optimization by combining multiple targets such as global value network evaluation task completion time and equipment load balancing.
Owner:HUNAN LIANGYUAN AUTOMATION EQUIP CO LTD

Quadruped robot robust motion control method based on deep reinforcement learning

The invention discloses a quadruped robot robust motion control method based on deep reinforcement learning, and belongs to the technical field of robot motion control, and the method comprises the steps: constructing a deep reinforcement learning model which comprises a state estimation network, a strategy network and a value network; interaction between the quadruped robot and the simulation environment is carried out, and standard observation information, historical observation information and privileged observation information of the quadruped robot at all moments are obtained; inputting the standard observation information, the historical observation information and the privilege observation information of the moment into a deep reinforcement learning model, and training based on a total loss function until convergence is carried out to obtain a trained deep reinforcement learning model; inputting standard observation information and historical observation information at corresponding moments in an actual scene into the trained deep reinforcement learning model to obtain output features of a corresponding strategy network; and the target position of each joint motor is calculated to complete the motion control of the quadruped robot. And efficient training and robust motion on various complex and unstructured terrains can be realized.
Owner:ZHEJIANG UNIV OF TECH

Motor defect identification method fusing time sequence space feature extraction and reinforcement learning

The invention provides a motor defect identification method fusing time sequence space feature extraction and reinforcement learning. The method comprises the following steps: building a Transform-GAT model architecture T-GAT, and learning time sequence relevance and spatial topological structure features by the T-GAT to form a space-time composite representation vector; designing a dual-network architecture reinforcement learning weighted fusion mechanism of a strategy network and a value network, and after the strategy network and the value network of a parallel structure receive the composite vector, constructing a strategy function to select a defect type with the maximum probability; designing potential energy function quantization parameters, and constructing a reward function to generate a reward in combination with a difference value; taking a reward function as a target, training and optimizing model parameters of the dual-network architecture in an off-line manner, running a real-time decision in an on-line manner, and storing a tetrad to an experience pool to form a'perception-decision-feedback-update 'closed-loop mechanism; according to the method, motor defect identification is realized, downtime is reduced, and motor operation reliability and equipment operation efficiency are improved.
Owner:长沙千之然信息科技有限公司

Power stealing risk threshold dynamic optimization and multistage early warning method based on reinforcement learning

The invention discloses an electricity larceny risk threshold dynamic optimization and multistage early warning method based on reinforcement learning, and the method achieves the dynamic adjustment of an electricity larceny risk threshold through constructing a multi-dimensional state space containing a risk perception percentage, a line loss rate deviation, a cost budget and a load feature vector, and employing a reinforcement learning algorithm. Safety constraints are introduced to ensure the rationality of threshold adjustment, a TD3 algorithm is adopted to train and update a strategy network and a value network, the learning efficiency is improved through priority experience playback, a multi-target weighted reward function is designed to balance the multi-target optimization requirements of detection precision, line loss control, inspection cost and the like, and a red, orange and yellow three-color early warning mechanism is constructed. Early warning actions of different levels are triggered according to the ratio of the dynamically updated reference threshold value to the risk perception percentage, and early warning visualization and strategy effect evaluation are achieved through a geographic information graph dynamic rendering module, a natural language report generation module and an anti-fact analysis module.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Water body cleaning equipment path optimization method based on deep reinforcement learning

The invention discloses a water body cleaning equipment path optimization method based on deep reinforcement learning, and the method comprises the following steps: S1, collecting water body environment data through a sensor, and carrying out the preprocessing; s2, establishing a state space and an action space for path optimization, and constructing a reward mechanism; s3, carrying out feature extraction on the preprocessed data by adopting a Transform network; s4, training the strategy network and the value network by using an improved trust region strategy optimization algorithm, optimizing an update step length of the strategy network based on KL divergence constraint, and optimizing a path planning strategy based on a strategy gradient method; s5, performing path planning by using the trained strategy network and value network; and S6, calculating the energy consumption of the water body cleaning equipment, and adjusting the operation mode or operation path of the water body cleaning equipment. According to the method, the path planning of the water body cleaning equipment is optimized by combining Transform and an improved trust region strategy optimization algorithm, and the method has the advantages of being high in environmental adaptability, low in energy consumption, high in cleaning efficiency and high in operation stability.
Owner:HEFEI OUWO ENVIRONMENTAL PROTECTION TECH CO LTD

Information retrieval system and method based on semantic normalization

The invention discloses an information retrieval system and method based on semantic normalization, and relates to the technical field of artificial intelligence information, and the method comprises the steps: collecting a semantic query record input by a user, carrying out the preliminary semantic analysis, and generating structured data; on the basis of the structured data, entity disambiguation is carried out by utilizing a knowledge graph, abstract classes are generated through a neural network, calibration and dynamic weight adjustment are carried out, and high-confidence entity abstract classes and confidence scores are generated; entity abstract classes and confidence scores are combined with user contexts, an action-value function is calculated through a value network, and an optimal action is selected by utilizing a-greedy algorithm; executing semantic normalization mapping according to the optimal action, and obtaining an intermediate expression by using a meta-symbol dynamic generator; and performing index retrieval and multi-dimensional sorting based on the intermediate expression to generate a sorted retrieval result list. According to the method, the semantic fragmentation problem of multi-modal query is solved, and deep semantic alignment and dynamic weight calibration of heterogeneous data are realized.
Owner:上海笑聘网络科技有限公司

Milling processing path automatic generation method based on reinforcement learning

The invention discloses a milling path automatic generation method based on reinforcement learning, and relates to the technical field of machine tool machining, and the milling path automatic generation method comprises the following steps: step 1, defining action output; step 2, defining state input including a cutter state, a global observation space and a local observation space; 3, defining a reward function; value feedback given to the intelligent agent by the environment according to the current state and action is determined through the reward function; 4, setting deep neural networks including a feature extractor, a strategy network and a value network; 5, constructing a process autonomous generation model and a machining process world model; step 6, strategy training; step 7, outputting a strategy, and autonomously completing a machining process flow; by means of the method, the problems that in the prior art, manual experience dependence is too heavy, local optimization limitation is remarkable, and physical simulation efficiency is insufficient are solved.
Owner:SHANGHAI JIAOTONG UNIV

Server fault diagnosis method and device, storage medium and program product

The invention discloses a server fault diagnosis method and device, a storage medium and a program product, and relates to the technical field of server fault localization, and the method comprises the steps: carrying out the preprocessing of original log data, determining a state space according to the obtained preprocessed log data, a fault classification model containing a value network and an action network is obtained through reinforcement learning in advance, the value network in the fault classification model is used for carrying out value evaluation on the state space, and the action network in the fault classification model is used for carrying out action probability distribution calculation on the state space according to a target value obtained through value evaluation; and determining the target fault type according to the calculated probability distribution result, thereby solving the technical problem that new features or anomalies cannot be effectively judged, effectively improving the quality and integrity of log data, improving the server fault diagnosis efficiency and accuracy, realizing intelligent identification of the fault type, and improving the fault diagnosis efficiency and accuracy. The method has the technical effects of good training stability and generalization ability.
Owner:ZHENGZHOU YUNHAI INFORMATION TECH CO LTD

Large language model security decision agent driven by security reinforcement learning

The invention discloses a security reinforcement learning-driven large language model security decision agent, and the decision agent comprises a high-level semantic planner which is used for receiving a target and constraint instruction in a text form, receiving a language or visual observation signal of an environment at the same time, and outputting text formatted security risk information and suggested action planning; the low-layer action actuator is used for receiving low-dimensional observation and semantic codes of the environment, and the semantic codes are output by the high-layer semantic planner after text embedding conversion; the strategy network of the low-layer action actuator outputs a final safety action; the training alignment module is used for optimizing the strategy network and the value network; a high-level semantic planner is fed back and prompted through reward and cost signals collected through environment interaction, and parameters of a strategy network and a value network are trained through a security reinforcement learning algorithm. According to the method, the decision cannot violate the given text security constraint while the decision of the given text target is completed.
Owner:BEIHANG UNIV

Crop nitrogen fertilizer management system and method based on multi-agent reinforcement learning

The invention discloses a crop nitrogen fertilizer management system and method based on multi-agent reinforcement learning, and belongs to the technical field of intelligent agriculture. Comprising the following steps: collecting and preprocessing multi-source data of a target area, calibrating a crop growth-nitrogen cycle model based on the multi-source data, and constructing a dynamic simulation environment; constructing a nitrogen fertilizer application strategy model and a multi-target award function, and performing agent reinforcement learning training by adopting a centralized training-decentralized execution architecture; introducing a large language model, and updating a strategy network and / or a value network of each agent in a training process according to a reward adjustment signal and a decision constraint; when the index fluctuation ratio in the continuous evaluation period is smaller than a preset threshold value, it is judged that the nitrogen fertilizer application strategy model is converged, and an optimal nitrogen fertilizer application strategy model is obtained; and generating a nitrogen fertilizer application scheme based on the optimal nitrogen fertilizer application strategy model in combination with the real-time state data, and generating a natural language interpretation and risk assessment report based on a large language model.
Owner:INST OF SOIL SCI CHINESE ACAD OF SCI

Relationship-driven multi-agent reinforcement learning method and system based on mixed game

The invention belongs to the field of multi-agent reinforcement learning, and discloses a relation-driven multi-agent reinforcement learning method and system based on a mixed game, and the method comprises the steps: enabling a strategy network to generate an agent action, and carrying out the interaction of an environment, so as to collect sample data; the centralized value evaluator calculates marginal influence values of the agents based on samples, and deduces a social influence weight set; assigning group external rewards as individual external social rewards based on the weights; processing the global state by using a random network distillation-driven method, updating a prediction network parameter to minimize a prediction error, and outputting an internal reward set; the external social rewards and the internal rewards are fused to form comprehensive rewards; and updating the strategy network and the value network by using the comprehensive reward, and circularly training until convergence. By adopting the method, accurate modeling of the individual interaction relationship is enhanced, invalid exploration is remarkably reduced, the learning efficiency and the strategy reliability are improved, and the overall performance of the system is enhanced through an explicit social mechanism.
Owner:XI AN JIAOTONG UNIV

Data center IT load and cooling system cooperative control method based on TD3 algorithm

The invention provides a data center IT load and cooling system cooperative control method based on a TD3 algorithm, and relates to the field of deep reinforcement learning, and the method comprises the steps: obtaining an initial data set representing the server state, environment and task characteristics of a data center, and processing the initial data set to obtain a target data set; building a deep reinforcement learning model which is based on a TD3 algorithm and comprises a strategy network Actor and a value network Critic, and determining a state space, an action space and a reward function of the model; training the model to obtain a target model; and outputting a work task allocation strategy and a cooling system regulation and control strategy through interaction of the target model and the environment. According to the target model, by outputting the work task allocation strategy and the cooling system regulation strategy, work task allocation can be guided, operation of the cooling system can be controlled, the IT load of the server and refrigeration balance of the cooling system are achieved, safe and stable operation of the server is ensured, and energy consumption of the cooling system is reduced.
Owner:HEFEI UNIV OF TECH

Robot motion control model training method, device and equipment based on deep reinforcement learning, robot and medium

The invention provides a robot motion control model training method, device and equipment based on deep reinforcement learning, a robot and a medium, and relates to the technical field of robots. The method comprises the following steps: acquiring a first potential vector obtained after a student encoder encodes robot body observation data, and a second potential vector obtained after a teacher encoder encodes privilege observation data; based on the current training step number and a preset probability function, calculating a sampling probability for controlling a fusion proportion of the first potential vector and the second potential vector; fusing the first potential vector and the second potential vector based on the sampling probability to generate a third potential vector, and inputting the third potential vector into a strategy network; and updating the parameters of the policy network based on the value estimation of the current state output by the value network and the action policy output by the policy network. According to the method, updating oscillation caused by sudden change of input distribution in the training process of the strategy network can be avoided, the training efficiency is improved, and the training cost is reduced.
Owner:SHENZHEN ZHUJI POWER TECH CO LTD

Man-machine cooperation task allocation method based on reinforcement learning

The invention discloses a man-machine cooperation task allocation method based on reinforcement learning, and the method comprises the steps: constructing a reinforcement learning model which comprises a task definition, a resource definition, a state space, an action set and a reward function, and carrying out the modeling of a cooperation relation among a task, a worker and a robot in man-machine cooperation through a heterogeneous graph neural network, and extracting high-order features of the state space. Based on the characteristics, a strategy network and a value network are trained through a near-end strategy optimization algorithm, and learning of an efficient task allocation strategy is achieved. The reward function combines the task completion duration and the resource utilization rate, and gives consideration to the production efficiency and resource balance. In the system operation process, workshop state data are collected in real time to dynamically update the state space, and the reinforcement learning model generates an optimal allocation action according to the latest state and issues an instruction to a worker or a robot for execution. The method has the advantages of high adaptability, high distribution efficiency, high scheduling intelligence degree and the like, and is suitable for task scheduling optimization in complex dynamic environments such as intelligent manufacturing and the like.
Owner:TONGJI UNIV

Aircraft engine maintenance policy optimization method based on nonparametric reinforcement learning

Disclosed in the present invention is an aircraft engine maintenance policy optimization method based on nonparametric reinforcement learning. Firstly, for sparse aircraft engine operation data, constructing an aircraft engine model by means of a Bayesian network and a Gaussian process; then, establishing a policy network and a value network, interacting with the aircraft engine model to form a state / action value group, storing same into a replay buffer, and performing random sampling for use in a training set; updating the policy network and the value network, and updating the training set; and finally, performing aircraft engine maintenance optimization policy decision-making. Provided in the present invention is, for the first time, a Gaussian process-based nonparametric reinforcement learning method for an aircraft engine, which improves the degree of fit between overall training data and a model by means of dynamic data updating while integrating system uncertainty, thereby improving the sampling efficiency of an algorithm. Action selection is performed on the basis of prior maintenance experience data fused with uncertainty, improving the safety and sampling efficiency of a system and thereby solving the problem of predictive maintenance of an aircraft engine system.
Owner:ZHEJIANG UNIV

Time-varying task scheduling method and system based on space-time constraint near-end strategy optimization

The invention discloses a time-varying task scheduling method and system based on space-time constraint near-end strategy optimization, and belongs to the technical field of information processing. Comprising the steps of predicting task resource requirements of future N time steps through an LSTM model, and generating a time dynamic mask matrix Mtime; generating a spatial dynamic mask matrix Mspace based on a three-level hierarchical constraint rule; the three-level hierarchical constraint rule comprises resource capacity, interference avoidance and important application protection; the Mtime and the Mspace are fused into a global mask matrix Mt; an improved PPO algorithm is adopted, masked action probability distribution is output through an Actor network, and a Critic network evaluates a state value containing constraint punishment; and dynamically iterating the task scheduling strategy by jointly optimizing the strategy network, the value network and the mask generator parameters. According to the method, efficient utilization and load balancing of resources can be realized, and the stability and the performance of time-varying task scheduling are remarkably improved.
Owner:CIVIL AVIATION UNIV OF CHINA

Display panel repair path optimization method and device based on deep reinforcement learning

The invention relates to a display panel repair path optimization method and device based on deep reinforcement learning. The method comprises the following steps: obtaining a target display panel image; performing preprocessing and feature extraction on the target display panel image to obtain related features; inputting the related features into a pre-established deep reinforcement learning model, outputting an optimized repair path by the deep reinforcement learning model, and repairing the display panel; the deep reinforcement learning model is a PPO algorithm model which is constructed based on deep reinforcement learning and takes a strategy network and a value network as intelligent agents. According to the method, the problems of path redundancy and planning time consumption in a multi-defect scene in a traditional method are solved. Through real-time data acquisition, a reward function and a strategy are continuously updated by using a deep reinforcement learning algorithm, so that a repair path and an on-site defect state are matched in real time, and the problem that dynamic defect changes cannot be responded in real time in a traditional method is solved.
Owner:GUANGDONG UNIV OF TECH

Flexible workshop production resource scheduling method based on multi-strategy deep reinforcement learning

ActiveCN120355202AData processing applicationsMachine selectionMachine
The invention discloses a flexible workshop production resource scheduling method based on multi-strategy deep reinforcement learning, and relates to the field of intelligent manufacturing, the method constructs three strategy networks and three value networks of work, machines and AGVs, and at each scheduling moment, the work selection strategy network selects a scheduling rule according to a work state and maps the scheduling rule to a specific operation; the AGV selection strategy network selects a scheduling rule according to the AGV state and maps the scheduling rule to a # imgabs0 # table AGV # imgabs1 # of transportation operation; a machine selection strategy network is mapped to a # imgabs2 machine # imgabs3 # of processing operation according to a machine state, it is determined that the operation is transported from an AGV # imgabs4 # to a machine # imgabs5 # for processing through multi-strategy cooperation, the starting time and the ending time of the operation are calculated, states, actions and rewards are stored in corresponding experience pools, and a value network samples data from the experience pools, calculates losses and updates network parameters. And after the experience pool is emptied, entering a next round of iteration. The method can dynamically respond to production demand changes, and optimize resource allocation and scheduling decisions in real time.
Owner:GANTRY LAB +2

Sparse reward environment optimization learning identification method and system based on demonstration data enhancement

The invention relates to the technical field of mechanical arm optimization learning, and discloses a sparse reward environment optimization learning method and system based on demonstration data enhancement, and the method comprises the steps: obtaining expert demonstration data of a grabbing task of a mechanical arm, generating an enhanced demonstration sample through track segmentation time sequence interpolation and state space neighborhood extension, and carrying out the reconstruction of the enhanced demonstration sample; constructing a demonstration experience playback buffer area; building a reinforcement learning framework containing a strategy network and a value network, and learning an optimal grabbing strategy by utilizing reinforcement demonstration; designing a lightweight collaborative knowledge fusion model to monitor an action change rate, and outputting a smooth action when the action change rate exceeds a threshold value; and a reward function containing a task target and dynamic constraints is constructed, knowledge fusion rewards are superposed when the fusion model is activated, and reward sparsity compensation and strategy optimization are realized. According to the method, the executing stability and safety of the mechanical arm are improved, and the parameter updating accuracy and the training convergence speed are improved.
Owner:CHANGCHUN UNIV OF TECH

Flow field video generation method based on policy value architecture and online physical exploration

The invention discloses a flow field video generation method based on a policy value architecture and online physical exploration, and belongs to the technical field of crossing of artificial intelligence and computational fluid dynamics (CFD), and the method comprises the following steps: step 1, constructing an unsteady flow field multi-modal training data set, step 2, constructing a generative network system based on an Actor-Critic architecture, step 3, constructing an unsteady flow field multi-modal training data set, and step 4, constructing an unsteady flow field multi-modal training data set. Step 4, supervised fine tuning training is carried out in the first stage; step 5, online physical exploration of a generator is carried out in the second stage; step 6, feedback co-evolution of a physical encoder is carried out in the third stage; and step 7, reasoning generation of an unsteady flow field video is carried out. According to the method, a reinforcement learning architecture containing an Actor and a Critic is constructed, a physical equation is packaged into a digital environment, and a training strategy of basic supervision fine tuning, online physical exploration of a generator and coevolution feedback of an encoder is adopted.
Owner:CALCULATION AERODYNAMICS INST CHINA AERODYNAMICS RES & DEV CENT

Multi-agent combat mission cooperation method of structure entropy guided graph neural network

The invention discloses a multi-agent combat task cooperation method for a structure entropy guided graph neural network, and the method comprises the steps: S10, each combat agent interacts with an environment according to an action generated by a strategy network, the environment comprises environment information, task parameters and a preset task target, and the strategy of each combat agent is completely executed in a decentralized manner; collecting complete empirical trajectory data; s20, using the collected data for centralized training; performing value evaluation on the global state of each time step by using a value network; s30, calculating strategy loss and value loss by using a multi-agent near-end strategy optimization algorithm in combination with the output of the strategy network and the value estimation of the output of the value network; updating parameters of the strategy network and the value network by using a gradient descent method; and S40, performing loop iteration. The problems that in a traditional method, the battlefield game dynamic structure sensing ability is insufficient, the hierarchical strategy learning and generalization ability is limited, the adaptability of a model in a small sample area is poor, and the migration efficiency is low are solved.
Owner:BEIHANG UNIV

Rail transit multi-scene intelligent inspection device, method and equipment based on reinforcement learning

The invention relates to a rail transit multi-scene intelligent inspection device, method and equipment based on reinforcement learning, and the device comprises a sensing and state characterization module which is responsible for collecting multi-modal information from a rail transit inspection environment, and processing and fusing the multi-modal information into a structured environment state; the reinforcement learning decision-making module is used for selecting and outputting a current optimal decision-making action by utilizing the trained strategy network and value network based on the current structured environment state and the inspection task; the action execution and control module is used for converting the decision action into a bottom layer physical instruction sequence and controlling an intelligent agent to execute; the reinforcement learning training and optimization module is used for carrying out iterative optimization on the strategy network and the value network according to the inspection task and empirical data, and updating model parameters; and the task planning and scheduling module is used for defining inspection tasks and task allocation. Compared with the prior art, the method has the advantages of improving inspection efficiency and quality, realizing risk active early warning and the like.
Owner:CASCO SIGNAL LTD

Network congestion control method based on multi-agent reinforcement learning

The invention provides a network congestion control method and system based on multi-agent reinforcement learning, computer equipment and a storage medium, and relates to the technical field of network congestion control, a communication network comprising a plurality of communication nodes is constructed, the communication nodes are mapped into independent agents, and the independent agents are used for controlling the network congestion. Designing a weighted composite reward function containing indexes such as throughput and delay by taking a flow-level index as a local state and a data sending rate adjustment coefficient as an action; the intelligent agent selects action execution according to a current state based on a strategy network, calculates an environment reward based on the current state, the action, a next state after the action is executed and a weighted composite reward function, and stores the environment reward in a shared experience playback pool; training a value network and a strategy network by using the data in the pool; the trained strategy network is deployed at each node, distributed congestion control is realized, the stability of a congestion control decision in a complex network environment is improved, and bandwidth resources are distributed fairly while the throughput is ensured.
Owner:SUN YAT SEN UNIV

Vehicle formation control method fusing graph nerve and reinforcement learning

The invention provides a vehicle formation control method fusing graph nerve and reinforcement learning, and the method comprises the steps: constructing a graph structure model between vehicles, extracting the structural interaction characteristics between a current vehicle and a neighbor vehicle through a graph attention network, and combining with the dynamic evolution characteristics of the modeling historical time steps of a long-short-term memory network, and predicting a neighbor aggregation state at a future moment. And processing the two-dimensional reflection image acquired by the vehicle sensor by using a convolutional neural network, and extracting the spatial semantic information of the surrounding environment. A vehicle state, neighbor features and environment features are fused as strategy input, joint training is performed based on a multi-agent depth deterministic strategy gradient algorithm, and end-to-end control strategy optimization is realized by adopting a centralized value network and distributed strategy network structure. According to the invention, the intelligent decision-making level of the formation system is improved, and the system has good environment adaptability and cooperative control performance.
Owner:SHANGHAI UNIV

Spacecraft attitude and orbit integrated game control method and equipment based on PPO

The invention provides a PPO-based spacecraft attitude and orbit integrated game control method and device, and the method comprises the steps: building a spacecraft attitude and orbit integrated model, and carrying out the modeling of the interaction process of a tracking spacecraft and a target spacecraft as a non-cooperative zero-sum game; setting a multi-agent deep reinforcement learning framework, wherein the tracking spacecraft and the target spacecraft are respectively configured with an independent strategy network and an independent value network; training is carried out in a double-agent PPO mode, control instructions of orbit thrust and attitude torque are generated through the strategy network, and the state value is evaluated through the value network; and controlling and guiding the behavior of the spacecraft according to the trained strategy. According to the method, the autonomous decision-making capability and cooperative control performance of the spacecraft in complex tasks are effectively improved, and the method is suitable for high-precision aerospace tasks such as interception of non-cooperative targets, autonomous trajectory planning and attitude control.
Owner:HUBEI LUOJIA LAB

Personalized control method for transverse and longitudinal cooperation of intelligent vehicle based on reinforcement learning

The invention discloses an intelligent vehicle transverse and longitudinal cooperative personalized control method based on reinforcement learning, and the method comprises the steps: 1, obtaining target path information and vehicle state information, and constructing an intelligent body state space; 2, vehicle control parameters are obtained, and an intelligent agent action space is constructed; 3, constructing a vehicle dynamics safety boundary according to the vehicle dynamics model; 4, according to different driving style types, designing differentiated reward functions in combination with path tracking targets and dynamic safety boundary constraints; and 5, constructing a strategy network and a value network, and training the strategy network and the value network based on the reward functions of the three driving styles to obtain an optimal control model. According to the method, the transverse and longitudinal cooperative control model considering the dynamic factors and the driving styles is trained by using the reinforcement learning method, personalized control requirements are met, and meanwhile, the control precision and stability of the intelligent vehicle are improved.
Owner:HEFEI UNIV OF TECH

Cold chain transportation path optimization method based on intelligent scheduling

The invention discloses a cold chain transportation path optimization method based on intelligent scheduling, and the method comprises the steps: collecting the environmental data of a transportation path, and constructing a directed graph structure containing the temperature and length attributes of a path segment; defining a cooling capacity consumption function, calculating an estimated cooling capacity consumption value of the path section, and forming an environment diagram containing the refrigeration cost; constructing a state vector and an action space, and designing a path selection strategy network and a state value network; setting a multi-objective reward function including cold consumption, delivery completion, cold chain failure and supply behaviors, performing multi-round strategy training by adopting an improved PPO algorithm, and optimizing and updating network parameters through a trust region strategy; and finally, calling an optimal strategy output path selection action in actual transportation to realize dynamic avoidance and supply insertion regulation and control of the cold chain interruption risk. According to the invention, intelligent, self-adaptive and energy-efficient optimization of cold chain path selection can be realized.
Owner:BEIJING LONGXUNDA COLD CHAIN TRANSPORTATION CO LTD

Layout automatic generation method based on experience module and related device

The invention discloses an automatic layout generation method based on an experience module and a related device, and relates to the technical field of computers.The method comprises the steps that a template file is extracted, an independent device is generated based on the template file, and an experience module library is constructed based on the independent device; determining coordinate mapping and size constraints of the device, and performing multi-channel chip canvas hierarchical construction based on the coordinate mapping and size constraints of the device to obtain multi-channel chip canvas hierarchical representation; based on the empirical module library and the multi-channel chip canvas hierarchical representation, a reinforcement learning model is used for layout wiring, a target layout is obtained, and the reinforcement learning model comprises a state generation model, a value network and a strategy network. Layout and wiring of devices can be controlled more accurately, stray capacitance and resistance are effectively reduced, and the performance and reliability of an analog integrated circuit are effectively improved.
Owner:SUN YAT SEN UNIV +1

Data dynamic transaction value evaluation method

The invention relates to the technical field of data processing, and discloses a data dynamic transaction value evaluation method, which is applied to a server, and comprises the following steps: S1, receiving a data set from a data source, the data set comprising data attributes, association relationships among data and historical operation records; s2, on the basis of the received data set, adopting a mixed model of a graph attention neural network and a long and short-term memory network to calculate the total transaction value # imgabs2 # of data # imgabs0 # under the operation of a user # imgabs1 # at the time t; and S3, outputting the total transaction value to a user terminal, carrying out real-time risk monitoring according to the dynamic risk, and sending early warning information to the user terminal when the risk exceeds a preset threshold value. By adopting the graph attention neural network and long short-term memory network hybrid model, the data intrinsic value, network value, time value and dynamic risk are comprehensively evaluated, and compared with an existing method based on static rules, the accuracy is higher.
Owner:CENTURY LONGMAI TECH +1