Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

179 results about "Decision-making models" patented technology

All people need to make decisions from time to time. Given limited time in formulating policies and addressing public problems, public administrators must enjoy a certain degree of discretion in planning, revising and implementing public policies. In other words, they must engage in decision-making (Gianakis, 2004). Over the years, many scholars tried to devise decision-making models to account for the policy making process.

Decision model construction method based on big data environment

The invention provides a decision model construction method based on a big data environment. The method belongs to the economic data decision field. The method comprises the following steps: firstly, acquiring multi-source heterogeneous data including sensor time sequence data, an expert rule base and equipment causal priori knowledge, and constructing a standardized training data set; then, extracting an explicit causal relationship from the historical data by utilizing a causal discovery algorithm, and generating an interpretable knowledge graph in combination with priori knowledge of an equipment manual; thirdly, constructing a neural symbol joint model; neural network parameters and rule confidence coefficients are synchronously adjusted, and collaborative learning of data driving and knowledge driving is realized. And then, carrying out multi-dimensional logic verification on the trained model, including rule conflict detection, anti-factual reasoning and decision path traceability verification. The method can effectively improve the accuracy and interpretability of the decision model, and is suitable for processing decision tasks in a big data environment.
Owner:YIBIN VOCATIONAL & TECH COLLEGE

Perception collaborative decision-making method and system based on multi-modal heterogeneous data fusion

The invention provides a perception collaborative decision-making method and system based on multi-modal heterogeneous data fusion, and relates to the technical field of artificial intelligence, and the method comprises the steps: inputting a global environment situation perception graph into a pre-trained multi-target collaborative decision-making model; the multi-target collaborative decision-making model forms a multi-target decision-making feature set by analyzing the resource entities and the incidence relation in the graph; based on the multi-target decision feature set, decision optimization is carried out to obtain a comprehensive collaborative scheduling scheme; performing instruction analysis and packaging on the comprehensive collaborative scheduling scheme to obtain an executable instruction sequence; and issuing the executable instruction sequence to a corresponding decision node and a control terminal in parallel through a distributed communication architecture to complete real-time scheduling of resources and collaborative issuing of control instructions. According to the invention, by constructing a linkage mechanism of multi-modal data fusion, dynamic environment perception and collaborative decision execution, intelligent perception and quick response to a complex environment are realized.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

Power system load dynamic optimization method based on reinforcement learning

The invention discloses a power system load dynamic optimization method based on reinforcement learning. The method comprises the following steps: S1, collecting power system data to construct a state space; s2, constructing a hierarchical reinforcement learning model based on the state space, and dividing a high-level decision and a low-level execution task; s3, training a high-level decision model, and outputting a scheduling task target category instruction in a high-level state; s4, training a low-layer execution model, and outputting a control action in combination with a current node state and a high-layer instruction; s5, introducing an evolutionary mechanism to generate a strategy population and optimizing a low-layer execution model; s6, fusing an evolutionary mechanism and a strategy gradient to synchronously optimize individuals with excellent performance; s7, deploying the trained model to a power dispatching system; s8, performing model parameter fine tuning based on scheduling feedback; and S9, continuously applying the fine-tuned model to load scheduling control. According to the invention, power load accurate scheduling and strategy efficient adaptive optimization are realized, and system responsiveness and operation stability are improved.
Owner:ZHEJIANG JUHUA THERMAL POWER CO LTD

Reinforcement learning method and system for source network load storage collaborative multi-scene optimization

The invention discloses a reinforcement learning method and system for source network load storage collaborative multi-scene optimization, and the method comprises the steps: constructing a power distribution network optimization model with the minimum cost, and converting a mixed integer nonlinear problem into a solvable mixed integer second-order cone optimization problem; converting a mixed integer second-order cone optimization problem into a reinforcement learning decision model, and performing multi-round assignment on boundary condition parameters by using different operation scene data to form a multi-scene training task set; designing a reinforcement learning decision model multi-scene training loss function, combining the reinforcement learning decision model to construct a strategy neural network, an evaluation network neural network and a scene representation embedded neural network, and completing reinforcement learning decision adaptive to multi-scene optimization; and training a power distribution network multi-scene optimization decision model based on the multi-scene training task set, the multi-scene training reinforcement learning loss function and the neural network structure, and deploying the power distribution network multi-scene optimization decision model to an actual system to complete reinforcement learning decision model application oriented to source network load storage collaborative multi-scene optimization.
Owner:CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +2

Digital twin workshop management and control method and system based on data multi-layer fusion

The invention relates to a digital twin workshop management and control method and system based on data multilayer fusion, and belongs to the field of service-oriented industrial system data science, and the method comprises the steps: obtaining the global multi-source heterogeneous data of a digital twin workshop, and carrying out the data classification, data processing and data storage of the global multi-source heterogeneous data, carrying out data-level fusion to obtain digital twin workshop standard data; extracting standard data of the digital twin workshop, performing semantic association, feature extraction and feature analysis on the standard data of the digital twin workshop, and performing feature level fusion to obtain key state features of the digital twin workshop; the method comprises the following steps: constructing a decision-making model oriented to a digital twin workshop, inputting key operation features of the digital twin workshop, performing decision-making level fusion based on the decision-making model, decision-making optimization and decision-making method solution, and outputting a workshop production management and control decision. The method can provide support for digital twin workshop operation data governance analysis and production management and control accurate decision.
Owner:BEIHANG UNIV

Industrial equipment energy efficiency evaluation and maintenance decision-making method and system based on reinforcement learning

The invention provides an industrial equipment energy efficiency evaluation and maintenance decision-making method and system based on reinforcement learning, and relates to the technical field of reinforcement learning, and the method comprises the steps: collecting equipment operation data and extracting features, constructing a maintenance time sequence decision-making model based on deep reinforcement learning, and achieving the model optimization through the combination of transfer learning; and generating a maintenance decision scheme according to the energy efficiency index analysis and the time sequence causal relationship, executing maintenance, and recording process data for online learning and updating of the model. The energy efficiency management level of industrial equipment is improved, the service life of the equipment is prolonged, and the operation and maintenance cost is reduced.
Owner:CHANGZHOU RUIWU TECH CO LTD

Coordinated control method and system for comprehensive energy multi-agent coordinated group control and autonomous decision

The invention discloses a coordinated control method and system for comprehensive energy multi-agent coordinated group control and autonomous decision making. The method comprises the following steps: establishing a relaxation strategy of an agent corresponding to a distributed power supply; the main coordinator iteratively solves to determine a global coordination signal according to the global operation data of the integrated energy system and the operation data of each agent by taking the lowest total cost, the highest system energy efficiency, the minimum carbon emission and the minimum global penalty coefficient meeting the global relaxation constraint as a relaxation global optimization target; the minimum operation cost of each intelligent agent, the minimum response deviation to a global coordination signal and the minimum deviation between an actual state and a reference state are taken as control targets; establishing an electricity price autonomous decision-making model of each agent and a response power boundary autonomous decision-making model for a global coordination signal; and based on the control target of each agent, the electricity price autonomous decision-making model and the response power boundary autonomous decision-making model, predicting to obtain a control target value of each agent, and realizing balance between a global optimization target and local autonomy.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +2

Training method for action decision-making model of intelligent agent, and action decision-making method and apparatus

PCT designated stage expiredWO2025148822A1Artificial lifeMachine learningEngineeringDecision taking
Disclosed in the present application are a training method for an action decision-making model of an intelligent agent, and an action decision-making method and apparatus. The training method comprises: constructing a state topology graph on the basis of historical trajectories; on the basis of the state topology graph, training an action feedback model of an intelligent agent, wherein the action feedback model is used for providing a feedback signal of an environment where the intelligent agent resides in respect of an action executed by the intelligent agent; on the basis of the state topology graph and the action feedback model, training an action value model of the intelligent agent, wherein the action value model is used for providing an estimated action value of the action executed by the intelligent agent, and the estimated action value indicates a metric value used for measuring the impact of the action executed by the intelligent agent on the environment; and on the basis of the action value model, training an action decision-making model of the intelligent agent, wherein the action decision-making model is used for deciding an action to be executed by the intelligent agent in a given state.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Bridge group maintenance priority dynamic decision-making method and device based on reinforcement learning

The invention provides a bridge group maintenance priority dynamic decision-making method and device based on reinforcement learning, and relates to the technical field of bridge intelligent maintenance. The method comprises the following steps: constructing a topological structure of a bridge network and a road; defining a state space, a maintenance action space and a state transition matrix of the bridge; defining a reliability index corresponding to the state of the bridge, and designing a comprehensive reward function based on maintenance cost, asset risk and traffic network capacity loss risk based on the topological structure; constructing a bridge maintenance decision problem; the method comprises the following steps of: describing a bridge maintenance decision problem as a Markov decision process, establishing a pointer network strategy model by adopting a pointer network, and training the pointer network strategy model by adopting an Actor-Critic algorithm to obtain a maintenance decision model based on reinforcement learning; and training the maintenance decision model based on reinforcement learning until convergence, and outputting a bridge maintenance action sequence under limited constraints. By adopting the method, the limitation problem of traditional single bridge assessment can be solved.
Owner:UNIV OF SCI & TECH BEIJING

Training method of action decision-making model of intelligent agent, action decision-making method and device

The invention discloses a training method of an action decision model of an intelligent agent and an action decision method and device, and belongs to the technical field of computers. According to the method, the state topological graph is constructed on the basis of the historical trajectory, experience distribution of actions of the intelligent agent can be fully reflected, the information utilization rate of the historical trajectory is higher, more information is brought, the action feedback model is guided and trained on the basis of the state topological graph, the accuracy of the action feedback model is improved, and the accuracy of the action feedback model is improved. A state topological graph, an action feedback model and a training process of a constraint action value model are combined to obtain an action value model with better accuracy and better performance, and the action value model is utilized to assist in training to obtain an action decision model with better accuracy, so that an accurate decision on which action is executed by an intelligent agent in a given state is facilitated; and the action decision model can be combined with a large model to mutually promote training, so that the performance of the two parties is improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Multi - supply - chain scheduling method and system based on global Critic multi - agent algorithm

The present invention relates to the technical field of supply chain scheduling, and discloses a multiple supply chain scheduling method and system based on a global Critic multi-agent algorithm. The method includes: constructing a multi-agent supply chain environment; introducing a global Critic network, and using the multi-agent deep deterministic policy gradient algorithm to perform reinforcement learning on the multi-agent supply chain environment to construct a multi-agent decision-making model; inputting supply chain demand data into the multi-agent decision-making model for optimized scheduling processing to generate a scheduling strategy that maximizes the overall profit of the supply chain. The present invention introduces a global Critic network to reduce multi-agent competition and optimize the overall performance. The multi-agent deep deterministic policy gradient algorithm is used to construct a decision-making model to adapt to complex supply chain environments. With the goal of maximizing the overall profit, it solves the problems of low efficiency and low returns. The environment has strong adaptability and can handle scheduling problems of different scales. By continuously learning to optimize decisions, the reliability is improved, and the efficiency and returns are significantly enhanced.
Owner:GUANGDONG UNIV OF TECH

Single-agent decision-making method and device based on deep reinforcement learning

The invention provides a single agent decision-making method and device based on deep reinforcement learning, and the method comprises the steps: obtaining agent sample data in an experience playback buffer pool, the sample data at least comprising an observation state, a next observation state, a decision-making action and a reward value of an agent in an environment; inputting the agent sample data into the estimation network and the target network for training; the estimation network comprises a strategy network and an evaluation network, and the target network comprises a target strategy network and a target evaluation network; respectively calculating a strategy gradient of the estimation network and a loss function of the target network; if the strategy gradient and the loss function meet preset conditions, the network model corresponding to the strategy gradient and the loss function is determined as the decision model of the agent, an actor-commentator reinforcement learning architecture is adopted, the problem that correlation exists before and after each parameter updating of a neural network in an actor-commentator structure is solved, and the decision model of the agent is determined to be the decision model of the agent. And meanwhile, the defect that deep reinforcement learning cannot be used for continuous actions is overcome.
Owner:CSSC SYST ENG RES INST

Green product full life cycle analysis system based on big data

The invention relates to the technical field of green product analysis, and discloses a green product full life cycle analysis system based on big data. The system comprises a data acquisition module, a preprocessing module, an environmental impact assessment module, a resource consumption optimization module, an association modeling and result generation module and a visual interaction module. The system collects data of all stages of a full life cycle and related element data of environment, resources, carbon emission and the like, after preprocessing of abnormal value elimination, entropy weight method weighting and the like, environment influence is dynamically evaluated through a fuzzy membership function, resource consumption is subjected to multi-objective optimization by utilizing a genetic algorithm, and the resource utilization rate is improved. A key stage is determined by means of grey relational degree, an analysis result is generated in combination with a multi-target decision model, and finally the analysis result is visually presented in the form of a three-dimensional map and a dynamic thermodynamic diagram. The system realizes efficient analysis and comprehensive evaluation of full-life-cycle data of green products, and provides scientific support for product environment performance optimization and decision making.
Owner:CHINA NAT INST OF STANDARDIZATION

Multi-robot cooperation interaction control method and system based on deep learning

The invention discloses a multi-robot cooperation interaction control method and system based on deep learning, and the method comprises the steps: obtaining historical interaction data, carrying out the environment state coding and feature extraction, and carrying out the correlation score matrix calculation, and obtaining a dynamic interaction weight coefficient; according to the situation representation vector and the interaction weight coefficient, performing future behavior prediction to obtain intention probability distribution data; performing encryption processing on the sensitive information according to historical interaction data to obtain encrypted shared data, performing weighted average aggregation in combination with pre-trained decision model parameters to generate a global collaborative decision model, and inputting the encrypted shared data into the global collaborative decision model to obtain a collaborative action scheme; and according to the cooperative action scheme and the intention probability distribution data, performing multi-constraint task allocation optimization to obtain an optimal task allocation scheme, and performing adaptive path planning to obtain a final action path. According to the method, real-time efficient collaborative decision making in a high-dynamic complex environment can be realized.
Owner:JIANGSU SMART WORKSHOP TECHNOLOGY RESEARCH INSTITUTE CO LTD

Patent value evaluation method based on Markov decision

The invention relates to the technical field of patent value evaluation. The invention provides a Markov decision-based patent value evaluation method, which is used for constructing a state-action-reward-transfer decision-making model for patent life cycle dynamic characteristics. Through multi-source data processing, state construction, strategy screening, value function iteration and optimal strategy generation, quantitative evaluation of patent dynamic values and operation suggestion output are realized, and the problem that strategy evolution and dynamic earnings cannot be described by an existing method is solved.
Owner:XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV

Client autonomous decision-making mechanism-based machine authority management method and system

The invention provides a machine authority management method based on a customer autonomous decision-making mechanism, comprising the following steps: firstly, acquiring a current account authority configuration set of a target machine, the current account authority configuration set comprising a plurality of authority level nodes and corresponding operation range identifiers, associated account identifiers and verification features, then performing authority feature extraction on the current account authority configuration set; generating a dynamic decision node set representing permission change trigger conditions and verification paths on the basis of a preset autonomous decision model according to hierarchical constraint features and operation behavior features of the permission level nodes, and calling a permission verification model to perform multi-dimensional matching processing on the related characteristics and the node set, generating a permission change verification result, updating a current account permission configuration set according to the permission change verification result, and finally synchronizing the updated configuration set to a target machine operation control module to enable the permission to take effect, thereby realizing flexible and intelligent machine permission management.
Owner:SHENZHEN BANGZHENG PRECISION MACHINERY CO LTD

Optimization system and method for participation of electric vehicle cluster in electric power standby market

The invention discloses an optimization system and method for participation of an electric vehicle cluster in an electric power standby market, and is applied to the field of an electric power auxiliary service market, and the system comprises a market uncertainty modeling module which is used for constructing a joint uncertainty model of market price and standby calling probability, wherein the model comprises a probability distribution model and a conditional probability model; the risk quantitative evaluation module is used for evaluating risk exposure degrees of different decision-making schemes, including establishing a multi-period risk accumulation model, and comprehensively evaluating balance indexes of expected income, risk openness and opportunity cost; the robust optimization decision-making module constructs a decision-making model according to the market uncertainty modeling module and the risk quantitative evaluation module, and generates an optimal strategy that the electric vehicle cluster participates in the standby market; the adaptive learning optimization module is used for optimizing the decision model through reinforcement learning; according to the invention, risk-controllable revenue maximization can be realized, so that the electric vehicle cluster can efficiently participate in the electric power standby market.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Task request processing method, computing device, electronic device and storage medium

The invention discloses a task request processing method, computing equipment, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: in response to a received task request sent by a target client, constructing a cue word based on the task request; the cue word is input into an intelligent decision model, decision reasoning is conducted on the cue word through the intelligent decision model, an initial processing result and a calling decision result of the task request are obtained, and the calling decision result is used for representing whether the task request needs to be processed through a calling tool or not; and processing the task request based on the initial processing result and the calling decision result to obtain a target processing result. The technical problem of relatively low task processing efficiency in related technologies is solved.
Owner:SHANGHAI TAOXINBAO NETWORK TECHNOLOGY CO LTD

Super-multi-target vehicle path planning method for deep reinforcement learning assisted evolution

The invention discloses a super-multi-target vehicle path planning method based on deep reinforcement learning assisted evolution, and aims to dynamically guide a search direction through deep reinforcement learning for multi-target optimization and path planning problems. The core of the method comprises the following steps: 1) constructing a super-multi-target vehicle path problem mathematical model; 2) designing a deep Q network as a decision model, and adaptively selecting a high-potential search direction in iteration; 3) designing a global search strategy, and balancing global search and local optimization; 4) generating a filial generation and evaluating the fitness, and executing environment selection to update the population; and 5) calculating a reward value according to a population evolution effect, and optimizing DQN network parameters. According to the method, through a collaborative mechanism of reinforcement learning and an evolutionary algorithm, a search strategy is dynamically adjusted, the convergence speed and stability of the algorithm are effectively improved, the adaptability to a complex path planning problem is enhanced, and the method has remarkable advantages in solving efficiency, solution set quality and dynamic environment response.
Owner:HUNAN UNIV

Communication radiation source target confrontation decision-making method based on unsupervised learning

The invention discloses a communication radiation source target confrontation decision-making method based on unsupervised learning. The method comprises the following steps of obtaining an IQ signal of a communication signal and performing signal preprocessing; training a contrast predictive coding CPC model based on the IQ data after signal preprocessing, and extracting IQ data features; inputting the IQ data features into a ResNet network for closed set individual classification, and outputting closed set scores; correcting the closed set score based on an OpenMax algorithm to obtain an open set score, and identifying a radiation source individual according to the open set score; constructing a target adversarial decision model based on a GBDT algorithm, extracting individual features of a communication radiation source and interference adversarial parameter features, and constructing a data set to train the target adversarial decision model; and inputting individual characteristic parameters of the communication radiation source to the trained target adversarial decision model, and outputting interference adversarial parameters. According to the method, the identification accuracy of the communication signals is improved, and the robustness in the face of signals of unknown categories is enhanced.
Owner:UNIT 75737 OF THE CHINESE PEOPLES LIBERATION ARMY

Deep counterfeit multimedia identification method and system based on artificial intelligence

The invention provides a deep counterfeit multimedia identification method and system based on artificial intelligence, and relates to the technical field of artificial intelligence security. According to the method, high-precision forged content detection is realized through multi-modal feature extraction and link generation reasoning. The method specifically comprises the steps of multi-modal data acquisition and feature extraction, generated link reasoning and knowledge graph adaptation, multi-modal correlation analysis and dynamic threshold setting, forged path rule construction and decision model training, and self-adaptive identification strategy generation. According to the method, multi-modal features are fused, links are generated based on dynamic analysis of the knowledge graph, adaptive threshold judgment and targeted countering are achieved, the problems that a traditional method is insufficient in generalization ability, low in calculation efficiency and the like are solved, and the method can be widely applied to content authenticity identification in the fields of social media, public opinion and the like.
Owner:TIANJIN NAT CYBERNET SECURITY CO LTD

Power grid intelligent decision-making method, model construction method, equipment and medium

The invention provides a power grid intelligent decision-making method, a model construction method, equipment and a medium, and the method comprises the steps: carrying out the feature extraction of historical power grid metadata and power grid regulation and control strategy data, carrying out the clustering of the obtained historical data features, generating a corresponding label, constructing a knowledge graph through employing the label and strategy features, and carrying out the calculation of the knowledge graph. And the power grid intelligent decision-making model is constructed based on the knowledge graph and the deep learning algorithm, so that the internal semantic relation of the historical power grid metadata can be accurately, fully and quickly mined without relying on the labeling of the rule basis by utilizing the automatic label generation and the knowledge graph construction; therefore, reliable basis and guarantee are provided for timely and accurate power grid intelligent decision making.
Owner:STATE GRID HUNAN ELECTRIC POWER COMPANY LIMITED +2

Business process configuration method based on Activiti and AI decision

The invention relates to the technical field of artificial intelligence, and discloses a business process configuration method based on Activiti and AI decision, and the method comprises the steps: collecting the historical execution data of a business process through a process instance monitoring module, generating a process behavior feature data set, carrying out the training processing of an AI decision model based on the process behavior feature data set, and carrying out the analysis of the AI decision model. A business process decision tree model is constructed, AI decision routing labels are configured for decision nodes in an Activiti process engine, non-intrusive intelligent decision integration is achieved, and when a process instance is executed to a key decision node, a system automatically calls the pre-trained business process decision tree model to conduct real-time path analysis and generate an optimal dynamic routing instruction. The problem of decision stiffness caused by the fact that a traditional workflow depends on artificial experience configuration is solved, the accuracy and adaptability of flow branch selection in a complex service scene are improved, and meanwhile the compatibility and stability of an original flow engine are guaranteed.
Owner:SHANGHAI CAPITAL SOFTWARE CO LTD

Game-oriented multi-agent decision model architecture method and device, equipment and medium

PendingCN120044981AInternal combustion piston enginesMachine learningHybrid systemDecision-making software
The invention provides a game-oriented multi-agent decision-making model architecture method and device, equipment and a medium, and the method comprises the steps: firstly constructing a multi-agent decision-making framework through combining a hierarchical conclusion thought with a containing structure according to a knowledge rule and data learning mixed system; and then, according to the operator type, the training direction of the reinforcement learning algorithm in the action is guided through a knowledge rule strategy, and a task decision is designed. Meanwhile, based on a multi-agent near-end strategy optimization algorithm, micro-operation actions of the agents are constructed through data learning, and action decisions are designed. Finally, a task decision and an action decision are introduced into a multi-agent decision framework, a multi-agent model is constructed, the model is verified through game confrontation, and architecture of multi-agent decision software is completed. According to the embodiment of the invention, hybrid-driven multi-agent construction is realized, and the efficient response capability of the agents in a complex wide-area environment is ensured. And the decision-making ability and the game level of the intelligent agent are improved.
Owner:CSSC SYST ENG RES INST +1

Traffic industry carbon reduction method and system based on multi-Agent-SD decision model simulation

The invention discloses a traffic industry carbon reduction method and system based on multi-Agent-SD decision model simulation, and relates to the technical field of urban traffic, and the method comprises the steps: obtaining a historical data set of the traffic industry in a to-be-researched region; processing the historical data set, and dividing the historical data set into a training data set and an inspection data set; constructing a multi-Agents-SD decision model, performing comparative analysis on simulation data simulated by the training data set and the test data set, and calibrating parameters of the multi-Agents-SD decision model; corresponding parameters are designed according to different development situations, the adjusted parameters are input into the multi-Agent-SD decision model for simulation, a differentiation result is obtained, and then the traffic industry carbon reduction strategy is obtained. According to the method, the advantages of the two models are combined, the multi-Agents model for exploring micro individual behavior changes and the SD model for a macroscopic complex system are combined, and a more suitable traffic carbon reduction scheme is researched from multiple angles.
Owner:TIANJIN UNIV +2

Government affair big model-based data government method and system

The invention discloses a data government method and system based on a large government affair model, and belongs to the technical field of deep learning, and the method comprises the steps: government affair data collection: obtaining basic data from a data source for implementing data government, the basic data comprising historical data and current data; government affair data treatment: preprocessing the collected basic data, and integrating the data to obtain a data sequence; a platform adaptively selects a preset deep learning framework according to the government affair scene features, and a staged parameter optimization strategy is combined to train and generate a government affair decision model with a domain cognitive ability; and intelligent analysis: deploying a semantic understanding engine, converting the unstructured government affair text into data vector representation which can be analyzed by a machine, and outputting a data analysis result through a multi-modal feature fusion technology. According to the invention, the data management difficulty in a government affair scene can be reduced, and deep transformation of intelligent government affairs and establishment of an intelligent decision support system are promoted.
Owner:SHANDONG INSPUR CLOUD GOVERNMENT INFORMATION TECHNOLOGY CO LTD

Unmanned aerial vehicle cluster formation cooperative multi-target distribution method and system based on multi-agent layered reinforcement learning

The invention discloses an unmanned aerial vehicle cluster formation cooperative multi-target distribution method and system based on multi-agent layered reinforcement learning. Based on the idea of hierarchical reinforcement learning and multi-agent collaborative decision-making, a hierarchical reinforcement learning architecture of top-layer decision-making and bottom-layer execution is provided, a top-layer decision-making model is a single-agent system, the position, target location and transportation target information of an unmanned aerial vehicle cluster are analyzed, and task targets of all unmanned aerial vehicles are decided; the bottom layer execution model is a multi-agent system, and each agent decides a moving path according to the landform state observed by the agent. According to the method, the multi-target allocation task is clearly divided into two sub-tasks, namely task allocation and autonomous execution, so that the training time is shortened and the execution capability of the unmanned aerial vehicle cluster is improved.
Owner:WUHAN UNIV

Method, apparatus, device, medium, and program product for training decision model

This disclosure provides a method, an apparatus, a device, a medium, and a program product for training a decision model. The method includes: determining a first policy using a supervised learning model and a second policy using a reinforcement learning model within the decision model based on training data; determining an imitation learning loss based on a difference between the first policy and the second policy; and training the decision model based on both the imitation learning loss and a reinforcement learning loss corresponding to the second policy. By combining the imitation learning loss and the reinforcement learning loss, a human-like decision model with excellent performance may be obtained, leveraging the expert data utilization capability of supervised learning and the strong generalization capacity of reinforcement learning. In some embodiments, the trained model is applied to autonomous driving for tasks such as lane-changing.
Owner:YINWANG INTELLIGENT TECHNOLOGIES CO LTD

Cognitive interference decision-making method based on PPO-RNN algorithm and interference end

The invention discloses a cognitive interference decision-making method based on a PPO-RNN algorithm and an interference end. The method comprises the following steps: acquiring real-time state information of a communication party serving as an interference target; the real-time state information is input into a trained decision model based on a PPO-RNN algorithm, a real-time interference strategy is obtained, the decision model comprises a strategy network and a value network, and the strategy network comprises two full connection layers and a recurrent neural network layer which are connected in sequence. According to the method provided by the invention, the full-connection network and the cyclic convolutional network are introduced into the decision model, so that the cyclic convolutional network can be utilized to capture the time sequence relevance of the electromagnetic signal of the interference target, and the decision effect is improved; and the training speed of the decision model can be improved and the convergence time of the decision model can be shortened by the cyclic convolutional network, so that the purpose of rapid decision can be realized, the real-time performance is enhanced, and the model stability can be improved.
Owner:XIDIAN UNIV

Multi-agent conflict-free path planning method based on directed space-time network and forward direction decision

The invention discloses a multi-agent conflict-free path planning method based on a directed space-time network and forward direction decision, which comprises the following steps: 1, acquiring environment information, constructing a grid network according to the environment information, and constructing a directed space-time road network for the grid network by using the space-time network, initializing starting and ending point information of the robot group according to the directed space-time road network; 2, establishing a forward direction decision model of the agent group, and avoiding conflicts by using constraint conditions of the model; and 3, solving a forward direction decision model of the agent group, and planning an overall optimal path. According to the method, the directed space-time road network can be constructed by utilizing the space-time network to accurately describe the environment, the agents and the tasks, the path of the agents is planned by utilizing the advancing direction decision model, and the possibility of multi-agent conflict is avoided, so that the feasibility of the path is ensured, and the efficiency and the safety of multi-agent path planning can be improved.
Owner:HEFEI UNIV OF TECH