Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

83 results about "Decision problem" patented technology

In computability theory and computational complexity theory, a decision problem is a problem that can be posed as a yes-no question of the input values. An example of a decision problem is deciding whether a given natural number is prime. Another is the problem "given two numbers x and y, does x evenly divide y?". The answer is either 'yes' or 'no' depending upon the values of x and y. A method for solving a decision problem, given in the form of an algorithm, is called a decision procedure for that problem. A decision procedure for the decision problem "given two numbers x and y, does x evenly divide y?" would give the steps for determining whether x evenly divides y. One such algorithm is long division. If the remainder is zero the answer is 'yes', otherwise it is 'no'. A decision problem which can be solved by an algorithm is called decidable.

Automatic polygon and quadrilateral mesh subdivision method based on deep reinforcement learning

The invention discloses a polygonal and quadrilateral mesh automatic subdivision method based on deep reinforcement learning, relates to the field of computational geometry and mesh generation, and constructs an automatic subdivision framework fusing geometric priori knowledge and a data driving strategy. According to the method, nine basic topology filling templates covering triangles to hexagons are predefined as discrete action spaces, and continuous geometric segmentation is converted into a sequence decision problem; learning a mapping relation between the polygon state characteristics and the optimal template selection strategy by using a deep Q network, and balancing exploration and utilization through a dynamic epsilon-greedy mechanism; and in combination with a vertex number priority scheduling strategy and a multi-dimensional quality award function, guiding an intelligent agent to adaptively generate a high-quality quadrilateral grid. According to the method, full-automatic and high-robustness subdivision of the complex polygon area is achieved, the grid orthogonality and the length-width ratio quality are remarkably improved, manual intervention is avoided, and the method is suitable for engineering scenes such as finite element analysis.
Owner:CALCULATION AERODYNAMICS INST CHINA AERODYNAMICS RES & DEV CENT

Cost evaluation reinforcement learning-based unmanned aerial vehicle trajectory optimization and power distribution method, system and device, and medium

The invention discloses an unmanned aerial vehicle trajectory optimization and power distribution method, system and device based on cost evaluation reinforcement learning, and a medium, and the method comprises the steps: constructing an unmanned aerial vehicle data service scene under safety constraints, modeling a multi-objective optimization problem of unmanned aerial vehicle trajectory optimization and power distribution, and designing an unmanned aerial vehicle data service evaluation index; converting a multi-objective optimization problem into a constrained Markov decision problem; respectively designing a state space, an action space, an instant reward function and a corresponding cost function; constructing a cost evaluation critic network parallel to the reward evaluation critic network, training and updating network parameters of the unmanned aerial vehicle by using a cost evaluation SAC algorithm according to a constrained Markov decision problem, and calculating trajectory optimization and power distribution parameters of the unmanned aerial vehicle; the system, the equipment and the medium are used for implementing the method. According to the invention, ground user data service requirements of a data service scene under random and dynamic security constraints are met.
Owner:XIDIAN UNIV

Enterprise appeal intelligent sensing and closed-loop processing system based on multi-mode AI

The invention discloses an enterprise appeal intelligent perception and closed-loop processing system based on multi-modal AI, and relates to the technical field of intelligent government affairs, the system realizes accurate semantic understanding and deep intention recognition of multi-modal fusion, and the intelligent perception and closed-loop processing of enterprise appeals are realized through a cross-modal attention mechanism and a unified semantic representation model. According to the method, deep fusion and complementary analysis of multi-source heterogeneous data such as voices, texts, images and the like are realized, the limitation of a traditional single-mode or simple splicing mode is broken through, the accuracy and robustness of intention recognition are remarkably improved, and misjudgment is fundamentally reduced; the method comprises the following steps: constructing a dynamic self-adaptive routing mechanism based on Actor-Critic reinforcement learning, constructing a work order assignment problem into a sequence decision problem, driving a model to learn a dynamic fusion and weight distribution strategy of multiple decision factors through a reward mechanism, and adopting an online strategy iteration optimization mechanism to realize an optimal assignment decision, so as to improve the work order assignment efficiency. And the shunting accuracy and efficiency are obviously improved.
Owner:SICHUAN ENRISING INFORMATION TECH CO LTD

E-commerce platform product selecting and pricing method based on multi-agent reinforcement learning

The invention discloses an e-commerce platform product selection and pricing method based on multi-agent reinforcement learning, aiming at the game decision problem between an e-commerce platform and a supplier, a platform agent and a supplier agent are constructed, the platform agent uses a deep Q network DQN to learn a product selection strategy, and the supplier agent uses an Actor-Critic algorithm to learn a pricing strategy. Through Stackelberg master-slave game modeling, a platform is used as a leader to make a decision for selection, and a supplier is used as a follower to make a decision for pricing after observation. A state space (including commodity sales volume, inventory, user score, supplier fulfillment rate and the like), an action space (discrete commodity selection and continuous pricing) and a multi-target reward function (balance profit, user satisfaction, inventory turnover and the like) suitable for an e-commerce scene are designed. The strategies of the two parties are converged to game equilibrium through alternate training, and online continuous learning is supported to adapt to a dynamic market environment. Experiments show that compared with an existing method, the method has the advantages that the overall income can be improved by 18-30%, the win-win situation of the platform and the suppliers is achieved, and the blank of multi-agent reinforcement learning in e-commerce product selection and pricing scenes is filled.
Owner:SHENZHEN WEIRUIHAO TECHNOLOGY CO LTD

Unmanned sweeping robot navigation method based on multi-sensor fusion

The invention discloses an unmanned sweeping robot navigation method based on multi-sensor fusion, and belongs to the technical field of sweeping robots. According to the method, an environment sensing system is constructed through multi-source sensor fusion, and efficient path planning and obstacle avoidance decision making in a dynamic environment are realized in combination with a quantum heuristic behavior decision engine; a traditional decision problem is converted into quantum state superposition operation through a quantum state probability amplitude model, and dynamic coupling of environment perception and a task target is achieved through the joint effect of a complex phase angle and a weight coefficient; and the dynamic collapse equation improves the adaptive capacity of behavior selection. The sweeping efficiency, safety and environment adaptability of the sweeping robot in a complex dynamic environment are remarkably improved, and the sweeping robot is suitable for various scenes such as families and offices.
Owner:HUAIYIN INSTITUTE OF TECHNOLOGY

Multi-hop reasoning method based on dynamic reasoning guidance and multistage self-feedback retrieval

A multi-hop reasoning method based on dynamic reasoning guidance and multistage self-feedback retrieval belongs to the field of natural language processing, and comprises the following steps: deconstructing a multi-hop reasoning process into a target-oriented sequence decision problem, carrying out dynamic reasoning guidance by using a large language model, generating a sub-problem sequence matched with a reasoning progress in real time, and carrying out multi-level self-feedback retrieval on the sub-problem sequence; target document retrieval is guided, and sub-questions are dynamically generated; according to the generated sub-questions, obtaining associated documents by adopting a three-level collaborative retrieval mechanism; and performing information refining on the associated document through a large language model, fusing the refined information into an inference chain, and performing inference to generate an answer. The invention further discloses a multi-hop reasoning system, a storage medium and a computer program product. The method aims at solving the complex multi-hop problem that multiple dispersed knowledge fragments need to be integrated, high-accuracy and high-efficiency reasoning is achieved, the retrieval requirement is dynamically generated through an explicit thinking chain guiding mechanism, and evidence obtaining is optimized and redundant information is filtered in combination with a three-level self-feedback retrieval mechanism.
Owner:XI AN JIAOTONG UNIV

Multi-agent task planning and executing method and decision making system

The invention provides a multi-agent task planning and executing method and a decision making system, and relates to the technical field of artificial intelligence and automatic control. According to the method, a complex multi-agent decision problem is decomposed into an upper-layer semantic planning subsystem and a lower-layer action control subsystem which are relatively independent and collaborative optimization; the upper-layer vision-language model planner generates sub-target sequences and track priori conforming to tactical intentions from a global perspective by utilizing the powerful vision understanding and language reasoning capabilities of the upper-layer vision-language model planner; and the reinforcement learning strategy controller at the lower layer focuses on executing specific control actions under the guidance of the sub-targets. Meanwhile, the system can optimize the final task performance and the process execution quality at the same time by introducing an evaluation mechanism and a combined reward signal based on the completion degree of the sub-targets. Environmental native rewards ensure that the system is optimized towards a final objective.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Flexible job shop multi-target scheduling method and system based on preference driving

The invention belongs to the technical field of workshop production scheduling, and discloses a flexible job workshop multi-target scheduling method and system based on preference driving, and the method comprises the steps: defining a target function and a constraint condition based on obtained information, and constructing a workshop scheduling model; converting a workshop scheduling problem into a Markov decision problem, designing a reward function, and creating a preference pool; selecting a preference vector from the preference pool, and respectively inputting the preference vector and the system state into a strategy network and a value network for iterative training to obtain a trained strategy network and a trained value network; and obtaining current preferences for different dispatches and a current state of the system, inputting the current preferences and the current state of the system into the trained strategy network, obtaining action probability distribution in a given state, and selecting a dispatching action with the maximum probability to obtain a corresponding optimal dispatching scheme. According to the method, the requirements of tire enterprises in different situations are met, high flexibility is achieved, the nonlinear target utility function is designed, and a tire workshop scheduling scheme with higher quality can be found out easily.
Owner:SHANDONG UNIV +1

Electric heavy truck cooperative scheduling optimization method based on multi-agent reinforcement learning

An electric heavy truck collaborative scheduling optimization method based on multi-agent reinforcement learning belongs to the technical field of demand side resource scheduling and intelligent transportation, and comprises the following steps: firstly, modeling an electric heavy truck and a vehicle network interaction system, and depicting a constraint relationship among a transportation task, charging and discharging and power grid interaction; secondly, constructing a Markov decision process, defining a state, an action, a state transition rule and a reward function of a single agent, and converting a scheduling problem into a decision problem solvable by reinforcement learning; then, based on an MAPPO algorithm, a distributed execution-centralized judgment architecture is adopted, a multi-agent collaborative decision-making model is trained, decision-making feasibility is ensured through action masks, and collaborative optimization is achieved through global value evaluation; and finally, realizing dynamic scheduling and optimization decision of the electric heavy truck based on the training model. According to the invention, the transportation efficiency and the power grid interaction cost can be balanced while the transportation timeliness and the battery safety constraint are satisfied, and the overall operation economy and collaboration of the electric heavy truck fleet are improved.
Owner:NORTH CHINA ELECTRIC POWER UNIV +1

Multi-target layered depth optimization method for illumination control

The invention discloses a multi-target hierarchical depth optimization method for illumination control, and belongs to the technical field of illumination control, and the optimization method specifically comprises the following steps: I, collecting and preprocessing various types of information at edge nodes of an illumination area, and extracting semantic information describing a real-time scene from the preprocessed various types of information; according to the method, the influence of sensing noise, missing data and heterogeneous sampling on decision is remarkably reduced, the accuracy and stability of environment sensing are improved, the illumination control strategy has the immediate response capability and the prospective adjustment capability, the problem that a traditional fixed weight method is insufficient in adaptability in different time and different scenes is solved, and the method is suitable for popularization and application. According to the method, the targets of energy conservation, safety, comfort and the like can be dynamically balanced according to the actual environment, the complex decision problem is effectively decomposed, the convergence speed and robustness of the strategy are improved, meanwhile, sudden performance drop caused by scene extrapolation is avoided, and the generalization ability of the system in the complex urban environment is remarkably improved.
Owner:NANJING LICON LOT TECH CO LTD

Method for dynamically updating housekeeping service knowledge graph based on reinforcement learning

The invention discloses a reinforcement learning-based household service knowledge graph dynamic updating method, and relates to the technical field of artificial intelligence and data processing. The method comprises the following steps: initializing a knowledge graph into a reinforcement learning environment; defining an agent state including a current entity, target information and a historical path; defining an action space of the intelligent agent; designing a reward function including a final reward and an intermediate reward; training an intelligent agent by adopting a strategy network; and inferring and verifying new knowledge by using the trained agent, and adding the new knowledge back to the knowledge graph. The knowledge graph updating process is modeled into a sequence decision problem, automatic discovery of knowledge and dynamic evolution of the graph are achieved through autonomous exploration and learning of the reinforcement learning agent, and the problems that a traditional updating method is high in cost and poor in timeliness are effectively solved.
Owner:LIAONING UNIVERSITY

A multi-modal collection and evolution graph analysis method and system of a policy text

The application belongs to the technical field of natural language processing and knowledge graph, and discloses a kind of multi-modal acquisition and evolution graph analysis method and system of policy text. In view of the passive decision problem caused by the rigidity of existing policy analysis, semantic understanding and lack of dynamic correlation, the application adopts visual and DOM feature weighted fusion and reinforcement learning agent to realize multi-modal adaptive acquisition and joint alignment in anti-crawling environment;Use large language model for semantic deconstruction, and convert implicit logic into computer-readable explicit logic expression through thought chain reasoning;Construct a policy evolution knowledge graph with time and validity level attributes, perform satisfiability analysis to output conflict detection results;Map the target object portrait to the graph and the logic expression to evaluate item by item, and output intelligent decision deduction results combined with quantitative dimensions. The application realizes the whole life cycle intelligent analysis of policy data acquisition, deep logic analysis and active deduction.
Owner:POWERCHINA ZHONGNAN ENG

Unmanned aerial vehicle resource allocation method, system and device based on multi-agent reinforcement learning, and medium

An unmanned aerial vehicle resource allocation method, system and device based on multi-agent reinforcement learning, and a medium, the method comprising: establishing an optimization problem of multi-unmanned aerial vehicle resource allocation for a multi-modal unmanned aerial vehicle assisted communication-sensing integrated data service scene, and designing an unmanned aerial vehicle data service evaluation index of average peak information age; expressing an optimization problem as a Markov decision problem, and designing a state space, an action space and a corresponding instant reward function; the method comprises the following steps: designing a conditional variation auto-encoder, representing a coupling relationship between discrete actions and continuous actions, combining an MATD3 multi-agent reinforcement learning algorithm to obtain an MATD3 algorithm represented by the coupling relationship, and training and updating network parameters of an unmanned aerial vehicle by using the MATD3 algorithm represented by the coupling relationship according to a Markov decision model. The method is applied to a data service scene for communication resource allocation; according to the method, the data service requirements of the ground users in a random and dynamic communication and sensing integrated scene can be better met.
Owner:XIDIAN UNIV

Traffic signal control optimization method and device and electronic equipment

The invention provides a traffic signal control optimization method and device and electronic equipment, and the method comprises the steps: obtaining multi-source traffic data, an exploration rate and a current parameter vector corresponding to a current time step of an optimization hierarchy, determining a large language model or a corresponding reinforcement learning model as an action source according to the exploration rate, determining a current action according to a current state and an action source, determining an adjusted parameter vector based on the current action and the current parameter vector, determining green light time of a target phase based on the adjusted parameter vector so as to control traffic signals of an optimization level to obtain new multi-source traffic data, and determining the adjusted parameter vector as a new current parameter vector, continuing to control the traffic signal of the optimization level, storing control experience data of the optimization level in a preset first time period, and optimizing the reinforcement learning model corresponding to the optimization level; according to the method, the black box decision problem is solved by adjusting parameters with physical meanings, the action source is controlled through the exploration rate, the advantages of the two models can be combined, and the disadvantages are reduced.
Owner:CHONGQING ZHILU YUNXING TECHNOLOGY CO LTD

A category-based 6g network multi-dimensional resource ai model dynamic deployment optimization method

The application discloses a kind of 6G network multidimensional resource AI model dynamic deployment optimization methods based on category theory, belongs to intelligent collaborative optimization technical field;Method is: the cross-layer consistency dependency of end-to-end AI reasoning service is formalized by functor form;Establish the joint optimization model with long-term average end-to-end delay minimization as target, while being constrained by multidimensional resource and service quality;Convert long-term random optimization problem into time-slot online decision problem;Get AI model dynamic deployment and task scheduling result.The application realizes cross-layer consistency description to task scheduling and model deployment through category theory unified modeling and functor composite mechanism, reduces the inconsistency and redundant constraint caused by hierarchical modeling, improves the structured degree and explainability of joint decision;Under the constraint of multidimensional resources such as calculation, memory, storage and bandwidth, dynamic adaptive optimization is realized, node resource over-limit and load imbalance are effectively avoided, and congestion and queuing delay are reduced.
Owner:NANJING UNIV OF POSTS & TELECOMM

Table tennis playing robot preset performance control method based on optimal human decision

The invention relates to the technical field of robot intelligent control, and discloses a table tennis playing robot preset performance control method based on an optimal human decision, which comprises the following steps: reconstructing a kinetic model of a table tennis playing robot system, and obtaining servo constraints after state transformation; constructing a hierarchical hybrid controller of a robot preemption algorithm and a human decision control algorithm, integrating a human fuzzy intention with mechanical control to optimize a ball hitting strategy, and converting a decision problem into a functional optimization problem; and obtaining an analytical expression of the optimal membership function to obtain weights or priorities of human decisions in different states. According to the method, machine motion guided by a clear physical law is combined with fuzzy representation of human behaviors and intentions to form a coherent framework, optimization of system performance is achieved by integrating a preemption algorithm and a human decision algorithm, and meanwhile safety and robustness are ensured.
Owner:ANHUI UNIV

A serverless computing based adaptive video streaming method and system

ActiveCN116962414BVideo deliveryReinforcement learning algorithm
The application relates to a kind of adaptive video streaming method and system based on serverless computing, belong to streaming media transmission technical field.System is realized by fine-grained serverless pipeline Video delivery, use stateless function to strengthen the response to video request event, use a kind of near-end strategy optimization PPO of three-end clipping based on deep reinforcement learning algorithm to solve the bit rate adaptive sequence decision problem in video playing process.In addition, dynamic video block quality factor is included in user experience QoE index, to configure QoE model, for each video block is assigned a priority weight, improves the robustness of video bit rate decision, thereby reduces video streaming delay, improves user viewing experience.
Owner:BEIJING INST OF TECH

Heterogeneous cluster model update task scheduling method based on reinforcement learning

The invention provides a heterogeneous cluster model update task scheduling method based on reinforcement learning, which mainly improves the scheduling efficiency of resources in an unmanned cluster, balances the task processing performance of each resource in the unmanned cluster, and solves the scheduling decision problem of model update tasks in a heterogeneous unmanned cluster. According to the scheme, before an unmanned cluster starts distributed training, scheduling decision making is firstly carried out on a training task of a single model, then task scheduling is executed, and finally distributed training is carried out; the system takes model training completion time and cluster total energy consumption as optimization targets; according to the method, the advantages of Dirichlet distribution and reinforcement learning are combined, task scheduling constraint conditions can be effectively met, meanwhile, the method has high exploration capacity, and therefore an efficient task scheduling strategy is generated; compared with a traditional heuristic algorithm and other reinforcement learning methods, the task scheduling scheme can be directly generated for the cluster without complex mathematical modeling, and the complexity and cost of implementation and maintenance are reduced.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Mass spatio-temporal data retrieval acceleration method and device based on imitation learning

The invention provides a massive spatio-temporal data retrieval acceleration method and device based on imitation learning, and the method comprises the steps: modeling a cache decision problem in retrieval acceleration into a Markov decision process, and constructing a CorrelaCache model which comprises two parallel sub-networks: on one hand, carrying out the time sequence coding of a short-term access sequence through a long and short-term memory network, and on the other hand, carrying out the time sequence coding of the short-term access sequence through a short-term memory network; capturing a local access mode; and on the other hand, an autocorrelation module is constructed based on fast Fourier transform, periodic dependency features in the access sequence are identified, and periodic phases are aligned. And the joint prediction layer fuses the features from the LSTM and the self-correlation module, and outputs the expelling probability distribution of each cache data unit through full connection layer mapping. In the training process, a DAgger algorithm is adopted to iteratively update the strategy so as to avoid distribution deviation between strategy training and actual operation; and meanwhile, a composite loss function containing sorting loss and reuse distance prediction loss is designed. The retrieval response efficiency and the resource utilization rate of the method are obviously superior to those of an existing method.
Owner:WUHAN UNIV

Multi-attribute scheme decision-making method and system based on fuzzy trusted cubic number

The invention provides a multi-attribute scheme decision-making method and system based on a fuzzy trusted cubic number, and relates to the technical field of data processing, and the method comprises the steps: obtaining a plurality of alternative multi-attribute schemes related to a multi-attribute scheme decision-making problem; constructing a fuzzy credible cubic number of each evaluation attribute; constructing a decision matrix of a multi-attribute scheme decision problem by taking the obtained fuzzy credible cubic number as an element; determining a decision variable of each alternative multi-attribute scheme based on the decision matrix; determining an ideal evaluation attribute value of each evaluation attribute from the decision matrix; calculating the similarity between the decision variable and the ideal decision variable; and selecting the decision variable corresponding to the maximum similarity, and outputting the alternative multi-attribute scheme corresponding to the decision variable as the optimal multi-attribute scheme. By distinguishing the credibility, the interference of unreliable attributes on the result can be reduced when the attribute data has uncertainty or fluctuation, and the decision stability and the decision accuracy under the condition that the attribute data is uncertain are improved.
Owner:SHAOXING UNIVERSITY

Substation flood prevention response decision-making method and device, electronic equipment and medium

The invention relates to a transformer substation flood prevention response decision-making method and device, electronic equipment and a medium, and belongs to the technical field of transformer substation intelligent decision-making, and the method comprises the steps: constructing a flood prevention decision-making problem, a standard answer of the flood prevention decision-making problem, and a flood prevention rule base according to transformer substation flood prevention data; inputting the flood prevention decision question into a preset large language model to obtain a flood prevention decision answer output by the large language model; determining a first loss value according to the difference between the flood prevention decision answer and a standard answer, determining a danger level of a flood prevention rule violated by the flood prevention decision answer according to a flood prevention rule base, and determining a second loss value according to the danger level; and performing fine adjustment on the large language model according to the first loss value and the second loss value, and performing transformer substation flood prevention response decision through the fine-adjusted large language model. Through the fine-tuned large language model, a more reasonable transformer substation flood prevention response decision can be generated based on flood prevention rule constraints.
Owner:WUHAN UNIV OF TECH

A cold and hot data recognition method and system based on streaming learning

ActiveCN121614936BData streamEngineering
The application discloses a cold and hot data recognition method and system based on streaming learning, and belongs to the field of computer storage. The system regards cold and hot recognition as a decision problem, extracts multi-dimensional features including data flow, control flow and system information through a feature extraction module to construct a feature vector, realizes online hotness evaluation and real cold and hot label generation through an online label module, adopts a streaming learning algorithm in a cold and hot recognition module, predicts the cold and hot state of a data block in the future in real time according to the feature vector, and regularly updates a model to cope with concept drift. In addition, the system adopts a dynamic adjustment mechanism of cold and hot perception threshold, can adaptively guide data migration, realizes online judgment and labeling of data cold and hot, and provides efficient and adaptive cold and hot data recognition services for user applications.
Owner:HUAZHONG UNIV OF SCI & TECH

Heterogeneous Internet of Vehicles resource scheduling method based on graph embedding hierarchical reinforcement learning

The invention belongs to the technical field of Internet of Vehicles (IoV) and distributed artificial intelligence, and discloses a heterogeneous Internet of Vehicles resource scheduling method based on graph embedding hierarchical reinforcement learning. The thought of learning value and scheduling overhead decomposition is migrated to a deep reinforcement learning algorithm, and a hierarchical quality-driven deep reinforcement learning (HQ-DRL) framework is provided; according to the invention, a hierarchical cooperative Actor-Critic (GE-HSAC) algorithm based on graph embedding is designed based on a framework, and the problem of short-view decision making caused by the fact that learning quality and communication cost cannot be considered at the same time in a traditional scheduling strategy is effectively solved; in addition, the defect that the convergence rate of the global model is slowed down due to the fact that communication resources are wasted for transmitting the low-value model is avoided; through collaborative optimization of the high-layer intelligent agent and the low-layer intelligent agent, the convergence speed of a distributed learning task can be remarkably accelerated in a dynamic and heterogeneous car networking environment finally, the utilization rate of communication resources is improved, and the performance advantage higher than that of a traditional single-target optimization method is shown.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Rag and preference alignment collaborative optimization method and system for power field

ActiveCN121980039BData setLinguistic model
The application relates to the technical field of artificial intelligence, and discloses a RAG and preference alignment collaborative optimization method and system for the electric power field. The method comprises the following steps: performing text segmentation on an electric power business decision knowledge base, and generating electric power business adaptation texts by vectorizing and representing each text block obtained through retrieval enhancement; based on the electric power business adaptation texts, combining a first large language model and retrieval enhancement to generate a plurality of electric power business decision problems and expert decision texts and ordinary decision texts corresponding to each electric power business decision problem; and based on a preference data set, fine-tuning a pre-trained second large language model by using a direct preference optimization and a probability ratio preference optimization weighted fusion mode to obtain a target large language model. The application provides a reliable intelligent modeling tool for complex decision-making of an electric power system, and provides a reusable optimization paradigm for landing of professional field modeling of reinforcement learning.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO LTD +1

A regional power grid source and load collaborative scheduling learning optimization method based on principal-agent game

The present application relates to the technical field of power system, especially to a regional power grid source-load collaborative scheduling learning optimization method based on master-slave game, comprising the following steps: S1, analyzing the physical architecture and logical relationship of the regional power grid source-load collaborative scheduling system, and proposing optimization strategy merit evaluation index; S2, establishing the master-slave game model between the market price setting institution and the load aggregator, the unit generation plan model and the constraint model under the condition of uncertain load response; S3, describing the scheduling decision problem as a learning optimization mechanism for random sequential decision of electricity price, and solving by using a typical reinforcement learning method.
Owner:HEFEI UNIV OF TECH +1

Queue awareness task unloading method and device, equipment, storage medium and product

The invention relates to the technical field of mobile edge computing, and discloses a queue-aware task unloading method and device, equipment, a storage medium and a product, and the method comprises the steps: collecting a current system state of a hierarchical air-based mobile edge computing system at the beginning of each time slot; the current system state is input into a preset decision model, a task unloading decision is output, the preset decision model splits a complex decision problem into deterministic sub-problems which can be solved by time slots, and an unloading decision of each time slot is formulated; and unloading the sensing task in each mobile terminal to the high-altitude platform and the low-altitude equipment based on the task unloading decision, and executing the sensing task according to an unloading result. A dynamic optimization problem of an unloading decision is decomposed into a time slot-by-slot sub-problem through a preset decision model, and under the condition that task random generation and arrival, multi-node movement and channel time varying exist at the same time, the system cost can be minimized, the network stability can be kept, and the efficient operation of the whole layered space-based mobile edge computing system can be ensured.
Owner:PENG CHENG LAB

Aviation heterogeneous network airborne dynamic switching method based on Q learning

The invention discloses an aviation heterogeneous network airborne dynamic switching method based on Q learning, and aims to solve the switching decision problem caused by high-speed movement and network dynamic change of a civil aircraft in a heterogeneous network composed of VHF, LDACS, GEO satellites and LEO satellites. The method comprises the following steps: firstly, predicting an aircraft trajectory by adopting a gating circulation unit neural network, and screening candidate links in combination with network node data; then, communication quality indexes of candidate links are obtained through real-time monitoring, and predictive indexes are calculated according to aircraft trajectories; constructing a comprehensive utility function by applying an analytic hierarchy process, and performing quantitative evaluation on each index; and finally, on the basis of a Q learning decision model, designing two intelligent switching schemes, namely a centralized intelligent switching scheme and a distributed intelligent switching scheme, and dynamically selecting an optimal communication link. According to the method, through active prediction and intelligent decision making, the communication interruption risk is remarkably reduced, and seamless, reliable and efficient switching of the communication links of the aircraft in the whole air route range is ensured.
Owner:BEIHANG UNIV

Vehicle-mounted augmented reality task unloading decision and resource allocation method based on multi-agent reinforcement learning

The invention discloses a mobile edge computing task unloading method in an Internet of Vehicles environment. According to the method, resource allocation optimization is carried out through a multi-agent depth deterministic strategy gradient task unloading algorithm. According to the method, aiming at a resource allocation problem in a task unloading process, the problem is split into two key sub-problems, namely a task unloading decision and resource allocation, by combining an augmented reality service model. The task unloading decision problem selects an optimal unloading strategy by using a deep reinforcement learning algorithm; according to the resource allocation problem, an optimal resource allocation scheme is determined through multiple traditional optimization methods. Simulation experiment results show that the algorithm is superior to an existing distributed deterministic strategy gradient algorithm in the balance of energy consumption and time delay, and lower energy consumption can be achieved while the service quality is guaranteed.
Owner:HOHAI UNIV

Electric energy quality regulation and control method, device, equipment, medium and program product

The invention discloses an electric energy quality regulation and control method, device and equipment, a medium and a program product. The method comprises the following steps: obtaining power distribution network observation data of a regulation and control node; inputting the observation data of the power distribution network into the multi-agent deep reinforcement learning model to obtain a power quality regulation action; wherein the multi-agent deep reinforcement learning model makes a decision based on a Markov decision process; the multi-agent deep reinforcement learning model is obtained through training by the regulation and control center according to empirical values determined by the agents based on local target sample data; and executing an electric energy quality regulation and control action on the power distribution network where the regulation and control node is located. A power quality regulation and control problem is converted into a Markov decision problem, a centralized training-decentralized execution framework is adopted, a control strategy is continuously optimized in continuous interaction between local observation data of each agent and a power distribution network, and accurate value estimation is provided according to the global evaluation capability of a regulation and control center. And multi-agent cooperative operation is realized while local observation is ensured.
Owner:BEIJING SMARTCHIP MICROELECTRONICS TECHNOLOGY CO LTD

Space target identification simulation method based on Bernoulli probability model

The invention discloses a space target identification simulation method based on a Bernoulli probability model, and the method comprises the steps: taking a temporary target as a random variable obeying Bernoulli distribution, determining a probability value of the random variable, calculating a distribution value through a function meeting Bernoulli distribution characteristics, and enabling the obtained distribution value to be a decision result of space target identification; the values of the decision result are 0 and 1, 0 represents that a new space target is detected, and 1 represents that the target is an existing in-orbit target subjected to maneuvering orbital transfer. According to the method, the simulation decision problem of difficulty in temporary target identification in a space target deduction simulation system is effectively solved, space target identification is performed by using probabilistic decision, the algorithm logic is simple, the calculation amount is small, and the engineering practicability is high.
Owner:XIAN TIANYU STAR CONTROL INFORMATION TECH CO LTD +1