Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

79 results about "Q learning algorithm" patented technology

Q-learning, is a simple incremental algorithm developed from the theory of dynamic programming [Ross,1983] for delayed reinforcement learning. In Q-learning, policies and the value function are represented by a two-dimensional lookup table indexed by state-action pairs. Formally, for each state and action let:

AI-based automatic production line scheduling system in industrial internet

The invention discloses an AI-based automatic production line scheduling system in an industrial internet, which relates to the technical field of production scheduling and comprises a production plan management module S1, a dynamic scheduling engine module S2, a resource scheduling module S3, a real-time monitoring system module S4, an exception handling center module S5 and a data optimization platform module S6. In the industrial internet, an AI-based automatic production line scheduling system, an X dynamic scheduling engine millisecond response and a multi-agent reinforcement learning engine based on a federated learning architecture realize millisecond response scheduling, each device is used as an autonomous decision-making unit, and dynamic coordination is performed through a distributed Q learning algorithm, so that the vacancy rate of the devices is greatly reduced, and the scheduling efficiency is improved. According to a long-short-term memory network deep analysis model of order delivery cycle compression, emergency order insertion response speed improvement, multi-modal AI quality monitoring, fusion of vibration, thermal imaging and current spectrum, the detection rate is greatly improved compared with a unified sensor, causal reasoning and root cause analysis are performed, a fault causal graph is constructed to position a deep problem, and the average repair time is shortened.
Owner:JIANGSU AOYILAN INTELLIGENT TECH CO LTD

Low-delay video stream code rate adaptation method and system based on offline element reinforcement learning

The invention provides a low-delay video stream code rate adaptation method and system based on off-line element reinforcement learning. The method comprises the following steps: constructing an off-line expert track data set; constructing a code rate adaptation strategy network, and outputting a downloading code rate selection result and a network throughput correction value by taking a video stream transmission state as input; and adopting a meta-implicit Q learning algorithm to perform double-layer alternate optimization of an inner layer and an outer layer on the policy network, the inner layer updating policy parameters of the policy network based on the offline expert trajectory data set, and the outer layer further adjusting the policy network parameters through meta reinforcement learning to improve the adaptability of the policy network to unseen tasks. According to the method, multiple expert algorithm experiences are integrated through offline reinforcement learning, high generalization and adaptive ability can be realized in multiple target time delays and complex network environments without online trial and error exploration, and finally, high-quality and smooth playing of video streams in a low-time-delay live broadcast scene is supported.
Owner:SHANGHAI JIAOTONG UNIV

Multi-agent system double-layer Q learning control method and system based on Warisstein distance

The invention discloses a multi-agent system double-layer Q learning control method and system based on Warisstein distance. The method comprises the following steps: establishing a multi-agent system state space model under fault and unknown distribution interference; constructing an inner-layer fault-tolerant control system and an outer-layer robust control system; designing an internal fault-tolerant control gain based on a Q learning algorithm; designing an outer layer robust control gain and a bias item based on the Warisstein distance and a zero-sum game framework; and designing a distributed consistency protocol in combination with neighborhood state information. According to the method, under the conditions that a system model is unknown, an executor has additive time-varying faults, external disturbance probability distribution is uncertain and only depends on limited samples, the influence of compensation faults and disturbance can be effectively reduced, the H infinity performance constraint is met, meanwhile, asymptotic state synchronization of all agents is achieved, and system stability is ensured.
Owner:NANJING UNIV OF SCI & TECH

Data center global temperature control optimization method and system based on hybrid reinforcement learning

The invention discloses a data center global temperature control optimization method based on online-offline hybrid reinforcement learning, and the method comprises the steps: constructing an interpretable global state space fusing physical information, and guaranteeing that a state variable has a clear physical meaning; a combined action space is defined, and cooperative control of the cold source side and the tail end side is achieved; designing a multi-objective reward function based on physical model driving, and comprehensively considering energy consumption, temperature stability, carbon emission and physical optimization indexes; obtaining a basic security policy through offline pre-training, and extracting reliable behaviors from historical data by using a conservative Q learning algorithm; strategy optimization is achieved through online safety fine adjustment, and gradual adjustment is conducted under multiple constraints to adapt to real-time changes; and finally, deploying an optimization strategy to realize closed-loop control of the system, and establishing a continuous learning mechanism to cope with long-term drift. According to the method, a complete offline-online mixed learning system is established, and the energy efficiency of the data center temperature control system is improved.
Owner:SOUTHEAST UNIV

Engine rotating speed compensation method for hybrid power excavator and hybrid power excavator

According to the method, a fuzzy subset is set according to the required torque, the current SOC value and the difference value between the engine rotating speed and the motor rotating speed, and the torque of an ISG motor is dynamically adjusted through a fuzzy control algorithm. The fuzzy control adopts a fixed fuzzy rule, so that a Q table is established through a Q learning algorithm, the compensation torque of the ISG motor is obtained, the final output torque of the ISG motor is determined to be the sum of the output torque of the fuzzy controller and the compensation torque output by the Q learning algorithm, and the torque output by the ISG motor is determined. According to the method, the advantage of high response speed of the torque of the motor is utilized, the difference value between the required torque and the actual torque is effectively compensated, the rotating speed fluctuation of the engine caused by overlarge change of the torque of the engine is avoided, and the fuel economy and the driving operation stability are improved.
Owner:JIANGSU XCMG CONSTRUCTION MACHINERY RESEARCH INSTITUTE LTD

High-order robust control method for mobile robot in combination with weighted double-Q learning algorithm

The invention provides a high-order robust control method for a mobile robot in combination with a weighted double-Q learning algorithm. The high-order robust control method comprises the following steps: step 1, constructing an intelligent trolley motion scene including time-varying friction, variable gradient, random external disturbance and sensor noise complex uncertainty; step 2, designing a weighted double-Q learning adaptive nonlinear expansion state observer; step 3, based on the step 2, constructing a reward function including a trajectory tracking error, a disturbance estimation error and a stability constraint; 4, designing an improved super-spiral sliding mode control law based on the weighted double-Q learning adaptive nonlinear extended state observer; and 5, setting a plurality of groups of working condition simulation comparison schemes, and carrying out simulation experiment verification. According to the invention, real-time adaptive control of the mobile robot on different working conditions is realized, the trajectory tracking precision of the system is improved, and sliding-mode control buffeting is inhibited at the same time.
Owner:NANTONG UNIV

Multimodal sequence recommendation method and system based on hyperbolic curvature adaptive learning

The invention discloses a multi-modal sequence recommendation method and system based on hyperbolic curvature adaptive learning, and the method comprises the steps: exploring the adaptive hyperbolic curvature of a time sequence behavior diagram in different modals based on a reinforcement learning frame through a curvature adaptive learning agent; modeling the multi-modal interest of the user in a hyperbolic space with a specific curvature through a multi-modal interest learning agent; a Nash Q learning algorithm is adopted to cooperatively update a curvature adaptive learning agent and a multi-mode interest learning agent, and when the two agents reach Nash equilibrium, adaptive curvatures in different modes are explored; and learning a user behavior interaction mode under the image mode and the text mode in the hyperbolic space, predicting the score of the user for the next multi-mode item, and carrying out multi-mode item recommendation based on the predicted score. According to the method, the hyperbolic space is used as the embedding space, and the user preference information is mined by using the text and image modal information of the project, so that the overall recommendation performance is improved.
Owner:ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY

Multi-objective optimization method and system for heating and ventilation system based on reinforcement learning

The invention discloses a multi-objective optimization method and system for a heating and ventilation system based on reinforcement learning, and relates to the technical field of building energy system optimizing.The method comprises the steps that a state space is constructed based on collected environment parameters, equipment state parameters and historical operation data of the heating and ventilation system; the state space is composed of state vectors generated after normalization and discretization; constructing an action space based on controllable equipment of the heating and ventilation system; the action space is generated by discretizing the operation mode and control parameters of the controllable equipment; constructing a multi-target reward function to evaluate a control effect based on the state vector and the action space; taking the state vector as input, and performing iterative training by using a Double Q learning algorithm until a trained control model is obtained; and inputting the currently acquired state vector into the trained control model, and outputting an optimal action to generate a control instruction for the heating and ventilation system.
Owner:GUANGZHOU SHANGHANG INFORMATION TECH CO LTD

Self-learning job shop scheduling method meeting waiting time constraint

The invention belongs to the technical field of job shop scheduling, and particularly relates to a self-learning job shop scheduling method meeting waiting time constraint, which comprises the following steps: S0, constructing a job shop scheduling problem model with waiting time constraint; the method comprises the following steps: S1, acquiring job shop scheduling problem data, a configuration algorithm and operation parameters; s2, constructing chromosome individuals, generating chromosomes and initializing a population; s3, calculating the fitness value of each chromosome individual; s4, forming a new generation of population; s5, combining the fitness information of the current population, dynamically selecting a crossover rate Pc through a Q learning algorithm, and performing crossover operation on the population; s6, dynamically determining a mutation rate Pm in the same parameter combination space by using a Q learning algorithm, and performing mutation operation on the crossover progeny to generate an updated population; and S7, judging whether an iteration termination condition is met or not. According to the method, a feasible and near-optimal scheduling scheme can be efficiently generated on the premise of ensuring that the inter-process waiting time constraint is met.
Owner:CHONGQING CITY MANAGEMENT COLLEGE

A micro-grid scheduling control method and device based on policy optimization, equipment and medium

The application provides a micro-grid scheduling control method and device based on policy optimization, equipment and medium, belonging to the power grid scheduling technical field, the method: collecting real-time operation data of new energy power generation equipment, power load equipment and energy storage system; receive user power generation and power consumption plan, combine real-time operation data and historical operation data, extract multi-resource dynamic response characteristics through LSTM model, output adjustment capacity evaluation matrix; receive power grid scheduling plan, based on the adjustment capacity evaluation matrix, match the optimal operation mode from the knowledge base, generate the scheduling instruction set; issue the scheduling instruction set to the target equipment and execute, collect the execution data in real time and calculate the operation deviation; when the operation deviation is over the threshold, the instruction weight is updated by using the online Q learning algorithm, and the optimization strategy is fed back to the knowledge base. The application realizes fine scheduling of micro-grid through real-time data collection and intelligent optimization algorithm, improves resource utilization and operation flexibility, and reduces operation cost.
Owner:INSPUR ARTIFICIAL INTELLIGENCE RES INST CO LTD SHANDONG CHINA

Hybrid multi-target coordination control system based on SVG and synchronous phase modifier

The invention relates to the technical field of power systems, and discloses a hybrid multi-target coordination control system based on SVG and a synchronous phase modifier, which significantly improves the decision quality and foresight of an intelligent agent by fusing a real-time state and future prediction, and improves the decision efficiency compared with a single system state vector. The enhanced state characterization not only comprises real-time parameters such as voltage deviation and reactive power difference, but also integrates future uncertainty probability information predicted by the random forest, so that the intelligent agent can simultaneously sense the current state and the expected evolution trend of the system, and the enhanced environment sensing capability enables the Q learning algorithm to make a more predictable decision, so that the method is more suitable for popularization and application. According to the method, the coordination strategy of the SVG and the phase modifier is adjusted in advance before voltage fluctuation occurs, state characterization is enhanced, comprehensive environment information is provided, an intelligent agent can better balance the relation between quick response and steady state support, the voltage stability is guaranteed, the economical efficiency of system operation is improved, and finally multi-target collaborative optimization in the true sense is achieved.
Owner:INNER MONGOLIA ELECTRIC POWER (GRP) CO LTD WUHAI UHV POWER SUPPLY BRANCH

A workflow scheduling method considering security risk and cost optimization in cloud environment

The application provides a workflow scheduling method considering security risks and cost optimization in a cloud environment, which reduces the overall security risks and execution costs of the workflow under the given deadline constraint of a user. Firstly, the cloud computing resources are defined, the workflow structure is abstracted, and a security risk calculation method is proposed. Then, a reinforcement learning model oriented to security risks and cost optimization is constructed. Next, an action selection algorithm in the learning process is determined. Finally, iterative learning is carried out based on the Q learning algorithm idea, and finally the task scheduling scheme is determined, and the overall security risks and total costs of the workflow are obtained.
Owner:NANJING UNIV OF POSTS & TELECOMM

False alarm suppression method for mobile satellite terminal direction finding based on multi-dimensional feature joint screening

PendingCN122640007ATime domainFeature Dimension
The present application belongs to the technical field of wireless communication signal processing, and particularly relates to a mobile satellite terminal direction finding false alarm suppression method based on multi-dimensional feature joint screening, aiming to solve the technical problems in the prior art that a single power dimension is used for signal judgment, the uplink signal and the downlink signal on the mobile satellite system cannot be distinguished, and the judgment threshold is fixed and cannot adapt to complex electromagnetic environments. A multi-dimensional signal feature library is constructed; front-end signal reception and preprocessing are performed; multi-dimensional features are extracted in parallel, and the confidence levels in the time domain, the frequency domain, the modulation domain and the feature code domain are calculated based on the multi-dimensional signal feature library; adaptive weighted fusion is performed based on information entropy, the weight calculation and fusion score of each feature dimension are automatically adjusted according to the real-time electromagnetic environment complexity; and a dynamic threshold adaptive mechanism based on the Q learning algorithm is used to realize closed-loop optimization and feedback adjustment of the judgment threshold, realize multi-dimensional joint screening of signals, and effectively suppress direction finding false alarms in combination with the accurate distinguishing mechanism of uplink and downlink signals.
Owner:CHENGDU GUDE TECH CO LTD

A power transmission network outage maintenance scheduling optimization method based on Q learning

The application provides a power grid outage maintenance scheduling optimization method based on Q learning, which comprises the following steps: generating a Gantt chart by acquiring an original power grid topology file, and establishing a state judgment rule and a target function of reinforcement learning; integrating maintenance plans of subordinate units to the Gantt chart to form an initial maintenance plan; calculating an initial power flow collection by using a built-in power flow solver of Matlab, and evaluating a power grid operation state; in the agent training process, an epsilon-greedy method is used to select a maintenance action, the Gantt chart and the topology file are updated, and the Q value is updated through a Bellman equation, so that the maintenance plan is optimized; and the Gantt chart of a final outage maintenance plan is output after the training is completed. The power grid scheduling problem is abstracted as a Q learning problem, the outage scheduling is optimized by using a Q learning algorithm through reasonable constraints and a target function, the work intensity of dispatchers is reduced, the talent training cost is reduced, the scheduling efficiency and reliability are improved, and the key technical problems of power grid scheduling are solved.
Owner:STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST +1

Power distribution network planning method and device based on Nash Q learning, equipment and medium

The invention is applied to the field of power distribution network planning, and discloses a power distribution network planning method and device based on Nash Q learning, equipment and a medium. The method comprises the following steps: acquiring a real-time state corresponding to a to-be-planned power distribution network, topological parameters including the number of nodes and the number of lines of the to-be-planned power distribution network, and resource parameters and policy parameters including a parameter set of new energy candidate nodes and energy storage candidate nodes; constructing a multi-target planning model according to the topological parameters and the policy parameters; constructing a multi-target planning model according to the multi-target planning model based on a Nash Q learning algorithm; initializing a multi-main-body Q matrix, and updating the multi-main-body Q matrix based on the real-time state, a preset exploration strategy and a multi-target planning model; and determining a planning strategy of the to-be-planned power distribution network based on the multi-main-body Q matrix. In this way, the global optimality and practicability of the planning strategy are improved by combining the three-dimensional coupling decision and the Nash Q learning with the multi-agent game.
Owner:STATE GRID SICHUAN ECONOMIC RES INST

Adaptive model predictive control method based on incremental double q-learning

The application belongs to the field of intelligent control, and particularly relates to an adaptive model predictive control method based on incremental double Q learning, which comprises the following steps: 1) establishing a kinematic model of a non-complete constraint differential mobile robot, and combining the kinematic model to construct a trajectory tracking error model; 2) based on the trajectory tracking error model, constructing a prediction model of a trajectory tracking model predictive controller of the mobile robot, and realizing high-performance trajectory tracking control according to the prediction model; 3) performing incremental discretization operation on the action space and the state space of the double Q learning agent; 4) through real-time interaction between the agent and the environment and online iterative learning, the state and the action space of the double Q learning agent are continuously optimized, and an online adaptive MPC controller is obtained to realize high-precision autonomous trajectory tracking control of the mobile robot. The application realizes online optimization of a weight matrix of the MPC by using an incremental double Q learning algorithm, and adjusts the MPC weight matrix in real time according to the system state to optimize the total reward.
Owner:SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI

Production line intelligent clamp control method and system

The invention relates to the field of production line intelligent clamp control, and discloses a production line intelligent clamp control method and system, and the method comprises the steps: collecting the 3D point cloud data of a workpiece, calculating the geometric features and shape errors of the workpiece through a point cloud processing algorithm, predicting and optimizing the clamping force, and optimizing the clamp task switching through combining PID control and a Q learning algorithm; compared with the prior art that the traditional fixture control method mostly depends on manual adjustment of the position and the clamping force of the fixture, the requirements of high precision and rapid adaptation are difficult to meet when the requirements of complex workpiece shapes and diversified operation are met, and particularly, the requirements of high precision and rapid adaptation are not met under the conditions that the shapes of workpieces are irregular and the surfaces have large errors. In order to solve the technical problem that the production efficiency is low due to the fact that a traditional method cannot efficiently adjust a clamp, clamping points are accurately selected through visual detection, workpiece form changes are automatically adapted, and the production efficiency of a production line and the workpiece clamping precision are remarkably improved.
Owner:JIANGSU YUXIN SENSOR TECH CO LTD

Multi-agent system double-layer q learning control method and system based on wasserstein distance

The application discloses a multi-agent system double-layer Q learning control method and system based on a Wasserstein distance, which comprises the following steps: establishing a multi-agent system state space model under faults and unknown distribution disturbances; constructing an inner fault-tolerant control system and an outer robust control system; designing internal fault-tolerant control gains based on a Q learning algorithm; designing outer robust control gains and bias terms based on a Wasserstein distance and a zero-sum game framework; and designing a distributed consistency protocol in combination with neighborhood state information. The method can effectively reduce the influence of compensation faults and disturbances under the condition that a system model is unknown, an actuator has additive time-varying faults, external disturbance probability distribution is uncertain, and only limited samples are available, and can realize asymptotic state synchronization of all agents while meeting H-infinity performance constraints, thereby ensuring system stability.
Owner:NANJING UNIV OF SCI & TECH

Omnidirectional leveling device for hilly and mountain tractor seat and control method

The invention belongs to the technical field of tractor seat adjustment, and particularly relates to an omni-directional leveling device for a hilly and mountainous region tractor seat and a control method, and the omni-directional leveling device comprises a posture sensing unit, a central controller and a hydraulic execution unit. The posture sensing unit is used for detecting transverse and longitudinal inclination angles of the seat in real time; the hydraulic execution unit comprises an oil source system, four symmetrically-arranged leveling hydraulic cylinders and corresponding electro-hydraulic proportional direction valves. The core of the control method is that a QBP-PID composite intelligent control strategy fusing a Q learning algorithm, a BP neural network and PID control is adopted. The ECU receives an actual inclination angle signal, optimizes PID control parameters in real time by using a QBP-PID composite intelligent control strategy with a horizontal attitude as a target, and drives the electro-hydraulic proportional directional valve to accurately control the cooperative telescopic movement of each hydraulic cylinder, so that the rapid, stable and self-adaptive adjustment of the transverse and longitudinal inclination angles of the seat is realized, and the safety of the seat is improved. The problem that the sitting posture of a driver is uncomfortable due to topographic relief of a hilly and mountain tractor is solved, and the operation comfort and safety are improved.
Owner:HENAN UNIV OF SCI & TECH

A method and system for frequency control of a thermal power plant based on reinforcement learning algorithm

The application relates to a frequency control method for a thermal power plant based on a reinforcement learning algorithm, wherein the method comprises the following steps: obtaining primary frequency regulation evaluation data and current combined cycle unit working condition data to form state data; determining a state-action pair in a preset algorithm model, wherein the preset algorithm model is obtained by training a Q learning algorithm through reinforcement learning based on the state data of the thermal power plant; obtaining evaluation values of each state-action pair based on a preset estimation rule through the preset algorithm model, and obtaining a target state-action pair according to the evaluation values; generating a control instruction according to an environmental interaction action in the target state-action pair, and adjusting the output power and action time of an electric energy storage device through the control instruction. Through the application, the problem that the service life of an electric energy storage device of a combined cycle unit is low due to long-term high-load working state in the related art is solved, accurate frequency regulation control is realized, power waste of the electric energy storage device is reduced, and the service life of the electric energy storage device is improved.
Owner:HUADIAN ELECTRIC POWER SCI INST CO LTD

Security control method for networked switching system based on encryption and decryption switching q-learning replay attack detection

The application discloses a kind of networked switching system security control method based on encryption and decryption switching Q learning replay attack detection, comprising: constructing switching system model and controller model, obtain state signal and mode signal;Event-triggered mechanism based on replay attack detection is constructed, including error detection condition, mode matching condition and attack detection condition;According to Q learning algorithm, replay attack detection mechanism based on switching Q learning is constructed, event-triggered mechanism and replay attack detection mechanism are used to process the mode signal and state signal of triggered time subsystem, state estimation value, mode estimation value and state estimation residual norm are obtained, the detection result of attack is determined;State estimation value and mode estimation value are substituted into controller model, and safety controller is obtained, the equation of closed-loop system is obtained by simultaneously solving the equation of safety controller, switching system model, state estimator and mode estimator;Switching rule based on replay attack detection is constructed, according to replay attack detection result, the switching behavior of closed-loop system under replay attack is analyzed, and the judgment condition of system safety and stability is obtained, the present application can accurately analyze whether replay attack is carried out, and give the safety control discrimination condition for ensuring that switching system under replay attack, realize the safety control of switching system.
Owner:DALIAN MARITIME UNIVERSITY

Remote area remote education heterogeneous network vertical switching method

The invention relates to the technical field of remote education, and discloses a remote region remote education heterogeneous network vertical switching method, system and device and a medium. The method comprises the following steps: acquiring a current teaching service type and a network type of a user terminal; according to a membership function of the fuzzy logic system, through the current teaching service type and a plurality of different network performance indexes of the current teaching service type under each network type, determining the membership of each network performance index of the current teaching service type under each network type; performing weighted summation on the membership degree of each network performance index under each network type to obtain a service quality score of the current teaching service type under each network type; based on a Q learning algorithm, according to the service quality score of the current teaching service type under each network type, a switching strategy of the current network type is determined, and a reward function of the switching strategy is constructed according to the change of the service quality score after the action is executed and the cost of executing the action.
Owner:INNER MONGOLIA UNIVERSITY

LTE-V dual-mode communication optimization and resource allocation system

The invention relates to an LTE-V dual-mode communication optimization and resource allocation system, which belongs to the technical field of communication, and comprises a mode selection module for selecting communication modes according to service types, vehicle distances and channel quality, the communication modes comprising a PC5 mode, a Uu mode and a dual-mode communication mode; the dual-mode communication mode is a combined communication mode of a PC5 mode and a Uu mode, and the resource allocation module is used for allocating communication resources of the PC5 mode or the Uu mode and improving the frequency efficiency; the interference coordination module is used for reducing dual-mode interference by controlling transmission power and resource isolation when the mode selection module selects the dual-mode communication mode; and the reinforcement learning optimization module is used for guiding and optimizing control parameters of the interference coordination module, a decision threshold value of the mode selection module and a resource allocation proportion of the resource allocation module based on a Q learning algorithm. The optimal communication mode is dynamically selected by combining the service type, the distance and the channel state, the reliability is improved, and resource waste is avoided.
Owner:CHENGDU TONGGUANG NETLINK TECH CO LTD

Clustering network inter-cluster time slot dynamic allocation method and device based on multi-constraint evaluation

The invention discloses a multi-constraint evaluation-based inter-cluster time slot dynamic allocation method and device for a clustering network. The method is applied to each cluster head node of a plurality of clustering networks, and comprises the following steps: acquiring state information of local and intra-cluster nodes and the total time slot number of the whole network; in response to detecting that the state information of any in-cluster node is missing, calculating an estimated value of missing state information by utilizing a Q learning algorithm to obtain updated state information; according to the updated state information and the total time slot number of the whole network, calculating the time slot preemption request number of the cluster where the cluster is located; according to the number of time slot preemption requests of the cluster and other clusters, determining the number of allocated time slots corresponding to each cluster through a distributed negotiation mechanism; and allocating time slot resources to the nodes in the cluster according to the allocated time slot number. According to the embodiment of the invention, the method can adapt to the distributed characteristics of a self-organizing network, and improves the time slot resource utilization rate and network service fairness.
Owner:THE 20TH RESEARCH INSTITUTE OF CHINA ELECTRONICS TECHNOLOGY GROUP CORP

Self-adaptive AI interview method and device, electronic equipment and storage medium

The invention provides a self-adaptive AI interview method and device, electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence, in the method, the next operation (namely rule instructions, or strategy optimization behaviors, specifically next interview questions, or interview termination) can be dynamically adjusted according to interview answers, the dynamic adaptability is good, the method is more flexible, and in addition, the method can be applied to the field of AI interview. In the comprehensive scoring process, a mixed scoring strategy is adopted, the method is more scientific, when the current state does not meet the trigger condition in the preset rule, the offline reinforcement learning agent is adopted to make an intelligent decision on the current state, then the strategy optimization behavior is obtained, the accuracy is better, in addition, the preset rule is pre-configured, the controllability is good, and the user experience is improved. And the offline reinforcement learning agent is a conservative Q learning algorithm, and during training, Q value overestimation of behaviors outside distribution of a training data set can be inhibited through a conservative regularization item, so that non-compliant behaviors can be inhibited, namely, the behaviors are more controllable and compliant.
Owner:BEISEN CLOUD COMPUTING CO LTD

Question answering method and device for beacon information, question answering equipment and storage medium

The invention relates to a question-answering method and device for beacon information, question-answering equipment and a storage medium, and belongs to the technical field of data processing.The question-answering method for the beacon information comprises the steps that a question text is obtained, and the question text is converted into a DQL query statement based on a semantic understanding model, the semantic understanding model is obtained by training a Transform model based on a sample question text carrying a DQL query statement tag; based on the DQL query statement, a candidate answer set is determined in a beacon knowledge graph, and the beacon knowledge graph is constructed based on a historical beacon text and a real-time beacon text; and determining a Q value of each candidate answer in the candidate answer set based on a Q learning algorithm, and determining the candidate answer with the highest Q value as the answer corresponding to the question text. The information query accuracy in the bidding and tendering field can be effectively improved.
Owner:SHIBO TECH (WUHAN) CO LTD

A dynamic traffic guidance and signal control coordination method under traffic accidents

The application belongs to the technical field of intelligent transportation, and relates to a dynamic traffic guidance and signal control coordination method under a traffic accident, comprising the following steps: constructing a directed weighted complex road network model, obtaining a real-time operation state of a road network and evaluating the real-time operation state; based on the complex road network model and the operation state, generating a path guidance module by using a double Q learning algorithm; based on a deep deterministic policy gradient algorithm, constructing a signal control agent and generating a signal control module; establishing a bidirectional coordination mechanism between the path guidance module and the signal control module, and executing coordination iteration in a rolling time window mode to form a closed-loop coordination optimization system of path guidance and signal control. The application solves the technical problem that existing traffic guidance and signal control are often managed and optimized as independent systems, lack effective coordination mechanisms and fail to form an integrated solution.
Owner:CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY

Route planning method and device for various transportation tools for transporting post-disaster emergency materials

The invention relates to the technical field of Internet of Things, in particular to a path planning method and device for various transportation tools for transporting emergency materials after disasters, and the method comprises the steps: firstly, preliminarily planning a transportation path according to the distance between an emergency material distribution center and a plurality of disaster-affected nodes and the bearing capacity of each transportation tool; subsequently, an operator pair set formed by a plurality of removal operators and a plurality of insertion operators in pairs is introduced, operator pairs are adaptively selected through a Q learning algorithm, the preliminarily planned transportation path is recombined and optimized, meanwhile, performance indexes after path recombination are used as reward signals to be fed back to a Q learning model, and the model is driven to iteratively update an operator pair selection strategy; and finally, an iteration ending condition is met, and the transportation path with the performance meeting a preset condition is selected as the final transportation path of the multiple transportation tools. According to the scheme, through collaborative optimization of Q learning and operator pairs, the algorithm is prevented from falling into a local optimal solution, and the finally obtained transportation path can well schedule the transportation means.
Owner:CHENGDU UNIVERSITY OF TECHNOLOGY

Unmanned aerial vehicle warehouse returning path optimization method and optimization system

The invention relates to the technical field of unmanned aerial vehicle scheduling, in particular to an unmanned aerial vehicle return path optimization method, which comprises the following steps: step 1, setting a preset range around an airport as a return airspace, and monitoring an unmanned aerial vehicle which needs to return to the airport outside the return airspace in real time, and taking the unmanned aerial vehicles which take the shortest path as the warehouse returning path and can return to the warehouse within the preset time range as a warehouse returning vehicle group. According to the method, a target cluster is screened in a preset return airspace, and a Q learning algorithm (generating a path point cluster for cooperative collision avoidance), a particle swarm algorithm (dynamically optimizing a flight speed and a path point execution strategy) and an NURBS curve algorithm (path smoothing) are fused, so that the target cluster is obtained on the premise of meeting double constraints of the shortest path and the minimum energy consumption. The collision risk in a multi-unmanned aerial vehicle dense warehouse returning scene is remarkably reduced, meanwhile, the overall warehouse returning efficiency and the smoothness and the flyability of the flight path are optimized, and safe, efficient and energy-saving unmanned aerial vehicle group collaborative warehouse returning is achieved.
Owner:HUNAN UNIV

Electric vehicle charging and discharging intelligent optimization decision-making method and system based on learning mechanism guidance

The invention discloses an electric vehicle charging and discharging intelligent optimization decision-making method and system based on learning mechanism guidance. Relates to the technical field of electric vehicle and power distribution network collaborative optimization scheduling. The method comprises the following steps: step 1, acquiring daily load data of an electric vehicle, distributed wind power and photovoltaic output respectively; 2, simulating a classic scene through a K-means clustering method based on the daily load data in the step 1 to obtain simulation data; 3, constructing a multi-target optimization model according to a time-of-use electricity price system and electric vehicle charging and discharging constraints; and 4, inputting the simulation data into a multi-objective optimization model, solving by adopting a double-Q learning algorithm and a multi-objective particle swarm algorithm, and determining an optimal scheduling strategy based on a TOPSIS decision method. According to the method, an EV charging and discharging cooperative scheduling mechanism of multi-interest subject demands can be considered, and accurate optimization of a multi-target Pareto optimal solution under complex constraints is realized through an efficient algorithm.
Owner:SHENYANG INST OF ENG