Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

105 results about "Action selection" patented technology

Action selection is a way of characterizing the most basic problem of intelligent systems: what to do next. In artificial intelligence and computational cognitive science, "the action selection problem" is typically associated with intelligent agents and animats—artificial systems that exhibit complex behaviour in an agent environment. The term is also sometimes used in ethology or animal behavior.

Dynamic scheduling optimization method and system for DAG application based on deadline constraint

The invention provides a deadline constraint-based DAG application dynamic scheduling optimization method and system, and the method comprises the steps: converting an application deadline into an instant reward of each task scheduling through employing a DAG structured encoder, a Transform encoding network based on gating feature fusion, and a multi-action selection deep reinforcement learning task scheduling method based on a pointer network, in combination with a dynamic mask scheme, the mobility of a DAG application and the dynamic nature of edge resources are dealt with, then priority subtask selection and real-time decision of the scheduling position of the priority subtask selection are made, and the completion time and execution energy consumption of the application are reduced. In order to stabilize and accelerate DRL scheduler training, task encoder training and reinforcement learning training are decoupled, and a DAG encoder is pre-trained based on self-supervised learning.
Owner:XINJIANG UNIVERSITY

Training reinforcement learning agents to learn farsighted behaviors by predicting in latent space

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training an action selection policy neural network used to select an action to be performed by an agent interacting with an environment. In one aspect, a method includes: receiving a latent representation characterizing a current state of the environment; generating a trajectory of latent representations that starts with the received latent representation; for each latent representation in the trajectory: determining a predicted reward; and processing the state latent representation using a value neural network to generate a predicted state value; determining a corresponding target state value for each latent representation in the trajectory; determining, based on the target state values, an update to the current values of the policy neural network parameters; and determining an update to the current values of the value neural network parameters.
Owner:GOOGLE LLC

Interpretable high-speed ramp intelligent driving vehicle confluence control method

The invention discloses an interpretable high-speed ramp intelligent driving vehicle confluence control method. The method comprises the steps that a state space is constructed, wherein the state space comprises the position and speed of a vehicle, lane information and the relative position and speed of adjacent vehicles; human-like decision guidance: training a human-like behavior discriminator based on real driving data by using a random forest model, screening action selection of reinforcement learning, and constructing a sample set containing state vectors and human-like action labels; reinforcement learning control; and carrying out interpretability analysis. Through an SHAP analysis method, an explainable DQN longitudinal acceleration control model in an automatic driving vehicle ramp merging process is provided, and an explainable automatic driving vehicle ramp merging control model is constructed in combination with an merging decision-making model constructed based on a random forest. According to the explainable high-speed ramp automatic driving vehicle confluence control method based on the confluence urgency degree and the environmental historical information, traffic conflicts during highway confluence can be relieved, and therefore the road passing efficiency is effectively improved.
Owner:CHANGAN UNIV

Heavy-load freight direct-current electric locomotive on-board platform and train operation curve calculation method

The invention discloses a heavy-load freight direct-current electric locomotive on-board platform and a train operation curve calculation method. Step-by-step iterative optimization is carried out through a dynamic programming algorithm. Presetting operation steps and step lengths, and defining a selectable action set in each step length; and with the current state of the train as a starting point, evaluating each action in the action set in each iteration step, and calculating the next train state which may be transferred to. And evaluating all possible train state changes by adopting a target cost function, sorting according to the cost values from small to large, and preferentially selecting the action with the minimum cost value for subsequent calculation. And if all the action cost values of the current step exceed the limit, backtracking to the previous iteration step for reselection. And after iteration to the maximum step number, storing all train state changes, and finally obtaining an optimal train operation curve under the constraint of the target function. The method optimizes a control strategy through a dynamic programming algorithm, improves prediction precision, optimizes action selection, accelerates convergence, reduces redundancy, and improves operation stability, safety and algorithm efficiency.
Owner:CASCO SIGNAL LTD

Q-Learning network path planning method based on multi-dimensional reward and adaptive exploration

The invention belongs to the technical field of computer network, artificial intelligence and network flow optimization, and discloses a Q-Learning network path planning method based on multi-dimensional reward and adaptive exploration, which comprises the steps of 1, network environment modeling and parameter initialization, 2, multi-strategy action selection, and 3, multi-strategy action selection. Step 3, calculating a multi-dimensional reward according to a multi-dimensional reward rule in a data packet transmission process; step 4, performing adaptive exploration adjustment, monitoring the search success rate in a predetermined period in real time, and dynamically adjusting the exploration intensity epsilon according to the search success rate; step 5, performing adaptive pruning optimization to finally obtain a candidate action set; and step 6, iteratively updating and finally outputting an optimal routing path and a Q value table. According to the invention, the routing exploration efficiency can be improved, dense learning signals are provided, and the calculation overhead is reduced.
Owner:NANJING UNIV OF POSTS & TELECOMM

Network slice selection method based on deep reinforcement learning of gated loop unit

The invention provides a network slice selection method and system based on deep reinforcement learning of a gated cycle unit, and relates to the technical field of wireless communication. The method comprises the following steps: constructing a simulation vehicle-mounted network scene and modeling the simulation vehicle-mounted network scene into a Markov decision process; a dynamic attenuation action selection strategy is adopted for action selection, a time sequence state, actions, rewards and an updated state form a tetrad to be stored in a time sequence experience pool, and the experience pool is managed through a first-in first-out and stratified sampling mechanism; and training a neural network containing a GRU layer based on a sample in the time sequence experience pool, taking a mean square error of a target Q value and a predicted Q value as a loss function optimization parameter, and outputting an optimal slice selection result by utilizing the trained network. According to the invention, through dynamic attenuation strategy balance exploration and utilization, the unbalance problem of a traditional fixed probability mechanism is solved; state time sequence association is reserved by means of a time sequence experience pool, and time sequence breakage caused by isolated samples is avoided; state long-term dependence is effectively captured by using the GRU network, the Q value estimation accuracy is improved, and an efficient solution is provided for slice selection in a vehicle-mounted network environment.
Owner:YANCHENG INST OF TECH

Labyrinth robot brain-like cognitive system based on emotion and memory mechanism

The invention provides a maze robot brain-like cognitive system based on an emotion and memory mechanism, and the system comprises a perception and emotion generation module which obtains the environment state of a robot, and calculates the reflective fear value and cognitive fear value of the robot; the memory module is used for recording short-time memory, working memory and long-time memory of the robot in combination with an anti-thought memory mechanism; the behavior decision-making module is used for comparing the evaluation function based on the cognitive fear value with the energy margin to obtain a behavior decision-making result; the learning module is used for exploring and learning a target node and an energy node to obtain an updated target guide table and an updated energy guide table; and the action selection module is used for updating the target guide table and the energy guide table according to the behavior decision result, and outputting the next action. According to the invention, the survival rate of the robot in the environment learning process can be improved, and long-time learning is realized; and the robot can be prevented from repeatedly exploring in an invalid area, the learning efficiency is improved, the energy consumption is reduced, and self-survival is ensured.
Owner:NORTH CHINA UNIVERSITY OF TECHNOLOGY

Multi-target traffic dynamic allocation method based on deep learning

The invention discloses a multi-target flow dynamic allocation method based on deep learning, which comprises the following steps: collecting network topology data, link operation data and service flow request data to form a network state sequence; a multi-target constraint set is obtained and written into a target constraint description table; constructing a state vector, and dividing the state vector into multiple groups of input sub-vectors; carrying out conjoint analysis on the service flow, and constructing a candidate allocation action set; constructing an improved CMAC network model, and outputting an action value vector; constructing an anti-multi-arm bandit decision-making module, and outputting action selection distribution; and executing multi-target feasible region projection processing to generate a feasible traffic distribution decision. According to the method, the improved CMAC network model is introduced, and the anti-multi-arm bandit decision-making method is combined, so that stable and self-adaptive dynamic distribution of the network traffic under the multi-target constraint and complex network state is realized.
Owner:NANJING JINSHUO TECHNOLOGY CO LTD

VVC code rate control algorithm based on deep reinforcement learning

The invention discloses a VVC code rate control algorithm based on deep reinforcement learning, and the algorithm comprises the following steps: importing a video sequence into an encoder, and enabling the video sequence to enter initial frame coding; after the encoder completes the default encoding of the first two frames, the subsequent frame prediction firstly extracts the encoding state information of the previous prediction frame, and a greedy strategy is adopted to perform action selection; an overall reward value is obtained through CTU-level code rate control and an actual coding process in sequence; observing the next state and the last reward, performing TD iteration on the Q value, and adding the Q value into the Q value network after TD iteration; setting an experience playback pool; after the capacity of the experience playback pool reaches a threshold value, randomly sampling from the experience playback pool in batches; resetting the Q value network; and carrying out coding test based on the Q value network obtained by training. The VVC code rate control algorithm based on deep reinforcement learning has a good capability of guiding code rate control coding, and compared with standard code rate control in VTM13.0, the VVC code rate control algorithm based on deep reinforcement learning can bring high rate distortion performance and improve code control precision.
Owner:HAINAN NORMAL UNIV

Intelligent routing traffic management method and system based on reinforcement learning

The invention discloses an intelligent routing flow management method and system based on reinforcement learning, and the method comprises the following steps: building a lightweight motion selection model at each network node, and mapping a communication link which can be used by the node into a plurality of forwarding options; the current network state is formed by collecting the delay, packet loss rate and utilization condition of each link in real time; and the system distributes priorities of different forwarding options by utilizing a probability strategy, and performs dynamic adjustment in combination with the change of the network state. And after each round of forwarding, the system updates and optimizes the strategy according to the communication effect, so that a routing decision mechanism which does not depend on central control and can be adaptively adjusted is realized. The method is suitable for a complex and frequently-changing network environment, and the overall transmission efficiency and the system stability are improved.
Owner:FOCALCREST LTD

A robot dance automatic generation method and system based on music feature analysis

PendingCN122156404AMetadata video data retrievalBiological modelsEngineeringMulti objective model
The application provides a robot dance automatic generation method and system based on music feature analysis, and relates to the technical field of robot control. First, a music feature spectrum covering rhythm, energy and other dimensions is constructed to convert abstract music into structured information understandable by machines, and then a dance semantic label sequence bound to a time axis is generated to provide accurate basis for action selection. By combining with a forward-looking window to predict future music features, the current and future candidate actions are filtered through a multi-objective model to realize the dual matching of action and music micro-rhythm and macro-structure. The finally generated control instruction sequence directly drives the physical robot, which breaks the dependence on video materials, realizes the automatic generation of dance with creativity and artistic expression, and improves the generation efficiency and adaptive flexibility.
Owner:ZHEJIANG SILICON ARK ROBOT CO LTD

Action proposal system, and action proposal method

To suggest an action executable during a user's free time by extracting the action according to difference in transit time due to a transit method, to a user.SOLUTION: An action proposal system 1 includes: a free time recognition unit 12 which recognizes free time of a user U; a free time starting site recognition unit 13 which recognizes a free time starting point being a site where the user is present at the starting time point of the free time; a candidate action selection unit 15 which selects, as a candidate action to be suggested to the user U, an action executable by the user U while using the free time, and containing a movement between the free time starting site and a destination site by any of a plurality of movement methods and a prescribed activity at the destination site; and a candidate action information providing unit 16 which allows a display device 50 used by the user U to display an action proposal window to display required time period for moving between the free time starting site and the destination site for each of the plurality of movement methods, and contents of the prescribed activity.SELECTED DRAWING: Figure 1
Owner:HONDA MOTOR CO LTD

A machining workshop scheduling optimization method with matched tooling sets

A kind of machining workshop scheduling optimization method with complete tooling matching, complete basic information set is established for the key influencing factors of machining workshop scheduling process, the correlation between workpiece and workpiece type, workpiece type and complete tooling is clear, deep reinforcement learning (DQN) algorithm is used as the core algorithm of intelligent scheduling optimization, compared with traditional algorithm, it has stronger dynamic adaptability and online learning ability, the state space, action space and reward function of deep reinforcement learning (DQN) model are constructed, three action selection trigger conditions based on the complete tooling matching state of machining production scheduling process are designed, and the scheduling results are analyzed comprehensively, this method significantly improves production efficiency and resource utilization, and provides an efficient and intelligent scheduling solution for intelligent manufacturing.
Owner:SHENYANG JIANZHU UNIVERSITY

A deep reinforcement learning-based robot arm cooperative autonomous grasping method

The application provides a kind of mechanical arm cooperation grabbing method based on deep reinforcement learning, it is related to robot application technical field, to solve the problem of network instability in the method of multi-object grabbing in complex scene, low sample efficiency, unreasonable behavior.The method comprises: constructing a cooperative autonomous grabbing decision network model, obtaining the state-action value distribution of pushing and grabbing action by pixel-by-pixel prediction;Introduce object mask function, filter invalid area, reduce the influence of negative samples on model convergence speed;Design different behavior constraint strategy to suppress unreasonable behavior in network prediction, optimize the action selection of mechanical arm;Design a multi-element reward function, dynamically adjust the reward mechanism, guide the model to learn more action strategy that meets the task goal.The application is verified through simulation and real experiment environment, which can improve the task success rate and completion efficiency of the mechanical arm in unstructured scene, and has good robustness and applicability.
Owner:NORTHEASTERN UNIV CHINA

Motion control interface for video generation interactive graphical user interface of electronic devices

1. Name of the product in this design: Action control interface for video generation interactive graphical user interface of electronic device. 2. Purpose of this design: An electronic device. 3. The key design features of this product are: the parts within the graphical user interface, with the area outside the dotted lines being the protected area. 4. The picture or photo that best illustrates the key design points: Design 1 front view. 5. Design 1 is designated as the basic design. 6. Uses of the graphical user interface: for interactive video generation; protected local uses are used for editing and controlling the actions of people in the video. 7. Human-computer interaction method of graphical user interface: In Design 1, users can click the editing module below "Upload Reference Action" / "Text Generate Action" in the upper right corner of the main view interface of Design 1 to upload video files or edit text content to generate custom actions, or they can click any action option in the action selection module in the lower right corner of the interface to inject the action into the character action preview model on the left side of the interface. The human-computer interaction description for Design 2 is the same as that for Design 1. In Design 3, users can click the "Select Character" control in the upper left corner of the Design 3 main view interface to select the character whose action needs to be edited, or click the editing module below "Upload Reference Action" / "Text Generate Action" in the upper right corner of the interface to upload video files or edit text content to generate custom actions, or click any action option in the action selection module in the lower right corner of the interface to inject the action into the character action preview model on the left side of the interface. In Design 4, users can click the "Select Character" control in the lower left corner of the Design 4 main view interface to select the character whose action needs to be edited, or click the editing module below "Upload Reference Action" / "Text Generate Action" in the upper right corner of the interface to upload video files or edit text content to generate custom actions, or click any action option in the action selection module in the lower right corner of the interface to inject the action into the character action preview model on the left side of the interface. The human-computer interaction description of Design 5 is the same as that of Design 1. The main view interface of Design 5 is currently in the state of having uploaded reference actions and generating custom actions. The human-computer interaction description for Design 6 is the same as that for Design 1. The human-computer interaction description of Design 7 is the same as that of Design 2. In Design 8, users can click on any character in the character selection pop-up window in the upper left corner of the Design 8 main view interface to select that character. The number of character action preview models on the left side of the interface will increase or decrease according to the number of selected characters. Users can also click the "Upload Reference Action" / "Text Generate Action" editing module in the upper right corner of the interface to upload video files or edit text content to generate custom actions, or click any action option in the action selection module in the lower right corner of the interface to inject the action into the character action preview model on the left side of the interface. In Design 9, when a user clicks the editing module below "Upload Reference Action" in the upper right corner of the Design 9 main view interface, they can upload relevant videos or files and generate custom actions, presenting the interface changes from Design 9 main view to Design 9 interface change state diagram 1 to Design 9 interface change state diagram 2. In Design 10, when a user clicks the "Select Character" control in the upper left corner of the Design 10 main view interface, a character selection pop-up window will be opened, showing the interface changes from the Design 10 main view to the Design 10 interface state diagram.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Autoregressively generating sequences of data elements defining actions to be performed by an agent

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for selecting actions to be performed by an agent to interact with an environment using an action selection neural network. In one aspect, a method comprises, at each time step in a sequence of time steps: generating a current representation of a state of a task being performed by the agent in the environment as of the current time step as a sequence of data elements; autoregressively generating a sequence of data elements representing a current action to be performed by the agent at the current time step; and after autoregressively generating the sequence of data elements representing the current action, causing the agent to perform the current action at the current time step.
Owner:GDM HOLDING LLC

A workflow scheduling method considering security risk and cost optimization in cloud environment

The application provides a workflow scheduling method considering security risks and cost optimization in a cloud environment, which reduces the overall security risks and execution costs of the workflow under the given deadline constraint of a user. Firstly, the cloud computing resources are defined, the workflow structure is abstracted, and a security risk calculation method is proposed. Then, a reinforcement learning model oriented to security risks and cost optimization is constructed. Next, an action selection algorithm in the learning process is determined. Finally, iterative learning is carried out based on the Q learning algorithm idea, and finally the task scheduling scheme is determined, and the overall security risks and total costs of the workflow are obtained.
Owner:NANJING UNIV OF POSTS & TELECOMM

Multi-agent electromagnetic spectrum control attack and defense method based on deep reinforcement learning

The invention particularly relates to a multi-agent electromagnetic spectrum control attack and defense method based on deep reinforcement learning, which comprises the following steps: constructing a channel model following a discrete time two-state Markov chain, and carrying out electromagnetic spectrum control multi-agent deep reinforcement learning training based on a centralized training-decentralized execution architecture; in the attack part, a quantitative trigger is designed according to the channel signal receiving strength, an induction reward is utilized to implant a back door, the situation that higher modification cost and complexity are possibly caused by modifying a global reward is avoided, a feasible action scheme is selected for a poisoning action, and multiple agents select the same channel to collide when the back door is triggered; in a defense part, multi-agent deep reinforcement learning backdoor detection is performed based on an abnormal reward value, whether an abnormal reward mode exists or not is detected by analyzing reward distribution of a strategy in different states, and whether the reward value deviates from a normal range remarkably or not under a specific triggering condition is detected, so that a trigger is reconstructed, and an input filtering method is adopted to eliminate a backdoor.
Owner:XIDIAN UNIV

Cartesian genetic programming energy scheduling method and related equipment

PendingCN122021815AData processing applicationsBiological modelsCartesian genetic programmingMicrogrid
The embodiment of the invention provides a Cartesian genetic programming energy scheduling method and related equipment, and belongs to the field of energy scheduling and evolution calculation. The method comprises the following steps: acquiring an environment state and a discrete action set of an energy system; a CGP population is initialized; adopting an explicit action selection strategy to splice the state and each candidate action into an augmented input vector, inputting the augmented input vector into a CGP individual to obtain a scalar score and preferentially execute the scalar score, and calculating the fitness; in the evolutionary process, an active sensing sampling mutation operator is applied to generate a filial generation: active nodes in a genotype are identified through reverse traversal, and a mutation budget is dynamically allocated according to the active nodes, so that resources are inclined to the active nodes, and meanwhile, a small amount of inactive node variation is reserved. According to the method, the search efficiency and the solution quality are remarkably improved through the directional variation mechanism, and meanwhile, the generated scheduling strategy is a completely symbolized mathematical expression, has high interpretability and easy deployment, and is suitable for real-time optimization scheduling of a multi-energy microgrid and an active power distribution network.
Owner:SOUTH CHINA UNIV OF TECH

Next state prediction using likelihood estimates

ActiveUS12639726B1MarketingState predictionSet estimation
Example implementations related to next state prediction and action selection are disclosed. In an example, an initial state is estimated based on a set of priors including at least one engagement opportunity. A diminishing effect of the at least one engagement opportunity is determined using a non-linear transformation and a set of likelihood estimates for the initial state is generated using a Bayesian steady state filter model that receives the initial state and the diminishing effect of the at least one engagement opportunity. A next state is predicted based on the initial state and the set of likelihood estimates.
Owner:WALMART APOLLO LLC

A td3 map-free navigation method based on dynamic window method guidance

The application discloses a TD3 map-free navigation method based on a dynamic window method, and particularly relates to the technical field of path planning, and acquires current state information and a target position of a robot; after the current state information and the target position information are preprocessed, the information is input into an improved TD3 network to obtain TD3 output actions, and improved DWA output actions are obtained by combining an improved evaluation function and an escape mechanism; the two obtained actions are input into an action selector to output an optimal action; the robot executes the optimal action to complete a map-free navigation task. The application improves the TD3 network structure, and applies DWA (dynamic window method) and LSTM (long short-term memory network) to map-free navigation. The network structure can directly output robot actions, namely linear velocity and angular velocity, according to the input target point relative position and the robot state, and realizes end-to-end map-free navigation.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

system

The system according to this embodiment aims to select and execute the optimal action according to the situation inside the store. [Solution] The system according to the embodiment comprises an intelligent unit, an action selection unit, an execution unit, a collection unit, and a provision unit. The intelligent unit understands the situation in the store. The action selection unit selects the optimal action based on the situation understood by the intelligent unit. The execution unit executes the action selected by the action selection unit. The collection unit collects information. The provision unit provides the results.
Owner:SOFTBANK GROUP CORP

A time sequence knowledge graph reasoning method and system based on fuzzy clustering and reinforcement learning

ActiveCN121413785BBiological modelsInference methodsFuzzy clustering analysisKnowledge graph
The application discloses a kind of based on fuzzy clustering and reinforcement learning's time sequence knowledge graph inference method and system, the method includes: based on fuzzy clustering algorithm to the entity in time sequence knowledge graph is fuzzy clustering analysis, each entity is mapped to multiple clustering clusters, reinforcement learning environment is constructed based on clustering cluster, and based on search strategy control agent searches reinforcement learning environment, candidate action space is generated based on the action in the process of agent search, and the candidate action in candidate action space is scored, obtains the target score of candidate action, target score is converted into the action selection probability of agent strategy network, the action selection probability is used to drive agent to select action, based on the target entity searched by agent to the missing component of quadruple in time sequence knowledge graph is predicted, to significantly reduce the redundancy of agent search action space, improve agent decision efficiency, effectively improve the inference efficiency and accuracy of time sequence knowledge graph.
Owner:NAT UNIV OF DEFENSE TECH

Decision-making device, unmanned system, decision-making method, and program

A decision-making device (2) comprising: an action selection unit (200) for selecting one of a plurality of actions that can be taken in a first state so that an environment performs the selected action; a state acquisition unit (201) for acquiring a second state indicating the state that follows the execution of the action; a reward acquisition unit (202) for acquiring a reward serving as an indicator for indicating whether or not the second state is desirable; and a storage processing unit (203) whereby experience data in which the first state, the action, the second state, and the reward are associated is stored in a storage unit (21) associated with the action, the storage unit (21) being one of a plurality of storage units.
Owner:MITSUBISHI HEAVY IND LTD

Two-dimensional packing optimization method based on improved skyline algorithm and reinforcement learning

The invention belongs to the technical field of two-dimensional boxing, and particularly discloses a two-dimensional boxing optimization method based on an improved skyline algorithm and reinforcement learning, and the method comprises the steps: receiving a size parameter and a maximum bearing parameter of a box body and a to-be-boxed article set, and initializing a skyline set and the accumulated weight of the box body; aiming at the size parameter of the current article to be placed and the current skyline state, extracting a state feature vector; and circularly executing state feature vector extraction, action selection and placement operation until the article set is empty or the accumulated weight of the box body reaches the maximum weight, and outputting a final boxing scheme and a final space utilization rate. The optimization effect and the execution efficiency of two-dimensional boxing can be improved through the method and the device.
Owner:UNIT 91428 OF THE CHINESE PEOPLES LIBERATION ARMY

Reinforcement learning with integration of discriminator model

This specification describes a method performed by one or more data processing devices, comprising: sampling latent variables from a possible set of latent variables; selecting an action to be performed by an agent to interact with an environment on a time-step sequence using an action selection neural network conditioned on the sampled latent variables; determining a corresponding reward to be received for each time step in the time-step sequence using an ensemble of discriminator models; and training the action selection neural network based on the reward using reinforcement learning techniques. Each discriminator model can process observations to generate a corresponding prediction output, the corresponding prediction output predicting which latent variable the action selection neural network conditioned on to bring the environment into a state represented by the observation.
Owner:GDM HOLDING LLC

Frequency control method and system for intelligent sensing sensors in smart cities

This invention relates to the field of intelligent sensing technology in the power Internet of Things (IoT) field, specifically disclosing a frequency control method and system for intelligent sensing sensors in smart cities. The method involves a terminal obtaining the sensor's corresponding sensing state and extracting relevant information from the sensing environment. Then, an action is selected according to an action selection strategy and applied to the sensor. A reward-penalty function is used to feed back the reward to a Q-learning-based intelligent decision-making algorithm. The Q-value table is updated based on the reward and the formula, enriching the correspondence between states and actions. Through Q-learning and the Q-value table, the relationship between the sensing terminal's energy consumption and the quality and accuracy of the sensing data is continuously adjusted. This invention is applicable to IoT-based intelligent sensing terminal scenarios, providing sensing terminals with better autonomous information processing and control capabilities, and offering more reliable and efficient data support for the development of smart cities.
Owner:CHONGQING UNIV

Autoregressively generating sequences of data elements defining actions to be performed by an agent

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for selecting actions to be performed by an agent to interact with an environment using an action selection neural network. In one aspect, a method comprises, at each time step in a sequence of time steps: generating a current representation of a state of a task being performed by the agent in the environment as of the current time step as a sequence of data elements; autoregressively generating a sequence of data elements representing a current action to be performed by the agent at the current time step; and after autoregressively generating the sequence of data elements representing the current action, causing the agent to perform the current action at the current time step.
Owner:GDM HOLDING LLC