Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

334 results about "Learning agent" patented technology

A deep learning agent is any autonomous or semi-autonomous AI-driven system that uses deep learning to perform and improve at its tasks. Systems (agents) that use deep learning include chatbots, self-driving cars, expert systems, facial recognition programs and robots.

AI-Enhanced Distributed Data Compression with Privacy-Preserving Computation

An AI-enhanced distributed system for neural network-based data compression leverages reinforcement learning optimization and privacy-preserving computation across edge and central computing devices to autonomously optimize efficiency and quality. The system includes a lightweight compression subsystem at edge devices that applies privacy-preserving preprocessing and partially compresses input data before securely transmitting it to central computing devices. A reinforcement learning agent continuously monitors system performance and automatically optimizes compression parameters, model selection, and task allocation based on multi-objective rewards. The central compression subsystem processes data using AI-optimized parameters and temporal modeling components. The system incorporates hardware detection capabilities that automatically select optimal compression models based on available processing resources and implements homomorphic encryption for computation on encrypted data while coordinating federated learning across distributed devices. This AI-enhanced distributed approach improves bandwidth efficiency, energy consumption, and adaptability while ensuring data privacy and security.
Owner:ATOMBEAM TECH INC

Road structure design method based on large model and reinforcement learning

The invention belongs to the crossing field of road engineering technology, artificial intelligence and engineering mechanics, and particularly relates to a road structure design method based on a large model and reinforcement learning. A technical closed loop of'natural language input-parameter automatic mapping-specification standard value acquisition-finite element verification-reinforcement learning optimization 'is constructed: by constructing a load parameter mapping table and a material semantic encoder, the system can automatically convert natural language input into engineering parameters such as load, material, layer thickness and the like, and the problem of adaptability to non-standard working conditions is solved. And a deviation characteristic space is constructed by further combining a mechanical theory solution and finite element simulation, and the reinforcement learning agent is driven to quickly optimize under the condition of meeting theoretical constraints. Compared with a traditional trial calculation method and an existing intelligent optimization scheme, the method has the advantages that the number of design iterations and the material cost are greatly reduced, the safety and compliance of an output scheme are improved, and intelligent spanning from experience driving to theory guiding is achieved.
Owner:TONGJI UNIV

Multi-round automatic machine learning agent system based on reinforcement learning optimization

The invention provides a multi-round automatic machine learning agent system based on reinforcement learning optimization. Comprising a task analysis module used for generating an initial prompt for an MLE agent to call; the MLE agent module is used for generating an executable code; the code executor is used for generating an execution result; the evaluator is used for outputting a normalized value of each index and a code correctness identifier; the reward construction module is used for generating a reward value; the reinforcement learning optimizer is used for calculating group average return and candidate advantages and updating strategy parameters of the MLE intelligent agent module based on the candidate advantages; and the multi-round interaction control module is used for feeding back the execution result of the previous round and the reward value to the MLE agent module in the multi-round interaction process, and controlling code generation of the next round until a preset termination condition is met. According to the invention, strategy adaptive evolution, reinforcement learning optimization of fine-grained credit distribution and multi-round closed-loop automatic process improvement can be realized.
Owner:北京衔远有限公司 +1

Inspection unmanned aerial vehicle autonomous navigation path planning and obstacle avoidance method and system

The invention discloses a routing inspection unmanned aerial vehicle autonomous navigation path planning and obstacle avoidance method and system. The method comprises the following steps: initializing; the reinforcement learning agent performs training through interaction with the environment, and in each time step, the current strategy network generates an action based on an inertial exploration strategy adopting OU noise enhancement; after the unmanned aerial vehicle executes the action, the environment updates the state, and an instant reward is calculated; storing the state, the action, the instant reward and the new state of each time step in an experience playback buffer area; updating double-Q network parameters based on the data in the buffer area and updating strategy network parameters according to a delay updating mechanism; and repeatedly training until the accumulated reward of the strategy network exceeds a threshold value, and outputting the trained strategy network to control the unmanned aerial vehicle in real time. The adaptive capacity of the unmanned aerial vehicle in a complex scene is improved, and the unmanned aerial vehicle is suitable for power equipment inspection tasks in complex terrains and dense obstacle environments.
Owner:STATE GRID JIBEI ELECTRIC POWER CO LTD TANGSHAN POWER SUPPLY CO +2

System And Method For Dynamic Hyperparameter Optimization For Large Language Models Using (Few-Shot) Reinforcement Learning

Techniques for increasing the quality of output from large language models using reinforcement learning to select inference-time hyperparameters are disclosed. The large language model is configured with a set of values corresponding to a set of inference-time hyperparameters that are used to influence the output of the machine learning model after the model has been frozen. After obtaining a set of performance metrics that indicate the quality of the output, a reinforcement learning agent computes an adjustment for one or more of the hyperparameters, resulting in a modification of the hyperparameter values. Applying the new hyperparameter values, the large language model is then applied to a new set of input to generate a second output. The process iterates until the performance metrics associated with the output are satisfactory.
Owner:ORACLE INT CORP

Machine vision production line efficiency evaluation and optimization management system

The invention relates to the technical field of industrial manufacturing digital management, in particular to a machine vision production line performance evaluation and optimization management system, which comprises a data acquisition module for triggering a high-speed industrial camera array, a vibration sensor and an RFID reader through a central synchronous controller to synchronously acquire product images, equipment operation and material circulation data; the data processing and fusion module extracts product quality features based on CNN, and fuses multi-modal data through time sequence alignment normalization and an attention mechanism; the dynamic efficiency evaluation module calculates OEE, FPY and a production line balance rate in real time by means of a deep neural network; the optimization strategy generation module is used for reinforcing the learning agent to output optimization instructions such as equipment parameter adjustment; and the control execution module converts the instruction into an industrial protocol format, issues the instruction to the PLC, and verifies the effect to form a closed loop. According to the method, the data relevance and the evaluation real-time performance are improved, the dynamic state of the adaptive production line is optimized, and the efficiency improvement is facilitated.
Owner:XIAMEN BOSHIYUAN MASCH VISION TECH CO LTD

Intelligent optimization method for multi-type well seam joint control fine injection-production mode

The invention discloses an intelligent optimization method for a multi-type well seam joint control fine injection-production mode, and relates to the technical field of oil-gas field development. The method comprises the following steps: setting a well seam joint control fine injection-production mode, establishing an oil reservoir numerical simulation model in oil reservoir numerical simulation software, obtaining multiple groups of oil reservoir injection-production schemes based on a Latin hypercube sampling method, performing simulation according to each group of oil reservoir injection-production schemes by utilizing the oil reservoir numerical simulation model, generating multiple pieces of sample data, and establishing a sample database; a deep learning agent model is established, after the sample database is utilized to train and train the deep learning agent model, a particle swarm optimization algorithm is adopted to carry out single-target pre-search global optimization to obtain a preferred reference strategy, a reinforcement learning dynamic decision model is established, and a reinforcement learning agent is obtained through training based on a PPO near-end strategy optimization algorithm; and the optimal injection-production development scheme of the oil reservoir is obtained by utilizing the reinforcement learning agent, so that rapid optimization and decision support of the oil reservoir injection-production scheme in a new multi-type well seam joint control mode are realized.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

General control system strategy optimization method based on reinforcement learning

The invention relates to the technical field of industrial automation and intelligent control, in particular to a general control system strategy optimization method based on reinforcement learning, and the method comprises the steps: firstly collecting the operation data of a system, constructing a state vector of reinforcement learning, and enabling a reinforcement learning agent to fully understand the current operation condition of the system; then, the state vector is input into a reinforcement learning strategy network, an action for adjusting a control strategy is generated by the network, and the action can be used for modifying the proportion, integral or differential coefficient of PID and can also be used for adjusting the prediction step length, weight coefficient or constraint strength of model prediction control, so that the adaptive capacity of a controller to external changes is enhanced; then, a reward signal is constructed according to a response result of the reference controller; the reward function comprehensively considers the error size, the steady-state characteristic, the system energy consumption, the control smoothness and the stability requirement, so that the reinforcement learning not only pays attention to the error minimization when optimizing the strategy, but also considers the low energy consumption, the smooth action and the anti-interference performance at the same time.
Owner:ZHONGBEI UNIV

Multi-fidelity physical field reconstruction method based on Fourier neural operator transfer learning

The invention discloses a multi-fidelity physical field reconstruction method based on Fourier neural operator transfer learning, and the method comprises the steps: obtaining training data which comprises low-fidelity data and high-fidelity data; preprocessing the constructed deep learning agent model by using a Fourier neural operator and low-fidelity data to obtain a low-fidelity model; training the deep learning proxy model by using a Fourier neural operator and taking the network parameters of the low-fidelity model as initial parameters of high-fidelity training to obtain a high-fidelity proxy model; performing fine tuning on the high-fidelity proxy model by using the high-fidelity data; and predicting the physical field by using the fine-tuned high-fidelity proxy model to obtain a prediction result corresponding to the physical field. According to the invention, the training of the deep learning model is completed by using a large amount of low-fidelity data and a small amount of high-fidelity data, so that the constructed deep learning model can ensure the prediction precision of a physical field, the demand of the deep learning model for the high-fidelity data volume is reduced, and the modeling cost is reduced.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

Intelligent optimization method and system for aluminum alloy die casting forming process

The invention discloses an intelligent optimization method and system for an aluminum alloy die casting forming process, and relates to the technical field of intelligent optimizing.The intelligent optimization method comprises the steps that the multi-element content in an aluminum alloy melt is collected in real time, and the multi-element content is combined into a three-dimensional material gene vector; inputting the three-dimensional material gene vector into a pre-trained graph neural network, and calculating to obtain an injection speed compensation coefficient and a boost pressure compensation coefficient; performing space-time alignment on the geometric center coordinate of the defect and die temperature field and pressure time sequence data recorded in the die casting process, inputting a reinforcement learning agent model constructed through a depth Q network algorithm, and reconstructing an evolution path of an internal defect area; and according to the defect formation time point and position point identified in the evolution path, reversely correcting the weight parameter of the graph neural network, and updating. According to the method, the content of multiple elements is fused into the three-dimensional material gene vector, and millisecond-level cooperative compensation of the injection speed and the boost pressure is achieved through the dynamic topology modeling capacity of the pre-training graph neural network.
Owner:EDT DIECASTING TECH SUZHOU CO LTD

Partitioned rapid inversion method for global structural mechanical parameters of concrete arch dam

The invention discloses a concrete arch dam global structural mechanical parameter zoning rapid inversion method, and relates to the technical field of dam operation safety monitoring and management, and the method comprises the steps: building a dam body and foundation three-dimensional finite element model through finite element software according to engineering design and monitoring data, and building a foundation three-dimensional finite element model according to damming material mechanical parameter information; partitioning the three-dimensional finite element model, analyzing the sensitivity of mechanical parameters of each region, determining sensitive mechanical parameters influencing the deformation of the concrete arch dam, and carrying out self-adaptive intelligent sampling on the sensitive mechanical parameters according to a sensitivity analysis result, and constructing a deep learning agent model reflecting a nonlinear relationship between the sensitive mechanical parameters of the dam and the deformation of each monitoring point, and carrying out deep learning inversion on the elastic modulus of the dam body and the deformation modulus of the bedrock. According to the invention, the method can efficiently and accurately invert and determine the structural mechanical parameters of the arch dam in the actual operation period, and provides a good basis for the safety analysis of the dam.
Owner:NANCHANG UNIV

Cross-border e-commerce advertisement accurate putting method and system based on user portraits

The invention relates to the technical field of cross-border advertisement putting, and provides a cross-border e-commerce advertisement accurate putting method and system based on a user portrait, and the method comprises the steps: building a unified data flow through multi-source data collection and standardization processing; calculating and fusing the real-time statistical features, the historical interest vectors and the context features in real time by using a stream processing engine to generate multi-dimensional feature vectors; inputting the feature vector into a multi-modal time sequence neural network model which captures a user behavior rule through time sequence coding and an attention mechanism and outputs a user behavior weight and an interest attenuation coefficient; dynamically updating the real-time interest score of the user by adopting an exponential decay model; and when the interest score reaches a threshold value, triggering a reinforcement learning agent to make a decision according to the comprehensive state information, and dynamically outputting an action strategy including release triggering, a bidding coefficient and a creative type. The problems that in the prior art, user interest modeling lags behind, and the self-adaptive capacity of the putting strategy is poor are effectively solved.
Owner:NANJING YUSITUOMENG INTERNATIONAL TRADING CO LTD

Neural network automatic pruning method based on GRPO reinforcement learning

The invention belongs to the technical field of artificial intelligence, and particularly relates to a neural network automatic pruning method based on GRPO reinforcement learning, and the method comprises the steps: introducing a dynamic scaling factor into a batch normalization layer of a to-be-pruned neural network, and calculating the importance score of each convolution layer channel of the to-be-pruned neural network in combination with an attention mechanism; s2, constructing a multi-dimensional state vector containing layer structure features based on an importance calculation result in the step S1; s2, inputting the multi-dimensional state vector constructed in S2 into a strategy network of a GRPO reinforcement learning agent, generating a pruning action by the strategy network according to state information of a current network layer, and defining the action to represent a pruning rate of the layer; according to the method, a GRPO reinforcement learning algorithm is adopted, a traditional Critic model is abandoned, the strategy calculation process is simplified through a group sampling-relative advantage estimation mechanism, and memory occupation is remarkably reduced.
Owner:SHANDONG UNIV

Machine-generated training examples for training machine-learned models

An example method includes obtaining a reference trace describing interactions between a machine-learned agent system and an environment, wherein the reference trace includes a reference sequence of interaction objects associated with performance of a task, each respective reference interaction object of the reference sequence of interaction objects corresponding to a respective reference environment state and a respective reference action executed on the respective reference environment state. The example method includes generating, using a machine-learned trace generation system, a forecasted trace based on the reference trace, the forecasted trace comprising a forecasted sequence of interaction objects that is predicted to continue the reference sequence of interaction objects from a branching position toward performance of the task, each respective forecasted interaction object of the forecasted sequence of interaction objects corresponding to a respective forecasted environment state and a respective forecasted action executed on the respective forecasted environment state. The example method includes training the machine-learned agent system using the forecasted trace sequence.
Owner:GOOGLE LLC

Deep learning assisted acceleration fracturing construction parameter intelligent real-time optimization method

The invention discloses a deep learning assisted acceleration fracturing construction parameter intelligent real-time optimization method, and belongs to the field of unconventional oil and gas field development, and the method comprises the steps: constructing a structured sample database covering geological attributes, engineering parameters and fracturing effect responses, and carrying out the unified preprocessing of multi-source data; constructing a multi-modal space-time collaborative prediction model, and combining a multi-expansion time permissible convolutional network, a three-dimensional residual network and a cross attention mechanism; supervised learning training is carried out based on a database, sparse regularization and learning rate scheduling are introduced, and the model precision and generalization ability are improved; a hierarchical collaborative optimization framework of outer-layer Bayesian search-inner-layer CMA-ES refinement is provided, and reservoir transformation volume maximization and construction feasibility are achieved under complex constraints; compared with an existing method, optimization parameters of the method include the cluster distance and the section distance and further include the displacement, the sand concentration, the fracturing fluid type and other construction parameters, and real-time optimization of the fracturing construction parameters can be achieved through the machine learning agent model.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

Dynamic computing power distribution method and system based on reinforcement learning

The invention belongs to the technical field of computing power distribution, and particularly relates to a dynamic computing power distribution method and system based on reinforcement learning, and the method comprises the following specific steps: S1, covering cloud, edge and end full-node scenes, and collecting computing power resource states, task demand features and cross-domain network condition data in real time; s2, on the basis of standardized data output by a cross-domain computing power sensing module, by constructing a state space fusing computing power, tasks and a network, defining an action space of computing power scheduling direction and proportion, and designing a multi-target reward function for balancing the resource utilization rate, the task satisfaction rate and long-term conflict avoidance; and realizing self-learning and self-iteration scheduling strategy generation based on a reinforcement learning algorithm. According to the invention, the reinforcement learning agent autonomously learns the computing power demand of the emergency scene and the new type of task, the rule does not need to be manually preset and modified, and the method has the advantage of realizing dynamic adaptation of computing power distribution to complex and changeable scenes.
Owner:BEIJING CENTURY FEIXUN TECH CO LTD

Multi-parameter and multi-field intelligent optimization method and system for press free forging large-scale die casting

The invention relates to a multi-parameter and multi-field intelligent optimization method and system for a press free forging large-scale falling die, and the method achieves the collaborative optimization of technological parameters, die geometry and microstructure through a heat-force-microstructure three-field coupling modeling, a deep kernel learning agent model and a digital twinning guided multi-objective evolutionary algorithm, improves the optimization efficiency, and improves the optimization precision. The mold testing times are reduced; and meanwhile, a real-time closed-loop control system is constructed, the grain size uniformity of forgings is improved, the forming load is reduced, the production period is shortened, the process stability is improved, and the die service life is prolonged.
Owner:ZHEJIANG JIEDE MASCH TECH CO LTD

Training reinforcement learning agents to learn farsighted behaviors by predicting in latent space

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training an action selection policy neural network used to select an action to be performed by an agent interacting with an environment. In one aspect, a method includes: receiving a latent representation characterizing a current state of the environment; generating a trajectory of latent representations that starts with the received latent representation; for each latent representation in the trajectory: determining a predicted reward; and processing the state latent representation using a value neural network to generate a predicted state value; determining a corresponding target state value for each latent representation in the trajectory; determining, based on the target state values, an update to the current values of the policy neural network parameters; and determining an update to the current values of the value neural network parameters.
Owner:GOOGLE LLC

Intelligent decision-making system and method for corn fertilization based on mechanism-data dual-drive fusion

The invention belongs to the technical field of unmanned aerial vehicle remote sensing and agriculture combination, and discloses an intelligent decision-making system and method for corn fertilization based on mechanism-data dual-drive fusion. The system comprises a mechanism simulation module, a data preprocessing module, a feature engineering module, a modeling module, a visualization module, a decision support module, a dynamic feedback correction module and a report generation and push module. According to the mechanism-data double-drive fusion-based intelligent decision-making system and method for corn fertilization, a large-area corn field block image is obtained in a short time through an unmanned aerial vehicle multispectral system, and the nitrogen diagnosis efficiency is improved; through a lightweight machine learning agent model, corn canopy leaf nitrogen nutrition parameters are accurately predicted, and a basis is provided for accurate fertilization; the system monitors the nitrogen nutrition status of the corn in real time, and provides possibility for dynamically adjusting a fertilization strategy; the multispectral remote sensing technology can perform nitrogen nutrition diagnosis under the condition of not damaging corn plants, and the corn growth environment is protected.
Owner:AGRI SCI RES INST OF THE SEVENTH DIVISION OF XINJIANG PROD & CONSTR CORPS

Mobile phone thermal simulation method and system based on multi-physics field coupling

The invention discloses a mobile phone thermal simulation method and system based on multi-physics field coupling, relates to the field of mobile phone thermal simulation, and constructs hybrid simulation coupling machine learning and physical solution. A machine learning agent model and a throttling logic script are innovatively introduced. The machine learning agent model is used for quickly predicting the temperature according to the current power consumption so as to instantaneously respond to the power consumption change; and the throttling logic script is used for simulating a real temperature control frequency reduction strategy and dynamically adjusting the target power consumption at the next moment according to the predicted temperature. In this way, the original black box temperature control logic is explicit, and a power consumption-temperature closed-loop feedback path is constructed. And finally, taking the adjusted power consumption as a heat source, and performing physical field solving by a traditional CAE solver. Therefore, the problem of time scale difference between heat conduction and electrical change is effectively solved, and efficient and accurate simulation of true performance and temperature performance of the mobile phone in a long-time high-load scene is realized.
Owner:SHENZHEN DUOKE ELECTRONICS CO LTD

One-dimensional pneumatic design method and device for axial flow compressor

The invention provides a one-dimensional pneumatic design method and device for an axial flow compressor, and belongs to the field of axial flow compressor pneumatic design. The method comprises the following steps: taking values of a plurality of design variables as the state of a reinforcement learning agent; the intelligent agent outputs the adjustment amount of the design variable according to the state, and the adjustment amount serves as the action executed by the intelligent agent; applying the action to the value of the current design variable to obtain an updated design variable; utilizing a one-dimensional inverse problem design program to generate a one-dimensional pneumatic design scheme of the gas compressor; utilizing a one-dimensional normal problem analysis program to output performance indexes; constructing a reward function according to the performance indexes; based on feedback provided by the reward function, a depth deterministic strategy gradient algorithm is adopted to train the intelligent agent; and performing one-dimensional pneumatic design of the gas compressor by using the trained intelligent agent. According to the method and device provided by the invention, the problems that the one-dimensional design of the axial flow compressor depends on experience, the efficiency is low, automatic optimization is difficult and the response is slow can be solved.
Owner:JINCHENG NANJING ELECTROMECHANICAL HYDRAULIC PRESSURE ENG RES CENT AVIATION IND OF CHINA

Explanatable farmland image enhancement method and system based on physical perception and reinforcement learning

ActiveCN120876343AImage enhancementImage analysisImaging processingIncrement threshold
The invention relates to the technical field of image processing, in particular to an interpretable farmland image enhancement method based on physical perception and reinforcement learning, and the method comprises the steps: obtaining an original farmland image, carrying out the global perception analysis, extracting multi-scale features, and recognizing a degradation region and feature distribution; on the basis of the global perception analysis result, constructing a semantic enhancement blueprint for quantifying the physical attributes and optimization requirements of the degradation area; initializing a reinforcement learning agent according to the semantic enhancement blueprint, and selecting an image processing operation sequence through a physical constraint reward function; executing the image operation sequence, and terminating enhancement processing according to the quality evaluation index increment threshold and the physical consistency of the semantic enhancement blueprint; and outputting an interpretability report of the enhanced image and associated image operation sequence physical basis traceability. The objective of the invention is to solve the technical problems of lack of interpretability, insufficient environmental adaptability and rigid decision-making mechanism of a farmland image enhancement technology.
Owner:CHINA TOWER CO LTD

Multi-objective optimization method and system for urban block form

The invention discloses a multi-objective optimization method for the form of an urban block, and the method comprises the following steps: S100, constructing a parameterized model, carrying out the experimental design, defining a group of geometric parameters used for describing the form of the urban block, and generating a plurality of sample points in the value range of the geometric parameters through employing a sampling method; s200, bidirectional coupling numerical simulation is executed, a data set is generated, at least one thermal comfort evaluation index and one air quality evaluation index are output, and therefore the data set containing input parameters and output indexes is formed; step S300, constructing and verifying a machine learning agent model to obtain an agent model capable of rapidly predicting the evaluation index; step S400, executing multi-objective optimization search, and automatically searching to enable the thermal comfort evaluation index and the air quality evaluation index to simultaneously achieve a better Pareto optimal solution set; and step S500, analyzing and outputting an optimization result, extracting an optimization value range of the geometric parameters, and generating a design guide.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Power transmission line project quality defect acceptance method based on generative adversarial and reinforcement learning

The invention discloses a power transmission line engineering quality defect acceptance method based on generative adversarial and reinforcement learning, and relates to the technical field of power engineering quality detection and intelligent image recognition, and the method comprises the steps: carrying out the sample amplification of original image data of a power transmission line tower through a generative adversarial network module, and obtaining an amplified training data set; constructing a PaFPN feature extraction network according to the amplification training data set, extracting multi-scale defect feature vectors and generating a coding feature matrix; establishing a reinforcement learning agent module, taking the coding feature matrix as state input, learning an optimal defect detection strategy through a Q-learning algorithm, and outputting a defect detection parameter combination; and performing feature fusion on the coding feature matrix to generate a fusion feature vector, inputting the fusion feature vector into a classifier network, and outputting a defect category label and a confidence score of the power transmission line project quality. According to the invention, the automation level and the detection precision of power transmission line project quality defect acceptance are improved.
Owner:SUZHOU POWER CONSTR ENG CO LTD

Power distribution network optimization scheduling method and system based on multi-target deep reinforcement learning

The invention discloses a power distribution network optimization scheduling method and system based on multi-target deep reinforcement learning, and relates to the technical field of power distribution network optimization operation, and the method comprises the steps: building a power distribution network operation model and a probabilistic load flow model considering the photovoltaic output and load demand uncertainty, completing the probabilistic load flow calculation under the uncertainty, and obtaining a power distribution network optimal scheduling model; analyzing to obtain a probability density function of node voltage and line power flow of the power distribution network; introducing a utility function based on preference, establishing voltage out-of-limit and line overload risk indexes considering severity weight, and constructing a risk-economic collaborative power distribution network multi-objective optimization operation problem; the decision process of the problem is modeled into a multi-target Markov decision process, a decomposition-based multi-target deep reinforcement learning algorithm is adopted, learning and training of the decision process are performed on reinforcement learning agents, a Pareto strategy set is obtained, an optimal operation strategy is screened out, and optimization regulation and control are performed on the power distribution network. And collaborative optimization of the operation risk and cost of the power distribution network system is realized.
Owner:SHANDONG UNIV

Big data-based collaborative self-optimization method and system for self-growth risk control system

The invention relates to the technical field of computer risk control, in particular to a big-data-based collaborative self-optimization method and system for a self-growth risk control system, and the method comprises the steps: obtaining the operation data of the risk control system, so as to construct a state vector; inputting the state vector into a reinforcement learning agent, and selecting one coordination action from a preset coordination action space for output; the coordination action is executed in the risk control system, and a scalar reward value used for evaluating the effect of the coordination action is obtained through calculation; and forming an empirical data tuple by using the state vector, the coordination action, the scalar reward value and a new state of the system after the action is executed, and training the reinforcement learning agent to update a decision strategy of the reinforcement learning agent. According to the risk control system, the rule engine and the model engine in the risk control system can be considered as a unified whole, and respective parameters are dynamically and cooperatively adjusted, so that the overall comprehensive efficiency of the system is maximized, and the self-adaptive capability and the iteration efficiency of the system are improved.
Owner:RED STAR MACALLINE GRP

Topological network global optimization method based on graph reinforcement learning

The invention discloses a topological network global optimization method based on graph reinforcement learning, and relates to the field of computational geometry and grid optimizing.The method comprises the steps that a topological optimization process is modeled into a Markov decision process, a topological network containing singular points and angular points is converted into heterogeneous graph representation, and a graph state is constructed; selecting a topology editing action by a reinforcement learning agent fused with the graph neural network based on the current graph state; after executing the action, updating a topological structure and a graph state, and calculating an instant reward based on the global grid quality index change; training intelligent agents in batches through experience playback and time sequence differential learning, so that the intelligent agents learn a mapping strategy from a graph state to an optimal action; and applying the trained policy network deployment to automatic optimization of the new topology. According to the method, full-process automation of topological optimization is achieved, non-local negative effects of local modification are effectively avoided, a global optimal optimization strategy is learned, the grid quality and optimization efficiency are remarkably improved, and the method is suitable for simulation grid generation of aircrafts.
Owner:CALCULATION AERODYNAMICS INST CHINA AERODYNAMICS RES & DEV CENT

Agent prompt generation using tool description embeddings

Methods, systems, and apparatuses include receiving, via a conversational interface, user input from a user of an online system. A user input embedding is generated for the user input. A vector store is retrieved including tool description embeddings. A similarity search is performed using the user input embedding and the tool description embeddings. A set of tool descriptions is determined using results of the similarity search. A prompt is generated using the set of tool descriptions and the user input. Machine learning agents are applied to the prompt to cause the machine learning agents to use tools associated with the set of tool descriptions. A response to the prompt is received, from the machine learning agents, in response to the machine learning agents using the tools. An output to the user input based on the response is sent, via the conversational interface, to the user of the online system.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Intelligent production line operation and maintenance strategy optimization method based on condition generation model

The invention discloses an intelligent production line operation and maintenance strategy optimization method based on a condition generation model, and the method comprises the specific steps: (1) collecting time series data of an equipment operation state, and constructing a sample data set; (2) preprocessing the sample data set to obtain a preprocessed data set, and dividing the preprocessed data set into a training set and a verification set; (3) a conditional VAE world model is constructed through a conditional variation auto-encoder, a Transform decoder and an RUL prediction module based on a historical sequence; (4) training a reinforcement learning agent based on the world model, and learning an optimal maintenance strategy; and (5) carrying out equipment state prediction and maintenance decision making by utilizing the trained world model and the intelligent agent. According to the method, state transition dynamics is modeled through conditional VAE, state transition is predicted by using a Transform decoder, an independent structure is designed to predict RUL and rewards, the prediction precision is improved in combination with historical sequence information, and decision support is provided for intelligent maintenance.
Owner:KUNMING UNIV OF SCI & TECH

Urban shallow lake ecological protection and restoration method and system integrated with intelligent management and control

The invention discloses an urban shallow lake ecological protection and restoration method and system integrated with intelligent management and control, and relates to the technical field of water environment ecological management and intelligent water affairs, and the method comprises the steps: constructing a three-dimensional monitoring network to synchronously collect water quality, image and meteorological data, and carrying out the fusion to generate a multi-modal data set; a deep learning model is utilized to realize pollution source tracing, water quality prediction and algae bloom early warning in parallel; based on the prediction result, outputting a multi-facility cooperative regulation strategy through a reinforcement learning agent; and converting the strategy into a hierarchical instruction to drive an execution unit, and feeding back the treated environment state to the model and the intelligent agent to form closed-loop optimization. According to the invention, whole-course intelligent management and control from monitoring to execution are realized, systematicness, accuracy and perspectiveness of lake treatment are effectively improved, and water environment risk response capability and ecological restoration effect are significantly enhanced.
Owner:GUANGDONG RES INST OF WATER RESOURCES & HYDROPOWER