Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

422 results about "Learning agent" patented technology

A deep learning agent is any autonomous or semi-autonomous AI-driven system that uses deep learning to perform and improve at its tasks. Systems (agents) that use deep learning include chatbots, self-driving cars, expert systems, facial recognition programs and robots.

AI-Enhanced Distributed Data Compression with Privacy-Preserving Computation

An AI-enhanced distributed system for neural network-based data compression leverages reinforcement learning optimization and privacy-preserving computation across edge and central computing devices to autonomously optimize efficiency and quality. The system includes a lightweight compression subsystem at edge devices that applies privacy-preserving preprocessing and partially compresses input data before securely transmitting it to central computing devices. A reinforcement learning agent continuously monitors system performance and automatically optimizes compression parameters, model selection, and task allocation based on multi-objective rewards. The central compression subsystem processes data using AI-optimized parameters and temporal modeling components. The system incorporates hardware detection capabilities that automatically select optimal compression models based on available processing resources and implements homomorphic encryption for computation on encrypted data while coordinating federated learning across distributed devices. This AI-enhanced distributed approach improves bandwidth efficiency, energy consumption, and adaptability while ensuring data privacy and security.
Owner:ATOMBEAM TECH INC

Water supply network water hammer control method and system based on multifunctional module fusion

The invention discloses a water supply pipe network water hammer control method and system based on multifunctional module fusion, and relates to the technical field of data identification. A time sequence diagram neural network prediction and traceability module which realizes water hammer risk prediction and propagation path traceability based on a causal constraint graph neural network model; the reinforcement learning intervention decision module is used for generating an active intervention strategy through a reinforcement learning agent, forming closed-loop control, abstracting a water supply pipe network into a graph structure, and modeling in combination with a time sequence, so that a propagation path of pressure waves can be comprehensively reflected, a blind area of traditional local modeling is overcome, and the comprehensiveness and accuracy of water hammer event detection are improved; and a causal analysis result is input as an adjacent matrix, so that the interference of irrelevant edges on prediction in topology is effectively eliminated, and the learning efficiency and causal traceability of the model are improved.
Owner:GREATER BAY AREA INST FOR INNOVATION HUNAN UNIV

Strategy optimization method and device based on interaction track, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of robot strategy training, financial science and technology, medical treatment and health and the like, and discloses a strategy optimization method and device based on an interaction track, equipment and a medium. The reinforcement learning agent interacts with the environment to generate an action and state sequence and record an interaction track, an optimal action strategy is generated based on a track optimization strategy, an original empirical data set is further generated, and a pre-training model is optimized through supervised learning to obtain a target strategy model. According to the method, a task-related high-quality trajectory is generated through reinforcement learning agent and environment interaction, original empirical data is extracted on this basis, and a pre-training strategy model is optimized in combination with a supervised learning mechanism, so that the sample utilization efficiency is effectively improved, and the generalization ability and execution robustness of the model under a multi-task condition are enhanced.
Owner:PING AN TECH (SHENZHEN) CO LTD

Road structure design method based on large model and reinforcement learning

The invention belongs to the crossing field of road engineering technology, artificial intelligence and engineering mechanics, and particularly relates to a road structure design method based on a large model and reinforcement learning. A technical closed loop of'natural language input-parameter automatic mapping-specification standard value acquisition-finite element verification-reinforcement learning optimization 'is constructed: by constructing a load parameter mapping table and a material semantic encoder, the system can automatically convert natural language input into engineering parameters such as load, material, layer thickness and the like, and the problem of adaptability to non-standard working conditions is solved. And a deviation characteristic space is constructed by further combining a mechanical theory solution and finite element simulation, and the reinforcement learning agent is driven to quickly optimize under the condition of meeting theoretical constraints. Compared with a traditional trial calculation method and an existing intelligent optimization scheme, the method has the advantages that the number of design iterations and the material cost are greatly reduced, the safety and compliance of an output scheme are improved, and intelligent spanning from experience driving to theory guiding is achieved.
Owner:TONGJI UNIV

Multi-round automatic machine learning agent system based on reinforcement learning optimization

The invention provides a multi-round automatic machine learning agent system based on reinforcement learning optimization. Comprising a task analysis module used for generating an initial prompt for an MLE agent to call; the MLE agent module is used for generating an executable code; the code executor is used for generating an execution result; the evaluator is used for outputting a normalized value of each index and a code correctness identifier; the reward construction module is used for generating a reward value; the reinforcement learning optimizer is used for calculating group average return and candidate advantages and updating strategy parameters of the MLE intelligent agent module based on the candidate advantages; and the multi-round interaction control module is used for feeding back the execution result of the previous round and the reward value to the MLE agent module in the multi-round interaction process, and controlling code generation of the next round until a preset termination condition is met. According to the invention, strategy adaptive evolution, reinforcement learning optimization of fine-grained credit distribution and multi-round closed-loop automatic process improvement can be realized.
Owner:北京衔远有限公司 +1

Inspection unmanned aerial vehicle autonomous navigation path planning and obstacle avoidance method and system

The invention discloses a routing inspection unmanned aerial vehicle autonomous navigation path planning and obstacle avoidance method and system. The method comprises the following steps: initializing; the reinforcement learning agent performs training through interaction with the environment, and in each time step, the current strategy network generates an action based on an inertial exploration strategy adopting OU noise enhancement; after the unmanned aerial vehicle executes the action, the environment updates the state, and an instant reward is calculated; storing the state, the action, the instant reward and the new state of each time step in an experience playback buffer area; updating double-Q network parameters based on the data in the buffer area and updating strategy network parameters according to a delay updating mechanism; and repeatedly training until the accumulated reward of the strategy network exceeds a threshold value, and outputting the trained strategy network to control the unmanned aerial vehicle in real time. The adaptive capacity of the unmanned aerial vehicle in a complex scene is improved, and the unmanned aerial vehicle is suitable for power equipment inspection tasks in complex terrains and dense obstacle environments.
Owner:STATE GRID JIBEI ELECTRIC POWER CO LTD TANGSHAN POWER SUPPLY CO +2

System And Method For Dynamic Hyperparameter Optimization For Large Language Models Using (Few-Shot) Reinforcement Learning

Techniques for increasing the quality of output from large language models using reinforcement learning to select inference-time hyperparameters are disclosed. The large language model is configured with a set of values corresponding to a set of inference-time hyperparameters that are used to influence the output of the machine learning model after the model has been frozen. After obtaining a set of performance metrics that indicate the quality of the output, a reinforcement learning agent computes an adjustment for one or more of the hyperparameters, resulting in a modification of the hyperparameter values. Applying the new hyperparameter values, the large language model is then applied to a new set of input to generate a second output. The process iterates until the performance metrics associated with the output are satisfactory.
Owner:ORACLE INT CORP

Electric power business data analysis processing method and system

The invention provides a power business data analysis processing method and system. The method comprises the following steps: receiving a business expansion work order containing a structured text and an unstructured image through a business hall terminal; analyzing text key entity fields, extracting image anti-counterfeiting feature points, and mapping the image anti-counterfeiting feature points to a unified vector space through a multi-modal alignment layer to generate semantic image feature flow; loading to a hardware acceleration chip to execute parallel matrix operation, identifying a logic binding relation between entities, and outputting a nested relation topological graph; calling a historical work order decision path to train a reinforcement learning agent to generate a control strategy; comparing the similarity between the to-be-anti-fake feature point and a pre-stored template, and outputting a verification result containing a missing identifier and anti-fake abnormity; and if the field is complete and the anti-counterfeiting similarity is greater than or equal to a threshold value, releasing the work order to a service system, otherwise, freezing the work order and triggering an alarm. According to the invention, the processing efficiency is improved, the abnormal generation risk is blocked, and automatic anti-counterfeiting verification and flow control of the electric power work order are realized.
Owner:BEIJING SHUYANG SMART TECH CO LTD

Machine vision production line efficiency evaluation and optimization management system

The invention relates to the technical field of industrial manufacturing digital management, in particular to a machine vision production line performance evaluation and optimization management system, which comprises a data acquisition module for triggering a high-speed industrial camera array, a vibration sensor and an RFID reader through a central synchronous controller to synchronously acquire product images, equipment operation and material circulation data; the data processing and fusion module extracts product quality features based on CNN, and fuses multi-modal data through time sequence alignment normalization and an attention mechanism; the dynamic efficiency evaluation module calculates OEE, FPY and a production line balance rate in real time by means of a deep neural network; the optimization strategy generation module is used for reinforcing the learning agent to output optimization instructions such as equipment parameter adjustment; and the control execution module converts the instruction into an industrial protocol format, issues the instruction to the PLC, and verifies the effect to form a closed loop. According to the method, the data relevance and the evaluation real-time performance are improved, the dynamic state of the adaptive production line is optimized, and the efficiency improvement is facilitated.
Owner:XIAMEN BOSHIYUAN MASCH VISION TECH CO LTD

Self-adaptive working condition sensing fuel cell hybrid tramcar hierarchical management method

The invention discloses a layered energy management method of a fuel cell hybrid tramcar with self-adaptive working condition perception. In the recognition layer, a sliding window mechanism is adopted to extract time domain and frequency domain features of load conditions, feature data are clustered based on a spectral clustering algorithm driven by a deep auto-encoder, a data set with category labels is obtained, and a deep dynamic learning vector quantization neural network classifier is trained; in the strategy layer, a double-delay depth deterministic strategy gradient reinforcement learning algorithm is adopted, a reward function is constructed, and lithium battery SOC fluctuation penalty term limit parameters in the reward function are adaptively adjusted according to the real-time load working condition category output by the recognition layer; training the reinforcement learning agent to obtain an optimal power distribution scheme between the multi-stack fuel cell power generation system and the lithium battery; and according to the performance degradation degrees of different fuel cell stacks, a distributed cooperative control strategy considering performance difference is adopted to distribute the output power of each stack, so that the coordinated control of the running state of the multi-stack fuel cell power generation system is realized.
Owner:SOUTHWEST JIAOTONG UNIV +1

Transmission path selection method and device based on reinforcement learning, and electronic equipment

The invention discloses a transmission path selection method and device based on reinforcement learning and electronic equipment, and relates to the technical field of big data, and the method comprises the following steps: obtaining node state information and link state information transmitted by each node in a target network topology, generating a routing selection action based on the node state information and the link state information, and sending the routing selection action to the target network topology; based on the node state information, the link state information and the corresponding route selection action, updating a target Q value table by adopting a reinforcement learning agent center; after a new route transmission request is received, based on the updated target Q value table, the reinforcement learning proxy center is adopted to select a target data transmission path for the current network node, and the target data transmission path is used for completing route transmission of a data packet in the new route transmission request. The technical problem that a static routing algorithm in the prior art lacks flexibility and cannot be adjusted in real time according to network topology and flow change is solved.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Intelligent optimization method for multi-type well seam joint control fine injection-production mode

The invention discloses an intelligent optimization method for a multi-type well seam joint control fine injection-production mode, and relates to the technical field of oil-gas field development. The method comprises the following steps: setting a well seam joint control fine injection-production mode, establishing an oil reservoir numerical simulation model in oil reservoir numerical simulation software, obtaining multiple groups of oil reservoir injection-production schemes based on a Latin hypercube sampling method, performing simulation according to each group of oil reservoir injection-production schemes by utilizing the oil reservoir numerical simulation model, generating multiple pieces of sample data, and establishing a sample database; a deep learning agent model is established, after the sample database is utilized to train and train the deep learning agent model, a particle swarm optimization algorithm is adopted to carry out single-target pre-search global optimization to obtain a preferred reference strategy, a reinforcement learning dynamic decision model is established, and a reinforcement learning agent is obtained through training based on a PPO near-end strategy optimization algorithm; and the optimal injection-production development scheme of the oil reservoir is obtained by utilizing the reinforcement learning agent, so that rapid optimization and decision support of the oil reservoir injection-production scheme in a new multi-type well seam joint control mode are realized.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

General control system strategy optimization method based on reinforcement learning

The invention relates to the technical field of industrial automation and intelligent control, in particular to a general control system strategy optimization method based on reinforcement learning, and the method comprises the steps: firstly collecting the operation data of a system, constructing a state vector of reinforcement learning, and enabling a reinforcement learning agent to fully understand the current operation condition of the system; then, the state vector is input into a reinforcement learning strategy network, an action for adjusting a control strategy is generated by the network, and the action can be used for modifying the proportion, integral or differential coefficient of PID and can also be used for adjusting the prediction step length, weight coefficient or constraint strength of model prediction control, so that the adaptive capacity of a controller to external changes is enhanced; then, a reward signal is constructed according to a response result of the reference controller; the reward function comprehensively considers the error size, the steady-state characteristic, the system energy consumption, the control smoothness and the stability requirement, so that the reinforcement learning not only pays attention to the error minimization when optimizing the strategy, but also considers the low energy consumption, the smooth action and the anti-interference performance at the same time.
Owner:ZHONGBEI UNIV

Hierarchical Machine-Learned Agents For Performing Mixed Sequence Processing Tasks

A computing device can obtain a first machine-learned sequence processing model configured to use a plurality of first tools, wherein at least one first tool of the plurality of first tools is a second machine-learned sequence processing model configured to use one or more second tools. The computing device can obtain an input context. The computing device can select, using the first machine-learned sequence processing model based at least in part on the input context, a first tool of the plurality of first tools, wherein the first tool selected is the second machine-learned sequence processing model. The computing device can select, using the second machine-learned sequence processing model, at least one second tool of the one or more second tools. The computing device can generate, using the at least one second tool of the one or more second tools, a first output.
Owner:GOOGLE LLC

Multi-fidelity physical field reconstruction method based on Fourier neural operator transfer learning

The invention discloses a multi-fidelity physical field reconstruction method based on Fourier neural operator transfer learning, and the method comprises the steps: obtaining training data which comprises low-fidelity data and high-fidelity data; preprocessing the constructed deep learning agent model by using a Fourier neural operator and low-fidelity data to obtain a low-fidelity model; training the deep learning proxy model by using a Fourier neural operator and taking the network parameters of the low-fidelity model as initial parameters of high-fidelity training to obtain a high-fidelity proxy model; performing fine tuning on the high-fidelity proxy model by using the high-fidelity data; and predicting the physical field by using the fine-tuned high-fidelity proxy model to obtain a prediction result corresponding to the physical field. According to the invention, the training of the deep learning model is completed by using a large amount of low-fidelity data and a small amount of high-fidelity data, so that the constructed deep learning model can ensure the prediction precision of a physical field, the demand of the deep learning model for the high-fidelity data volume is reduced, and the modeling cost is reduced.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

Engine turbine shaft sequential optimization design method based on ensemble learning

The invention discloses an engine turbine shaft sequential optimization design method based on ensemble learning, and the method comprises the steps: carrying out the parametric modeling of an aero-engine turbine shaft through employing a finite element method, and determining a design variable; determining an optimized objective function and constraint conditions based on the turbine shaft structure; constructing a training point set by using the target function and the constraint condition, and establishing an ensemble learning agent model by using the training point set through an ensemble learning algorithm; executing a sequential optimization design process based on the ensemble learning agent model; searching a current optimal sequence design point by utilizing an intelligent algorithm and a sequential updating criterion, and judging whether the sequence design point meets a threshold value requirement or not; if the current optimal sequence design point meets the threshold requirement, outputting the current optimal sequence design point as a design result; otherwise, adding the sequence design points as new training points into the training point set to update the training point set, updating the ensemble learning agent model by using the updated training point set, and iteratively calculating the sequence design points.
Owner:XIAN MODERN CONTROL TECH RES INST

Two-dimensional principal component analysis method for reinforcement learning of multi-section airfoil optimization strategy

The invention discloses a two-dimensional principal component analysis method for reinforcement learning of a multi-section airfoil optimization strategy, and belongs to the technical field of aircrafts, and the method comprises the following steps: S1, defining a multi-section airfoil optimization problem; s2, establishing a pneumatic data sample library by using Latin hypercube sampling; s3, performing dimension reduction on the velocity field matrix by adopting a two-dimensional principal component analysis method; s4, establishing a reinforcement learning model of the optimization strategy; and S5, training the reinforcement learning model to obtain an optimal optimization strategy with direct migration capability. According to the method, while the effectiveness of the reinforcement learning agent in observing the environment state is ensured, the dimensionality of the state is effectively reduced, the number of layers of the neural network and the number of parameters required by optimization strategy learning are reduced, and the trained optimization strategy is suitable for various design working conditions and multi-section airfoil profiles.
Owner:BEIHANG UNIV

Intelligent optimization method and system for aluminum alloy die casting forming process

The invention discloses an intelligent optimization method and system for an aluminum alloy die casting forming process, and relates to the technical field of intelligent optimizing.The intelligent optimization method comprises the steps that the multi-element content in an aluminum alloy melt is collected in real time, and the multi-element content is combined into a three-dimensional material gene vector; inputting the three-dimensional material gene vector into a pre-trained graph neural network, and calculating to obtain an injection speed compensation coefficient and a boost pressure compensation coefficient; performing space-time alignment on the geometric center coordinate of the defect and die temperature field and pressure time sequence data recorded in the die casting process, inputting a reinforcement learning agent model constructed through a depth Q network algorithm, and reconstructing an evolution path of an internal defect area; and according to the defect formation time point and position point identified in the evolution path, reversely correcting the weight parameter of the graph neural network, and updating. According to the method, the content of multiple elements is fused into the three-dimensional material gene vector, and millisecond-level cooperative compensation of the injection speed and the boost pressure is achieved through the dynamic topology modeling capacity of the pre-training graph neural network.
Owner:EDT DIECASTING TECH SUZHOU CO LTD

Partitioned rapid inversion method for global structural mechanical parameters of concrete arch dam

The invention discloses a concrete arch dam global structural mechanical parameter zoning rapid inversion method, and relates to the technical field of dam operation safety monitoring and management, and the method comprises the steps: building a dam body and foundation three-dimensional finite element model through finite element software according to engineering design and monitoring data, and building a foundation three-dimensional finite element model according to damming material mechanical parameter information; partitioning the three-dimensional finite element model, analyzing the sensitivity of mechanical parameters of each region, determining sensitive mechanical parameters influencing the deformation of the concrete arch dam, and carrying out self-adaptive intelligent sampling on the sensitive mechanical parameters according to a sensitivity analysis result, and constructing a deep learning agent model reflecting a nonlinear relationship between the sensitive mechanical parameters of the dam and the deformation of each monitoring point, and carrying out deep learning inversion on the elastic modulus of the dam body and the deformation modulus of the bedrock. According to the invention, the method can efficiently and accurately invert and determine the structural mechanical parameters of the arch dam in the actual operation period, and provides a good basis for the safety analysis of the dam.
Owner:NANCHANG UNIV

Cross-border e-commerce advertisement accurate putting method and system based on user portraits

The invention relates to the technical field of cross-border advertisement putting, and provides a cross-border e-commerce advertisement accurate putting method and system based on a user portrait, and the method comprises the steps: building a unified data flow through multi-source data collection and standardization processing; calculating and fusing the real-time statistical features, the historical interest vectors and the context features in real time by using a stream processing engine to generate multi-dimensional feature vectors; inputting the feature vector into a multi-modal time sequence neural network model which captures a user behavior rule through time sequence coding and an attention mechanism and outputs a user behavior weight and an interest attenuation coefficient; dynamically updating the real-time interest score of the user by adopting an exponential decay model; and when the interest score reaches a threshold value, triggering a reinforcement learning agent to make a decision according to the comprehensive state information, and dynamically outputting an action strategy including release triggering, a bidding coefficient and a creative type. The problems that in the prior art, user interest modeling lags behind, and the self-adaptive capacity of the putting strategy is poor are effectively solved.
Owner:NANJING YUSITUOMENG INTERNATIONAL TRADING CO LTD

Neural network automatic pruning method based on GRPO reinforcement learning

The invention belongs to the technical field of artificial intelligence, and particularly relates to a neural network automatic pruning method based on GRPO reinforcement learning, and the method comprises the steps: introducing a dynamic scaling factor into a batch normalization layer of a to-be-pruned neural network, and calculating the importance score of each convolution layer channel of the to-be-pruned neural network in combination with an attention mechanism; s2, constructing a multi-dimensional state vector containing layer structure features based on an importance calculation result in the step S1; s2, inputting the multi-dimensional state vector constructed in S2 into a strategy network of a GRPO reinforcement learning agent, generating a pruning action by the strategy network according to state information of a current network layer, and defining the action to represent a pruning rate of the layer; according to the method, a GRPO reinforcement learning algorithm is adopted, a traditional Critic model is abandoned, the strategy calculation process is simplified through a group sampling-relative advantage estimation mechanism, and memory occupation is remarkably reduced.
Owner:SHANDONG UNIV

Artificial intelligence (AI) multi-agent framework

According an embodiment of the present invention, a system processes requests to perform projects. The system comprises one or more memories, and at least one processor coupled to the one or more memories. The at least one processor processes a request in a natural language to perform a project via a hierarchy of machine learning agents. One or more machine learning agents of a management layer of the hierarchy determine and assign tasks for the project to one or more machine learning agents of an operation layer of the hierarchy based on the request. The at least one processor performs the assigned tasks by the one or more machine learning agents of the operation layer to perform the project. Embodiments of the present invention further include a method and computer program product for processing requests to perform projects in substantially the same manner described above.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Machine-generated training examples for training machine-learned models

An example method includes obtaining a reference trace describing interactions between a machine-learned agent system and an environment, wherein the reference trace includes a reference sequence of interaction objects associated with performance of a task, each respective reference interaction object of the reference sequence of interaction objects corresponding to a respective reference environment state and a respective reference action executed on the respective reference environment state. The example method includes generating, using a machine-learned trace generation system, a forecasted trace based on the reference trace, the forecasted trace comprising a forecasted sequence of interaction objects that is predicted to continue the reference sequence of interaction objects from a branching position toward performance of the task, each respective forecasted interaction object of the forecasted sequence of interaction objects corresponding to a respective forecasted environment state and a respective forecasted action executed on the respective forecasted environment state. The example method includes training the machine-learned agent system using the forecasted trace sequence.
Owner:GOOGLE LLC

Bridge group multi-target maintenance decision-making method fusing evolutionary algorithm and artificial intelligence

The invention provides a bridge group multi-target maintenance decision-making method fusing an evolutionary algorithm and artificial intelligence, and relates to the technical field of civil engineering and artificial intelligence crossing. The method comprises the steps of defining bridge group maintenance cost and structure failure risks, representing preferences of decision makers for different decision targets by weight combinations, and constructing a multi-target maintenance decision optimization model; encoding the weight combination into an individual of a multi-objective evolutionary algorithm, and randomly generating an initial population; aiming at each generation of weight combination, constructing a bridge group Markov decision-making environment; learning an optimal maintenance strategy by adopting an A2C training reinforcement learning agent; the optimal maintenance strategy is evaluated, and an evaluation result is used as individual fitness to be fed back to the multi-objective evolutionary algorithm; using a multi-objective evolutionary algorithm to perform evolutionary search on the multi-objective weight combination; and through a closed-loop feedback mechanism, outputting a Pareto optimal solution set containing an optimal maintenance strategy under various weight combinations, thereby realizing collaborative optimization of weight optimization and strategy learning.
Owner:UNIV OF SCI & TECH BEIJING

Deep learning assisted acceleration fracturing construction parameter intelligent real-time optimization method

The invention discloses a deep learning assisted acceleration fracturing construction parameter intelligent real-time optimization method, and belongs to the field of unconventional oil and gas field development, and the method comprises the steps: constructing a structured sample database covering geological attributes, engineering parameters and fracturing effect responses, and carrying out the unified preprocessing of multi-source data; constructing a multi-modal space-time collaborative prediction model, and combining a multi-expansion time permissible convolutional network, a three-dimensional residual network and a cross attention mechanism; supervised learning training is carried out based on a database, sparse regularization and learning rate scheduling are introduced, and the model precision and generalization ability are improved; a hierarchical collaborative optimization framework of outer-layer Bayesian search-inner-layer CMA-ES refinement is provided, and reservoir transformation volume maximization and construction feasibility are achieved under complex constraints; compared with an existing method, optimization parameters of the method include the cluster distance and the section distance and further include the displacement, the sand concentration, the fracturing fluid type and other construction parameters, and real-time optimization of the fracturing construction parameters can be achieved through the machine learning agent model.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

Dynamic computing power distribution method and system based on reinforcement learning

The invention belongs to the technical field of computing power distribution, and particularly relates to a dynamic computing power distribution method and system based on reinforcement learning, and the method comprises the following specific steps: S1, covering cloud, edge and end full-node scenes, and collecting computing power resource states, task demand features and cross-domain network condition data in real time; s2, on the basis of standardized data output by a cross-domain computing power sensing module, by constructing a state space fusing computing power, tasks and a network, defining an action space of computing power scheduling direction and proportion, and designing a multi-target reward function for balancing the resource utilization rate, the task satisfaction rate and long-term conflict avoidance; and realizing self-learning and self-iteration scheduling strategy generation based on a reinforcement learning algorithm. According to the invention, the reinforcement learning agent autonomously learns the computing power demand of the emergency scene and the new type of task, the rule does not need to be manually preset and modified, and the method has the advantage of realizing dynamic adaptation of computing power distribution to complex and changeable scenes.
Owner:BEIJING CENTURY FEIXUN TECH CO LTD

Multi-parameter and multi-field intelligent optimization method and system for press free forging large-scale die casting

The invention relates to a multi-parameter and multi-field intelligent optimization method and system for a press free forging large-scale falling die, and the method achieves the collaborative optimization of technological parameters, die geometry and microstructure through a heat-force-microstructure three-field coupling modeling, a deep kernel learning agent model and a digital twinning guided multi-objective evolutionary algorithm, improves the optimization efficiency, and improves the optimization precision. The mold testing times are reduced; and meanwhile, a real-time closed-loop control system is constructed, the grain size uniformity of forgings is improved, the forming load is reduced, the production period is shortened, the process stability is improved, and the die service life is prolonged.
Owner:ZHEJIANG JIEDE MASCH TECH CO LTD

Training reinforcement learning agents to learn farsighted behaviors by predicting in latent space

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training an action selection policy neural network used to select an action to be performed by an agent interacting with an environment. In one aspect, a method includes: receiving a latent representation characterizing a current state of the environment; generating a trajectory of latent representations that starts with the received latent representation; for each latent representation in the trajectory: determining a predicted reward; and processing the state latent representation using a value neural network to generate a predicted state value; determining a corresponding target state value for each latent representation in the trajectory; determining, based on the target state values, an update to the current values of the policy neural network parameters; and determining an update to the current values of the value neural network parameters.
Owner:GOOGLE LLC

Intelligent decision-making system and method for corn fertilization based on mechanism-data dual-drive fusion

The invention belongs to the technical field of unmanned aerial vehicle remote sensing and agriculture combination, and discloses an intelligent decision-making system and method for corn fertilization based on mechanism-data dual-drive fusion. The system comprises a mechanism simulation module, a data preprocessing module, a feature engineering module, a modeling module, a visualization module, a decision support module, a dynamic feedback correction module and a report generation and push module. According to the mechanism-data double-drive fusion-based intelligent decision-making system and method for corn fertilization, a large-area corn field block image is obtained in a short time through an unmanned aerial vehicle multispectral system, and the nitrogen diagnosis efficiency is improved; through a lightweight machine learning agent model, corn canopy leaf nitrogen nutrition parameters are accurately predicted, and a basis is provided for accurate fertilization; the system monitors the nitrogen nutrition status of the corn in real time, and provides possibility for dynamically adjusting a fertilization strategy; the multispectral remote sensing technology can perform nitrogen nutrition diagnosis under the condition of not damaging corn plants, and the corn growth environment is protected.
Owner:AGRI SCI RES INST OF THE SEVENTH DIVISION OF XINJIANG PROD & CONSTR CORPS

Mobile phone thermal simulation method and system based on multi-physics field coupling

The invention discloses a mobile phone thermal simulation method and system based on multi-physics field coupling, relates to the field of mobile phone thermal simulation, and constructs hybrid simulation coupling machine learning and physical solution. A machine learning agent model and a throttling logic script are innovatively introduced. The machine learning agent model is used for quickly predicting the temperature according to the current power consumption so as to instantaneously respond to the power consumption change; and the throttling logic script is used for simulating a real temperature control frequency reduction strategy and dynamically adjusting the target power consumption at the next moment according to the predicted temperature. In this way, the original black box temperature control logic is explicit, and a power consumption-temperature closed-loop feedback path is constructed. And finally, taking the adjusted power consumption as a heat source, and performing physical field solving by a traditional CAE solver. Therefore, the problem of time scale difference between heat conduction and electrical change is effectively solved, and efficient and accurate simulation of true performance and temperature performance of the mobile phone in a long-time high-load scene is realized.
Owner:SHENZHEN DUOKE ELECTRONICS CO LTD