Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

13results about How to "Improve exploration ability" patented technology

A Smart Path Planning Method for Heavy-Duty Robotic Arms for Retired Photovoltaic Modules

PendingCN122077647AImprove exploration abilityImprove the randomness of actionsProgramme-controlled manipulatorRobotic armControl engineering
This invention discloses an intelligent path planning method for heavy-duty robotic arms used in the dismantling of decommissioned photovoltaic (PV) modules. The method involves modeling and training the dismantling task of decommissioned PV modules using an actor-critic algorithm, followed by intelligent path planning for the heavy-duty robotic arm based on the constructed model, generating locally optimal work trajectories. Heavy-duty robotic arms face scenarios with few obstacles in the dismantling of decommissioned PV modules, and their path planning problem can be reduced to an optimization task for a single robotic arm. This solution employs an actor-critic algorithm to achieve efficient and stable path generation, supporting automated dismantling. Furthermore, it introduces the concept of entropy to enhance exploratory capabilities, avoid local optima, and ensure the robotic arm can adapt to the irregular shapes of the modules and dynamic environmental changes.
Owner:NANJING UNIV OF SCI & TECH +1

Trajectory coverage sensor placement method based on hierarchical deep reinforcement learning

PendingCN122263193AImprove stabilityReduce decision-making dimensionsGeometric CADBiological modelsAlgorithmEngineering
The application discloses a trajectory coverage sensor layout method based on layered deep reinforcement learning, comprising the following steps: constructing a trajectory dataset of a target area and a sensor deployment constraint model, including a deployable area, a minimum spacing constraint and a coverage criterion; decomposing a global layout task into two levels of hotspot center selection and local offset optimization: in the high-level strategy, the global trajectory density is analyzed based on environment observation, and a soft actor-critic algorithm is used to generate hotspot center instructions; in the low-level strategy, a continuous action optimization method with a mixed density network and a Gumbel-Softmax selector is used to solve the multi-modal of the sensor offset in each hotspot. The application introduces a random network distillation module at the high and low levels, and through a cross-layer reward injection mechanism, the high-level strategy can be adaptively updated according to the actual coverage gain of the low level, so as to improve the coordination and convergence efficiency of the overall layout.
Owner:SOUTH CHINA UNIV OF TECH

A mechanical arm motion optimization method and system based on an improved sparrow algorithm

ActiveCN118305805BGuaranteed uptimeOptimize runtime
The application provides a kind of motion optimization method and system of mechanical arm based on improved sparrow algorithm, based on the kinematic model of mechanical arm, the trajectory of mechanical arm is planned using polynomial interpolation method, and the joint motion trajectory is obtained;Based on the joint trajectory, the objective function and the constraint condition of the mechanical arm are constructed;The sine-cosine mapping, improved elite reverse learning strategy and Cauchy Gaussian variation principle are integrated into the sparrow algorithm to obtain the improved sparrow algorithm;Based on the objective function and the constraint condition, the improved sparrow algorithm is applied to the trajectory planning of the mechanical arm, and the motion of the mechanical arm is optimized.The application integrates the sine-cosine mapping, improved elite reverse learning strategy and Cauchy Gaussian variation strategy into the sparrow algorithm to form the improved sparrow algorithm, improves the diversity of population, avoids falling into local optimum, and can effectively find the global optimal solution.The improved sparrow algorithm is applied to the mechanical arm, which improves the running efficiency and running accuracy of the mechanical arm.
Owner:XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY

An entropy-regularized driving method for cooperative non-cooperative target capture of aircraft cluster

ActiveCN121050458BImprove exploration abilityImprove robustnessLocal optimumTarget capture
The application discloses an aircraft cluster cooperative non-cooperative target capturing method driven by entropy regularization, applies the entropy regularization thought to reinforcement learning, optimizes the updating method of the parameters of an evaluation network and a policy network, so that a UAV can obtain a more optimal maneuvering strategy, and the UAV performs actions according to the maneuvering strategy, and realizes the interception of an attacking aircraft. The aircraft cluster cooperative non-cooperative target capturing method driven by entropy regularization can not rely on a deterministic strategy in the interception process of multiple UAVs, but adopts a random strategy, avoids the training from falling into a local optimal point, and the actions of the aircraft are randomized as much as possible while the aircraft completes the task, so that the exploratory and robustness can be practically improved.
Owner:BEIJING INST OF TECH

Cloud large model and marginal small model collaborative terminal device with body and control method thereof

The invention discloses a cloud large model and marginal small model collaborative terminal device with a body and a control method of the cloud large model and marginal small model collaborative terminal device. The body terminal device comprises a cloud large model intelligent training module which is used for training and optimizing a TD3 strategy model through an immune algorithm so as to improve the exploration capability and the control precision. The edge reasoning and control module is used for deploying a lightweight TD3 model, sensing the environment in real time under the conditions of low power consumption and low delay and outputting a control action; the communication and cooperation module is used for synchronizing data and control strategies between the cloud and the edge; and the execution and feedback module is used for converting the control action into actual flight operation and returning information such as position, speed and attitude so as to support continuous updating and control.
Owner:CHINA TELECOM UNMANNED TECH (JIANGSU) CO LTD

Reinforcement learning simulation step control method based on closed-loop adaptive noise injection and entropy increasing optimization strategy

ActiveCN120255376BImprove exploration abilityExplore efficiencySimulator controlTransient analysisAlgorithm
This invention discloses a reinforcement learning simulation step size control method based on closed-loop adaptive noise injection and entropy increase optimization strategy, belonging to integrated circuit computer-aided design technology. Specifically, the method involves: first, inputting the circuit netlist and interacting with the simulator using file read / write; next, establishing two new network output layers to convert the deterministic actions of the policy network's output into a probability distribution; then, obtaining the entropy regularization term based on the probability density function and weight coefficients, and adding it to the policy loss function; next, generating Gaussian-distributed exploration noise and adjusting the standard deviation of the noise at each step using a PID controller; finally, adjusting the policy network parameters using a gradient update method and injecting the exploration noise into the output actions to obtain the final time step. Using this invention helps enhance the simulation stability of pseudo-transient analysis, improves simulation efficiency, and provides a new method for DC analysis.
Owner:SOUTHEAST UNIV

Serverless mapreduce multi-objective scheduling optimization method

The present application relates to cloud computing and distributed computing scheduling technical field, specifically belong to a kind of serverless MapReduce multi-objective scheduling optimization method, including: constructing directed acyclic graph, resource allocation decision is encoded as the solution from source node to target node, initial population is generated using hybrid initialization strategy, fitness evaluation, the crowded distance of solution in each non-dominated front is calculated, in turn execute tournament selection, adaptive crossover strategy, topological repair mutation strategy, periodically execute heuristic search guided solution injection strategy, crowded distance selection is carried out, the global Pareto optimal solution set is updated, the global Pareto optimal solution set of current time output is achieved, otherwise return iteration. The present application has the positive effect of balancing the execution cost and execution time of MapReduce job, improving the global Pareto optimal solution set quality of multi-objective scheduling scheme, enhancing the convergence speed and robustness of algorithm.
Owner:LIAOCHENG UNIV

Agent-based industrial malt manufacturing system and automatically optimizing operation of industrial malt manufacturing equipment and method thereof

An automated industrial malt manufacturing device (11) and an agent- and digital twin-based malt manufacturing system (1) are proposed for the automated operation of the industrial malt manufacturing device (11) of the agent- and digital twin-based malt manufacturing system (1). The malt manufacturing device (11) converts supplied grains (2, 21, 21i) into output malt (3). The malt manufacturing device (11) includes: a soaking device (111) for wetting the supplied grains (2, 21, 21i) until the moisture content of the supplied grains (2, 21, 21i) reaches a certain level (111). 1) Reaching a defined moisture content level (1112); a germination device (112) configured to receive moistened grains (22) after being moistened by a soaking device (111), for germinating the moistened grains (22) by activating enzymes in the moistened grains (22) and converting the moistened grains (22) into germinated green malt (23); and a drying kiln (113) configured to receive germinated grains (23) after being germinated by the germination device (112), for drying the germinated green malt (23) to a defined moisture content level (1131). The system (1) based on intelligent agents and digital twins includes a digital twin structure (12) providing a virtual digital representation of the physical malt manufacturing equipment (11). In addition, it includes sensors and / or measuring devices (13) associated with the physical malt manufacturing equipment (11), which transmit measured sensory parameter values ​​(131), wherein the system (1) dynamically monitors the transmitted sensory parameter values ​​(131) and updates the digital twin (12) based on the parameter values.
Owner:BUHLER AG

Multi-machine multi-missile collaborative air combat decision-making method based on SM-HAPPO deep reinforcement learning

The invention discloses a multi-machine multi-missile collaborative air combat decision-making method based on SM-HAPPO deep reinforcement learning, and relates to the technical field of air combat intelligent decision-making. Starting from a multi-aircraft multi-missile collaborative air combat decision-making method, a dynamic model of an aircraft, a missile guidance model and an air combat situation model are established, and simulation of a complex environment is achieved. For collaborative decision-making of isomorphic or heterogeneous aircrafts, a PPO algorithm is expanded to the field of heterogeneous agents, a GRU-based state memory module is introduced, and an SM-HAPPO algorithm suitable for deep reinforcement learning of the heterogeneous agents is provided. In the strategy network, a loss function with a mask mechanism and a strategy entropy regular term is introduced, and the stability and exploratory performance of the algorithm are improved. Finally, a simulation experiment of aircraft collaborative air combat decision making is completed in a python environment, and it is proved that compared with a traditional intelligent algorithm, the collaborative decision making capacity and the air combat winning rate of the unmanned aerial vehicle can be further improved.
Owner:XIAMEN UNIV

Gain updating method and device of virtual resources, storage medium and computer equipment

According to the gain updating method and device of the virtual resources, the storage medium and the computer equipment provided by the invention, the first virtual role is determined from the plurality of virtual roles which are randomly generated by the system and do not belong to the user in response to the triggering of the role activation instruction by the user, and the production attribute of the first virtual role is changed into the target attribute with the gain endowing permission; associated first gain parameter values and target production units are then determined to update the overall gain efficiency. According to the method, the traditional single mode of performing gain configuration depending on existing roles of players is broken through, the first virtual role randomly generated by the system provides an additional gain source for the players, the players can obtain gains without accumulating enough roles, and the operation efficiency and the plot propelling efficiency of the players are effectively improved. Meanwhile, due to the randomness of the first virtual character, the gain parameter value of the first virtual character and random or dynamic generation of the target production unit, the interactivity and interestingness of the game are remarkably improved, and the experience feeling of players is improved.
Owner:GUANGZHOU YIWAN NETWORK TECH CO LTD

Offshore wind power structure finite element model updating method and system based on sampling

The invention discloses an offshore wind power structure finite element model updating method and system based on sampling, and belongs to the technical field of wind power generation. The system comprises a sensing monitoring module and an updating module. The sensing monitoring module is used for collecting data of a bearing part of the offshore wind power structure represented by the finite element model to be updated; the updating module is used for acquiring various modal parameters for updating the finite element model based on the data acquired by the sensing monitoring module; the updating module carries a to-be-updated finite element model and processes the modal parameters and preset prior distribution of a plurality of to-be-updated parameters of the finite element model based on the finite element model and an enhanced transitional Markov chain Monte Carlo method to obtain posterior distribution of the to-be-updated parameters; and obtaining a posteriori mean value of each to-be-updated parameter based on the posteriori distribution, and updating the corresponding parameter in the finite element model to complete updating of the finite element model. The method provides a data basis for reliability analysis of a wind power structure.
Owner:ZHEJIANG UNIV

Inference method, system, device and medium based on heterogeneous model reinforcement learning

This application discloses a reasoning method, system, device, and medium based on heterogeneous model reinforcement learning, belonging to the field of artificial intelligence technology. This method generates structured tree-like reasoning knowledge at low cost using a Monte Carlo tree search mechanism, introduces computational value scoring to balance accuracy and computational overhead, and applies the high-order reasoning patterns obtained from the tree-like reasoning knowledge to the reinforcement learning training process. This effectively internalizes the high-order knowledge of the heterogeneous model group into the first policy model, thereby solving the dual technical challenges of high cost in constructing high-quality distilled data and the lack of multi-source external knowledge guidance in traditional reinforcement learning algorithms. Furthermore, it enables the first policy model to reason about answers to natural language tasks through a tree-like exploratory reasoning process during the reasoning phase, significantly improving the generalization ability, exploration ability, and decision confidence of the first policy model, thus effectively enhancing its reasoning ability.
Owner:TSINGHUA UNIVERSITY

A method for visual text search based on hybrid data model

This invention relates to the field of information retrieval and data visualization technology, and discloses a visual text search method based on a hybrid data model. The method includes: constructing a hypergraph hybrid model by associating a graph database and a vector database using global identifiers; identifying the intent of user queries; generating structured search instructions for the graph database and semantic similarity search instructions for the vector database; fusing and expanding the two types of search results based on global identifiers to form a search result set containing an extended knowledge graph, a list of semantic vectors, and traceable links between them; and finally, converting the result set into a dynamic and interactive visual view according to predefined mapping rules and a hybrid mechanics layout algorithm. This invention effectively unifies semantic retrieval and relational retrieval, improves the comprehensiveness and accuracy of retrieval, and significantly enhances the depth of information understanding and user experience through intelligent visualization and interaction mechanisms.
Owner:YIMING UNIVERSAL TECH (BEIJING) CO LTD