Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

227 results about "Parallel optimization" patented technology

Regional energy internet hierarchical control method considering flexible load

The invention discloses a regional energy internet hierarchical control method considering flexible loads, which comprises the following steps of: constructing a global layer, dividing a regional energy internet into a plurality of electrically strongly coupled sub-regions by adopting a dynamic partitioning algorithm, aggregating equivalent adjustable capacities of the flexible loads in the sub-regions, and performing rolling optimization in an hour-level time scale, an improved Benders decomposition algorithm is adopted to accelerate solving and generate a partition-level power regulation instruction; constructing a regional layer, designing a thermodynamic-economic hybrid constraint model according to flexible load physical characteristics and user behavior elasticity, introducing a user comfort recovery time constraint and elastic electricity price response function, and decomposing an adjustment instruction issued by the global layer to each flexible load cluster in a minute-level time scale by adopting a parallel ADMM optimization algorithm; and constructing a load layer, constructing a millisecond-level adjustable load priority switching mechanism based on FPGA hardware acceleration and finite-state machine logic, calculating a load adjustable power boundary in real time, and dynamically adjusting a response queue in combination with system frequency deviation.
Owner:STATE GRID CORPORATION OF CHINA +2

Multivariable energy efficiency optimization control system for heating furnace

The invention relates to the technical field of control, and particularly discloses a multivariable energy efficiency optimization control system for a heating furnace, which is used for solving the problems of local overheating, non-uniform temperature and difficulty in accurate positioning and compensation of heat loss in the operation of the existing cracking heating furnace. Comprising a parameter detection module, a multivariable coupling modeling and simulation module, an optimization control module and an execution and feedback module. According to the method, dynamic digital twinning is constructed through multi-modal online sensing and data assimilation, a Pareto frontier solution is generated based on model prediction control and improved NSGA-II parallel optimization, and the weight is adaptively adjusted; and when the hot spot / cold spot is triggered, a quadric surface fitting compensation strategy is implemented and issued for execution, so that high-precision simulation prediction, precise closed-loop control and real-time online energy efficiency optimization are realized.
Owner:ANHUI ZHONGKE WEIDE DIGITAL TECH CO LTD +1

Satellite networking communication method, device and equipment

The invention provides a satellite networking communication method, device and equipment, and the method comprises the steps: carrying out the space-time window decomposition processing of a satellite constellation orbit parameter based on orbit phase synchronization, and decomposing a long-period scheduling problem into a plurality of short-period sub-problems; performing state transition conflict detection processing on the double-feed antenna system according to the space-time window decomposition result, and determining the working state of the feed antenna; parallel neighborhood search processing based on reinforcement learning is carried out according to the feed antenna resource configuration scheme, and a communication link scheduling strategy is generated through multi-thread collaborative optimization; and performing dynamic fusion processing based on state prediction according to the parallel optimization scheduling scheme, and adjusting a satellite-ground station communication link establishment time sequence through learning type template matching. Through space-time decomposition, intelligent conflict detection, reinforcement learning optimization and dynamic fusion, the problems of complexity and real-time performance of large-scale satellite networking communication scheduling are solved, and the scheduling efficiency and the communication quality are remarkably improved.
Owner:SHEN ZHEN MORNSUN ELECTRONICS CO LTD

Graphene film surface performance detection device

The invention relates to the technical field of electrical property detection, and discloses a graphene film surface property detection device, which comprises a data acquisition module used for acquiring electrical property data of the surface of a graphene film through a multi-point array electrode probe; the vector grid modeling module is used for constructing a vector grid model for representing the surface performance of the graphene film according to the electrical characteristic data; the iterative inversion calculation module is used for performing iterative inversion calculation on the vector grid model to determine the surface performance distribution of the graphene film; the parallel optimization fusion module is used for performing parallel processing and optimization fusion on the multi-scale performance data through a distributed computing framework; the three-dimensional visualization module is used for generating a three-dimensional visualization model of the surface performance of the graphene film according to the optimized and fused performance data; the problems of insufficient precision and low efficiency of a traditional detection method in graphene film performance detection are solved.
Owner:SHENZHEN THIN CONDUCTOR TECH CO LTD

Large model and multi-agent collaborative decision-making method based on dynamic knowledge flow

The invention relates to the technical field of artificial intelligence, in particular to a large model and multi-agent collaborative decision-making method based on dynamic knowledge flow, which comprises the following steps of: analyzing a static knowledge and dynamic information fusion relationship through joint modeling, extracting a hierarchical structure of equipment constraints and environment variables, identifying constraint conflicts and deviations in task decomposition, and obtaining a multi-agent collaborative decision-making result; screening a consistency decomposition direction, extracting a task constraint parallel optimization theory, correcting constraint conflicts, evaluating consistency changes, and outputting a collaborative task convergence robust state identifier. According to the method, by integrating static knowledge and dynamic information and optimizing understanding of task decomposition and equipment constraints, the collaborative decision-making capacity of multiple agents in a complex environment is enhanced, the conflict and deviation processing capacity in the task execution process is improved, the accuracy and consistency of tasks are enhanced, and the convergence and stability of task targets are improved; the robustness of multi-agent cooperative work is promoted, and finally more efficient resource utilization and task completion effects are achieved.
Owner:JINJIELI TECH (BEIJING) CO LTD

Parallel optimization method and system for GPGPU (General Purpose Graphics Processing Unit) instruction execution stage

The invention provides a parallel optimization method and system for a GPGPU (General Purpose Graphics Processing Unit) instruction execution stage, and the method comprises the steps: splitting an input instruction into a plurality of sub-instructions which can be independently scheduled based on an instruction type, operand dependency analysis and a hardware resource real-time state, and generating a dependency table to record the input and output association of the sub-instructions; acquiring real-time load data of the calculation unit, the storage unit and the communication unit, and distributing the sub-instructions to the calculation unit, the storage unit or the communication unit by adopting a priority dynamic adjustment algorithm and combining instruction key path analysis; according to the dependency relationship of the sub-instructions and the availability of hardware resources, the execution sequence of the sub-instructions is adjusted to maximize the degree of parallelism; and analyzing the dependency relationship among the sub-instructions in real time to maximize the degree of parallelism to dynamically adjust a parallel execution strategy so as to realize serial execution of dependent instructions and parallel execution of non-dependent instructions. According to the method, the complexity of instruction compiling is reduced, the efficiency of instruction compiling is improved, and meanwhile extra consumed time is reduced.
Owner:HEFEI SUMICROELECTRONICS TECH CO LTD

Lightweight structure multi-scale parallel optimization method based on lattice discrete optimization

The invention belongs to the technical field of additive manufacturing, and particularly relates to a multi-scale parallel optimization method for a lightweight structure based on lattice discrete optimization. According to the method, through multi-scale parallel optimization design, a macro structure and a micro lattice structure are described through geometric parameters of a movable deformation rod piece, an improved two-value coding parameterization method, called a BCP method for short, is adopted to solve discrete optimization of the micro lattice structure, and geometric parameters of an optimization result are easy to extract.
Owner:CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI

Cholesky decomposition heterogeneous parallel optimization method and system based on SW architecture

The invention provides a Cholesky decomposition heterogeneous parallel optimization method and system based on an SW architecture, and relates to the technical field of high-performance computing. The method comprises the following steps: performing sub-block division on a symmetric positive definite matrix based on a distributed parallel distribution scheme, and performing iteration to complete matrix decomposition; each sub-block is distributed to different processes through an MPI programming model, data exchange is carried out between the processes through asynchronous communication, and coarse-grained task-level parallel acceleration is carried out; performing two-stage parallel acceleration on four operations in Cholesky decomposition by utilizing the acceleration parallel characteristic of a master core and a slave core of the SW architecture; wherein for GEMM and SYRK operations, column vectors of a matrix are mapped to a slave core array, and columns are divided according to the number of slave cores; the calculation process is optimized through a double-buffering mechanism, vectorization operation and a loop expansion technology, and the parallel efficiency is improved; and for the TRSM operation, the TRSM operation is decomposed into a plurality of TRSV operations, the TRSV operations are allocated to the slave cores for parallel execution, and a circular reading and data broadcasting mode is adopted to reduce data dependence and realize efficient parallel calculation.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Power distribution network photovoltaic openable capacity dynamic evaluation method fusing voltage stability margin and neural network optimization

The invention discloses a power distribution network photovoltaic openable capacity dynamic evaluation method fusing voltage stability margin and neural network optimization, and aims to solve the problems that a traditional method does not fully consider dynamic stability and is low in calculation efficiency. Firstly, a prediction method combining kernel density estimation and quantile regression is adopted to accurately quantify the uncertainty of distributed photovoltaic output. The core innovation of the method is that on the basis of traditional security constraints, a static voltage stability margin (VDSM) is introduced as a key constraint condition, and a double-layer interval analysis model capable of guaranteeing the dynamic stability of a power grid is constructed. Secondly, in order to efficiently solve, the invention provides a framework of'neural network pre-screening + parallel optimization ': after a model is decomposed into optimistic sub-problems and pessimistic sub-problems through an interval decoupling technology, massive candidate solutions are quickly screened by utilizing a neural network model, so that a feasible solution space is greatly reduced, and the solution efficiency is improved; and carrying out parallel optimization solution on the sub-models in combination with an improved particle swarm optimization algorithm.
Owner:INNER MONGOLIA POWER (GRP) CO LTD XUEJIAWAN POWER SUPPLY BUREAU

Joint optimization distributed training method based on heterogeneous hybrid parallel and communication

The invention relates to the technical field of model training, and particularly provides a joint optimization distributed training method based on heterogeneous hybrid parallel and communication. The method comprises the following steps: constructing a heterogeneous parallel optimization strategy through a heterogeneous parallel optimization algorithm HPO; constructing a dynamic dual-heterogeneous synchronous communication mechanism through a dynamic dual-heterogeneous synchronous strategy DHS; according to the heterogeneous parallel optimization strategy and the dynamic dual-heterogeneous synchronous communication mechanism, a hybrid parallel and communication distributed training framework in the dual-heterogeneous environment is constructed to perform joint optimization distributed training, the method optimizes computing resource scheduling, reasonably and dynamically solves the problem of communication delay, and the efficiency of distributed training is improved.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Large model prompt word engineering method and system

The present invention discloses a large model prompt word engineering method and system, the method comprising: receiving a task requirement input by a user for interacting with a large model, generating a set of initial prompt word sets based on the task requirement; performing feature mapping on the initial prompt word set, constructing a prompt word feature vector set, decomposing the feature vectors in the prompt word feature vector set into a number of sub-feature vectors to capture the different characteristics of the prompt words; performing multi-objective optimization on the sub-feature vectors using a parallel optimization algorithm to generate an optimized prompt word set; screening the optimized prompt word set for semantic consistency and generation diversity, and determining a number of final prompt words as the large model input corresponding to the task requirement. By using the embodiments of the present invention, it is possible to realize the intelligent generation and optimization of the prompt word set according to the user's task requirement, improve the efficiency and quality of prompt word construction, and provide more flexible support for the practical application of the large model.
Owner:GUANGZHOU ZHONGCHANG KANGDA INFORMATION TECH

CNN-oriented batch matrix multiplication parallel optimization method and system on SW architecture

The invention provides a CNN-oriented batch matrix multiplication parallel optimization method and system on a SW architecture, and belongs to the technical field of artificial intelligence parallel optimization. Comprising the following steps: respectively converting an input feature map and a convolution kernel in a convolution layer into an input matrix and a weight matrix, and processing the input matrix and the weight matrix into a plurality of groups of independent matrix multiplication tasks in batches; the main core encapsulates a matrix multiplication task into a parameter structure array, the parameter structure array is transmitted to the slave core through single DMA, and the slave core divides rows of an input matrix into row block tasks by adopting a dynamic row block division algorithm according to the total number of threads and the height of the matrix; and executing sub-matrix multiplication calculation on the distributed independent row blocks, asynchronously prefetching matrix sub-blocks by adopting a double-buffer DMA (Direct Memory Access), and executing matrix multiply-accumulate calculation. The parallel processing efficiency of batch matrix multiplication between the master core and the slave core of the SW processor can be improved, and the algorithm performance is optimized.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Photovoltaic optimization regulation and control method and equipment based on MPPT (Maximum Power Point Tracking) technology

The invention discloses a photovoltaic optimization regulation and control method and device based on an MPPT technology, and particularly relates to the technical field of photovoltaic power generation control, and the method specifically comprises the following steps: S1, full-link operation state synchronous collection, S2, DC bus ripple feature extraction, S3, multi-source disturbance quantitative evaluation, S4, cooperative control parameter decision making, and S5, grid-connected signal synthesis modulation. A multi-source disturbance quantitative evaluation system is adopted, the irradiance change gradient, the frequency ratio and the switching noise contribution ratio are subjected to multi-dimensional fusion, a dynamic weight distribution mechanism is established, and parallel optimization of MPPT mode selection, disturbance step length adjustment and PWM phase shift angle is achieved through a cooperative control parameter decision architecture. A response delay bottleneck existing in a traditional serial decision mode is broken through, and a ripple compensation component is dynamically injected while stable output of fundamental current is maintained based on a grid-connected signal modulation strategy of vector synthesis.
Owner:叶春

Non-uniform grid data calculation method based on hierarchical calling rule and related device

The invention discloses a non-uniform grid data calculation method based on a hierarchy calling rule and a related device, and the method comprises the steps: firstly analyzing an input parameter of a calculation task, and selecting corresponding calculation models according to a parameter type, the calculation models comprise a C-type model, a B-type model and an A-type model, and the models are divided according to a dependency relationship; thirdly, constructing a directed acyclic graph according to a hierarchical calling rule, representing a dependency relationship between tasks and excluding cyclic calling; then, a task execution sequence is generated based on the graph, and is ranked according to a hierarchical call rule. And performing parallel optimization on the task execution sequence according to the computing resource state, and allocating the task execution sequence to a distributed computing engine for execution. And finally, storing the result in a database, and supporting a subsequent task to directly call a historical result. The method aims at solving the problems of disordered dependency, high redundancy, low efficiency and the like of calculation tasks, and the organization efficiency and performance of building energy-saving calculation are improved.
Owner:XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY

Multi-scale parallel optimization method for lightweight structure based on lattice discrete optimization

The present application belongs to the technical field of additive manufacturing, and particularly relates to a lightweight structure multi-scale parallel optimization method based on lattice discrete optimization. The present application uses the geometric parameters of movable and deformable rods to describe macroscopic structures and microscopic lattice structures through multi-scale parallel optimization design, and uses an improved double-value coding parameterization method, referred to as BCP method, to solve the discrete optimization of the microscopic lattice structure and easily extract the geometric parameters of the optimization results.
Owner:CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI

Attitude control method, system and equipment of underwater leveling machine and underwater leveling machine

The invention provides an attitude control method, system and equipment of an underwater leveling machine and the underwater leveling machine, and relates to the technical field of ocean engineering. The method comprises the following steps: establishing a dynamic model; constructing a physical information neural network based on the dynamic model; constructing a digital twinborn body of the leveling machine, and performing simulation verification on the attitude error prediction result in the digital twinborn body; based on a simulation verification result, taking control parameters of the underwater leveling machine as optimization variables, performing parallel optimization on a plurality of control targets of the underwater leveling machine, and generating a control solution set; and determining a target control parameter combination from the control solution set, and generating a control instruction for controlling the underwater leveling machine to perform attitude adjustment based on the target control parameter combination. According to the method, the reliability of attitude control of the underwater leveling machine can be improved through prediction of the integrated physical information neural network, simulation verification of the digital twin and multi-objective optimization.
Owner:CHINA COMM FOURTH NAVIGATION BUREAU EIGHTH ENG CO LTD +1

Deep learning model reasoning method and device, computer equipment and storage medium

The invention relates to a deep learning model reasoning method and device, computer equipment and a storage medium. The method comprises the following steps: calling an inference framework matched with a domestic accelerator, and performing compilation optimization, parallel optimization, memory hierarchical optimization and deep fusion of calculation acceleration components on a deep learning model to obtain a target optimization model; compiling the target optimization model into an executable machine code of the domestic accelerator; and loading the executable machine code on the domestic accelerator, and executing a data reasoning process of the executable machine code on to-be-reasoned data to obtain a model reasoning result. By adopting the method, the problem that an existing general reasoning framework is difficult to directly adapt to the domestic accelerator can be solved, the calculation potential of the domestic accelerator is completely released, and the reasoning efficiency and the resource utilization rate are remarkably improved.
Owner:JIANGNAN INST OF COMPUTING TECH

Numerical control machine tool synchronous evolution control method and system associated with dynamic and static errors

The invention discloses a numerical control machine tool synchronous evolution control method and system associated with dynamic and static errors, relates to the technical field of information processing management, and is used for solving the problem that a machine tool structure optimization method is poor in structural integrity, structural connectivity and dynamic response characteristic analysis. By monitoring and analyzing geometric structure parameter changes of all parts in the machine tool operation process, effective geometric structure parameters with the high influence degree are screened out, preliminary optimization is carried out, a geometric structure optimization evaluation model is established, and an optimization result is evaluated; the method comprises the following steps of: acquiring local rigidity and weight, establishing an optimization evaluation model, defining geometric structure parameters of a plurality of components as an integral assembly structure, constructing a parallel optimization model by utilizing a grid division and rigidity matrix calculation method, and finally obtaining optimization results of a macroscopic unit and a connecting unit by taking minimization of structural flexibility as a target. Key parameters can be accurately monitored and optimized, the structural connectivity is considered, and the overall performance and stability are improved.
Owner:SHENZHEN HUAYA CNC MASCH CO LTD

Overdetermined equation parallel random solving method and system for aircraft structure grid

The invention discloses an aircraft structure grid-oriented overdetermined equation parallel random solving method and system, and relates to the technical field of grid generation, and the method comprises the steps: obtaining a numerical aircraft parameter model; generating a grid by adopting a variational harmonic partial differential equation on the basis of an aircraft parameter model, discretizing the variational harmonic partial differential equation, and establishing an overdetermined linear equation set taking a grid vertex coordinate value as an unknown number after discretization; performing low-rank approximate processing on the overdetermined linear equation set to obtain an overdetermined linear equation set after dimension reduction; a greedy random hybrid Kaczmarz algorithm is adopted to carry out parallel solution on the overdetermined linear equation set after dimension reduction, and an optimal solution is obtained; and mapping the optimal solution to a physical space to generate a hybrid grid, and importing the hybrid grid into an aircraft simulation platform for heat flow analysis. According to the method, parallel optimization is carried out on the solution of an overdetermined equation set generated by a grid, and rapid solution of an overdetermined linear equation set is realized through a greedy random hybrid Kaczmarz algorithm, so that the purpose of improving the grid quality is achieved.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Renewable energy power system supply-transmission-demand-storage collaborative optimization configuration method

The invention provides a renewable energy power system supply-transmission-demand-storage collaborative optimization configuration method, and relates to the technical field of smart energy, and the method comprises the steps: obtaining basic data of a plurality of provinces and cities, including renewable energy resource data and auxiliary data; establishing a collaboration degree calculation model, respectively calculating the collaboration degree of each subsystem, calculating the overall collaboration degree based on the collaboration degree of each subsystem, and constructing a subsystem collaboration strength matrix to quantify the coupling relationship among the subsystems; establishing a multi-sub-population collaborative optimization architecture, and respectively generating an initial solution set of each sub-population by adopting a collaborative degree-oriented initialization strategy; adopting an improved sparrow optimization algorithm to carry out parallel optimization on each sub-population, and obtaining an optimal configuration result when a convergence condition is met; and outputting a supply-transmission-demand-storage collaborative optimization configuration scheme of the renewable energy power system according to the optimal configuration result. The overall operation efficiency and economical efficiency of the system can be improved.
Owner:HUBEI UNIV OF ECONOMICS

Self-adaptive supervision control method for logistics storage scheduling

The invention discloses a self-adaptive supervision control method for logistics storage scheduling, and relates to the field of intelligent scheduling, and the method comprises the steps: S1, collecting an observation value of a multi-source sensor, and calculating a confidence coefficient and a state estimation vector in combination with a storage state pattern library; s2, constructing capacity and order perturbation factors based on the global confidence and the state estimation vector, and generating a scene set; s3, calculating a delay index and congestion penalty in combination with a scheduling rule, forming a scheme score, and generating a main scheduling scheme and an alternative scheme; and S4, constructing a real-time key index according to the order execution difference value, monitoring execution in combination with a dynamic threshold value, and triggering scheme switching when the deviation continuously exceeds the limit. Through the synergistic effect of multi-scene adaptive modeling, real-time supervision closed-loop control and a parallel optimization mechanism, the adaptability, robustness and global optimality of the scheduling process are remarkably improved.
Owner:FUJIAN ZHILIAN ALL THINGS TECH CO LTD

Water-energy-medicine collaborative optimization method and system for sewage plant

The invention provides a sewage plant water-energy-drug collaborative optimization method and system, and the method comprises the steps: obtaining the data of a technological process and a material transfer relationship of a sewage plant, constructing a graph network structure model with a technological unit as a node and material flow as an edge, and carrying out the preprocessing, thereby obtaining dynamic coupling graph structure data; a dynamic coupling graph neural network model containing a node feature coding layer, a time sequence coding module, a space message passing layer and an attention mechanism layer is constructed based on the data, and a prediction model capable of representing the dynamic coupling relation of the process unit is obtained through historical data training; then, a multi-objective optimization function which takes the lowest ton water treatment cost as an objective and covers water quality standard reaching, energy consumption and medicament dosage constraints is constructed; and in combination with the prediction model and the optimization function, the optimal operation parameters are solved through an optimization algorithm, and a whole-plant collaborative optimization decision scheme is generated. According to the invention, the dynamic coupling GNN prediction model and the virtual element multi-domain parallel optimization technology are integrated, and intelligent operation management of the sewage treatment plant is realized.
Owner:GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS

Image compression method based on JPEG-LS parallel optimization algorithm

The invention discloses an image compression method based on a JPEG-LS (Joint Photographic Experts Group-Least Squares) parallel optimization algorithm, which comprises the following steps: calculating local gradient values of to-be-coded data and then merging to obtain a context index Q value, the to-be-coded data being a prediction error of a current pixel; grouping the data to be coded according to the context index Q value; if the to-be-coded number contained in the key group affects the parallelism degree, a controllable distortion value is introduced into the pixel value of the to-be-coded data by using an equalization algorithm, and the group serial number corresponding to the data in the key group is corrected; deploying processing units of parallel channels with the number equal to that of the groups, and scheduling the grouped data to be coded to the processing units to realize pipeline processing; integrating each group of data into coded and compressed data through code stream splicing; the image compression method based on the JPEG-LS parallel optimization algorithm is used for image compression, the system data size is low, and the satellite-ground transmission bandwidth pressure is small.
Owner:XIDIAN UNIV

Self-adaptive collaborative parallel optimization aviation complex structural member production scheduling method, system and program product

PendingCN120762877AData processing applicationsResource allocationAviationParallel algorithm
The invention provides a self-adaptive collaborative parallel optimization aviation complex structural member production scheduling method and system and a program product. The method comprises the following steps: presetting parallel algorithm related parameters; initializing a population, wherein the initialized population comprises a first scheduling scheme generated by using a random greedy heuristic algorithm; performing evolutionary optimization on the initialized population or the derivative population of the initialized population in parallel by utilizing a plurality of configured sub-threads to generate a new scheduling scheme; according to a comparison result of the current CPU utilization rate and the memory utilization rate and the target CPU utilization rate and the target memory utilization rate, adjusting the number of sub-threads which are running; according to the method, the problems that in the prior art, a scheduling scheme is not high in quality and low in optimization efficiency, and the utilization rate of computing resources is difficult to consider are effectively solved, the high-quality and high-adaptability scheduling scheme can be generated, the solving time is shortened, the computing resources are fully utilized, and the scheduling efficiency is improved. The production efficiency and the resource utilization rate are improved.
Owner:SHANGHAI UNIV

Parallel optimization method of low-rank adapter and task perception scheduling system

The invention relates to the technical field of large-model lightweight fine tuning, and discloses a parallel optimization method of a low-rank adapter and a task awareness scheduling system.The parallel optimization method comprises the steps that an increment matrix of the low-rank adapter is fragmented to a tensor parallel group according to rows, the tensor parallel group comprises a plurality of computing devices, and the computing devices are used for computing the increment matrix of the low-rank adapter; the fragmentation granularity is dynamically determined according to the equipment hardware capability and the dimension of the increment matrix; dynamically scheduling the tasks with strong conflicts to different task parallel groups based on the calculated inter-task gradient conflict coefficient; and for a plurality of tasks scheduled to the same task parallel group, adapter calculation is merged into unified large matrix operation by adopting a micro-batch processing technology, and LayerNorm and adapter projection calculation are merged into a single calculation kernel by adopting a kernel fusion technology. According to the method, communication redundancy and synchronization overhead in distributed training are reduced, task interference during multi-task parallel is effectively eliminated, and the problem of computing resource fragmentation is solved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Communication calculation parallel optimization method, multiprocessor system, medium and program product

The invention discloses a communication computing parallel optimization method, a multiprocessor system, a medium and a program product, and the method comprises the steps: configuring a first computing core and a second computing core on a single task flow for a general matrix multiplication subtask allocated to a single processor; wherein the first calculation core is used for executing a general matrix multiplication subtask, and the second calculation core is used for executing a set communication task; the set communication task comprises a full accumulation operator or a protocol dispersion operator; then, starting scheduling is conducted on the first calculation core and the second calculation core according to a preset dependency mechanism, so that the general matrix multiplication subtask and the set communication task are executed asynchronously in an overlapped mode; according to the method, the problem of poor reusability of the original Kernel caused by intrusive modification of the GEMM or re-implementation of the Kernel can be effectively avoided through parallel optimization of GEMM calculation and ensemble communication operation, and the performance overhead of the processor is reduced.
Owner:SHANGHAI BIREN TECH CO LTD

Business distributed parallel processing method and device

The embodiment of the invention provides a service distributed parallel processing method and device, for a predetermined service abstracted as tensor calculation, a calculation graph can be constructed according to operators contained in the predetermined service, then in-layer parallel optimization is carried out on the operators of each layer to obtain an in-layer parallel strategy, and then the in-layer parallel strategy is obtained. And distributed business processing can be carried out by utilizing a corresponding in-layer parallel strategy. Wherein the intra-layer parallel strategy comprises a parallel calculation strategy of operators and a parallel transmission strategy of tensors among the operators. According to the embodiment, the parallel strategy of distributed service processing based on the tensor can be effectively optimized, and then the distributed service processing efficiency is improved.
Owner:TSINGHUA UNIVERSITY +1

Fortran program parallel optimization method based on intelligent dependency analysis

The invention discloses a Fortran program parallel optimization method based on intelligent dependency analysis. The method comprises the following steps: firstly, constructing a parallel optimization system consisting of a loop extraction module, a loop nested relation analysis module, a semantic analysis module, an intelligent dependency analysis engine, a variable classification module and an instruction generation and injection module; the loop extraction module analyzes a nested relation and a variable action range of loops; the loop nesting relation analysis module determines a nesting relation between loops; the semantic analysis module constructs a row-level data access view; the variable classification module identifies private variables and reduction variables; the intelligent dependency analysis engine executes loop type check, I / O operation check and data dependency check; and the instruction generation and injection module generates a parallelization instruction and inserts the parallelization instruction into the source code to obtain a parallelized program source code. The method can solve the problems that an existing parallelization method is low in cyclic dependency relation recognition accuracy and safety, and parallelization errors cannot be accurately recognized.
Owner:NAT UNIV OF DEFENSE TECH

A processing system for improving server data storage speed

The application relates to the technical field of server data storage, and discloses a processing system for improving server data storage speed, which comprises a data blocking module, a parallel optimization module, a performance monitoring module, a strategy matching module and a storage execution module, and can comprise a load self-learning module. The data blocking module blocks data streams according to a dynamic strategy and generates a distribution queue; the parallel optimization module generates an optimal storage node combination and a scheduling strategy according to the queue; the performance monitoring module collects real-time performance data and screens for abnormalities; the strategy matching module determines a target adjustment rule through multidimensional matching; the storage execution module drives a node to correct a path; and the load self-learning module optimizes a scheduling factor according to feedback data. The system realizes dynamic blocking, adaptive scheduling and real-time optimization, improves storage efficiency and stability, and is suitable for a server data storage scene.
Owner:GUIZHOU POLYTECHNIC COLLEGE OF COMM +1