Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

78 results about "Parallel computing architecture" patented technology

Heterogeneous computing thread block optimal scheduling method and system based on dynamic topology mapping

The invention belongs to the field of parallel computing architecture optimization, and relates to a matrix multiplication acceleration method and system based on dynamic computing resource mapping, and the method comprises the steps: constructing a dynamic topology model driven by tensor dimension features, and generating a thread block distribution mode according to matrix parameters and GPU hardware information; constructing a multi-dimensional resource scheduling strategy library, dynamically selecting an optimal thread block distribution strategy from the multi-dimensional resource scheduling strategy library, and generating a binding relationship between the thread blocks and the data blocks; calculating collaborative access logic of thread blocks and storage hierarchies based on block parameters and dynamic mapping function optimization; distributed calculation is carried out, calculation and data transmission are parallelized through pipelining and a double-buffering mechanism, and result aggregation across calculation units is completed synchronously through atomic operation and a barrier. According to the method, discontinuous memory access conflicts can be effectively reduced, the execution efficiency of the calculation instruction and the utilization rate of the cache space are improved, the parallel calculation process of accelerating and optimizing the general matrix multiplication is realized, and the data processing efficiency is improved.
Owner:SOUTH CHINA UNIV OF TECH

Mini LED lamp panel needle mark defect detection method and system

The invention discloses a Mini LED lamp panel needle mark defect detection method and system, and the method comprises the following steps: S1, carrying out the preprocessing of a Mini LED lamp panel surface image, suppressing high-frequency noise, and enhancing the defect features; s2, extracting defect features by adopting a self-adaptive convolution adjacency analysis method: S21, dynamically adjusting the shape and size of a convolution kernel through a deformable self-adaptive convolution kernel, and matching the local form of the needle mark defect; and S22, carrying out correlation analysis on the pixel points and neighborhoods thereof, and eliminating isolated noise interference. According to the method, defect forms are dynamically matched through the self-adaptive convolution kernel, isolated noise is eliminated through adjacency analysis, defects of different sizes are covered through multi-scale fusion, wavelet filtering and a parallel computing architecture are combined, the detection precision and the anti-jamming capability are remarkably improved, stable recognition of the tiny needle marks in the complex environment is achieved, meanwhile, high-speed real-time processing is guaranteed, and the detection accuracy and the anti-jamming capability are improved. The method is suitable for diversified industrial scenes and material requirements, reduces the false detection and omission ratio, and promotes the intelligent upgrading of precision manufacturing quality inspection.
Owner:SHENZHEN NANOVISION CORP

GPU (Graphics Processing Unit) accelerated multi-particle coupling Monte Carlo calculation method and device

The invention discloses a GPU accelerated multi-particle coupling Monte Carlo calculation method and device. The method comprises the following steps: asynchronously migrating geometric material parameters subjected to CPU serial processing and a basic database to a pre-allocated memory space in a GPU; operation parameters of the GPU are determined, and Monte Carlo simulation of the multi-particle coupling transportation process is realized in the GPU based on the operation parameters and the neutron nuclear reaction database; on the basis of an atom addition function of a CUDA parallel computing architecture, energy transfer data, obtained through Monte Carlo simulation, of particles and voxel grids in a multi-thread environment is efficiently summarized, energy deposition of the particles is determined, and a high-resolution three-dimensional energy deposition spatial distribution data set is generated. According to the method, high precision can be maintained, and meanwhile, the parallel computing capability of the GPU is fully utilized to realize multi-particle coupling transportation synchronous simulation, so that the computing efficiency of a complex radiation field is improved, and the dual requirements of different fields on high precision and real-time performance are met.
Owner:TSINGHUA UNIVERSITY

Plate defect detection system based on edge calculation

The invention relates to the technical field of image pattern recognition, in particular to a plate defect detection system based on edge calculation, which comprises a convolution kernel directivity screening module, a convolution kernel separable reconstruction module, a model structured constraint fine tuning module and an edge end parallel reasoning module. According to the method, calculation cores sensitive to specific directional defects are automatically screened, the accuracy of recognizing key features such as scratches is enhanced, complex detection operation is decomposed into two simplified vector convolution and one residual error correction calculation, the design greatly compresses the operation complexity and parameter quantity, and the detection accuracy is improved. By means of a parallel computing architecture of an edge end, separated computing tasks are synchronously processed, results are combined, efficient reasoning is achieved on equipment with limited computing resources, real-time on-site detection of plate defects is achieved, delay caused by a data remote transmission center server is effectively avoided, and the detection accuracy is improved. And the response speed and deployment flexibility of the whole detection process are improved.
Owner:FOSHAN POLYTECHNIC

Instruction-level simulation and performance modeling system for parallel computing architecture

The invention provides an instruction-level simulation and performance modeling system for a parallel computing architecture, and belongs to the technical field of computer architecture and simulation verification, and the system comprises an instruction modeling layer which is used for analyzing and executing an intermediate instruction set defined by the architecture; the scheduling execution layer is used for simulating a multi-thread and multi-core parallel execution process; the storage access layer is used for constructing a hierarchical storage access and bandwidth and delay model; and the performance analysis layer is used for collecting and counting key indexes such as an execution period, an instruction utilization rate and memory access delay, and realizing accurate performance modeling of the parallel architecture. According to the method, the performance bottleneck of the design scheme can be rapidly evaluated in the early stage of architecture design, the simulation speed is high, the module configurability is high, the modeling precision is adjustable, and the method is suitable for functional verification, micro-architecture exploration and compiler performance analysis of parallel computing architectures, accelerator chips, heterogeneous multi-core processors and the like.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Power supply network structure weakness detection method based on multi-diagonal-block matrix decomposition

The invention relates to a power network structure weakness detection method based on multi-diagonal-block matrix decomposition, and belongs to the technical field of super-large-scale integrated circuits. According to the method, an original large-scale sparse matrix is converted into a band edge diagonal block structure and divided into a plurality of sub-matrixes capable of being solved independently by constructing a layering and blocking strategy for eliminating tree drive, and redundancy of full-matrix calculation is avoided. Meanwhile, a local approximate inverse algorithm of column norm truncation is designed, and target elements are calculated on the premise that a preset error threshold value is met. According to the method, the parallel computing architecture of the multi-core processor is fully utilized, the computing complexity is reduced by 2-3 orders of magnitude, the computing efficiency is effectively improved, the memory occupation is remarkably reduced, and a high-precision and high-efficiency solution is provided for detecting the weakness of the power network structure of the super-large-scale integrated circuit.
Owner:SHANGHAI LIXIN SOFTWARE TECH CO LTD

Refined collaborative prediction method and system for global goaf subsidence

The invention, which relates to the technical field of mine geology and urban safety, discloses a refined collaborative prediction method and system for subsidence of a global goaf, comprising the steps of establishing a hierarchical storage system, and correcting a working face calculation boundary; extracting the maximum influence radius of all the working faces and four-dimensional coordinates of the mining boundary of each working face, and calculating to obtain a global prediction range to generate a global surface prediction grid; point set insertion and Delaunay triangulation are adopted to obtain an effective triangular unit, a local influence domain range of the effective triangular unit is calculated, and an influenced sub-grid is screened from the global surface prediction grid; and calculating a predicted point movement deformation value based on a parallel computing architecture, accumulating the predicted point movement deformation value to a global result array through a coordinate index mechanism, synthesizing a deformation component in a specific direction, generating a visual result, and performing statistical analysis. According to the method, the problems of poor adaptability and low efficiency of a traditional scheme are effectively solved, and the prediction precision and efficiency are remarkably improved.
Owner:SHANDONG LUNAN GEOLOGICAL ENG SURVEY INST

Power system wiring detection method based on block parallel processing

The invention discloses an electric power system wiring detection method based on block parallel processing, and relates to the field of current detection, and the method comprises the steps: collecting real-time operation data through a sensor disposed at a key node, and constructing a basic data set after Kalman filtering denoising and normalization processing; establishing a sub-block identification system based on topological structure blocks, and extracting electrical and topological feature vectors; a master-slave parallel computing architecture is adopted, a master node allocates tasks according to priorities, and an improved random forest model is utilized to perform sub-block level anomaly preliminary screening; cross-sub-block collaborative verification is realized by verifying the boundary interconnection switch state, the branch current and the bus voltage; and finally, constructing a global wiring state map to perform difference analysis, accurately positioning abnormal equipment and types, and verifying a result through high-precision retest. The method has the advantages that the efficiency is improved through scientific partitioning and distributed parallel computing, the accuracy is guaranteed in combination with multi-dimensional detection, boundary verification and a retest mechanism, wiring abnormity is efficiently positioned, and reliable support is provided.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

New energy station cluster hierarchical simulation modeling system and real-time verification method

The invention discloses a hierarchical simulation modeling system for a new energy station cluster and a real-time verification method, belongs to the technical field of new energy power generation simulation, and solves the problems of high simulation calculation complexity, difficulty in real-time operation and closed-loop interaction verification with an external control system of a large-scale new energy station cluster. According to the technical scheme, the method comprises the steps that a digital dynamic real-time simulation platform configured with a multi-thread parallel computing architecture is adopted, and a new energy station detailed model comprising a wind power generation unit model, a photovoltaic power generation unit model, an energy storage power generation unit model and a current collection system model is deployed on the platform; generating an independent wind speed sequence for each fan through a wind resource distribution model; and constructing a large-base station cluster model composed of a plurality of station simplified models through aggregation equivalence. The system is mainly used for real-time simulation testing of a large-scale new energy base and closed-loop verification of a control system.
Owner:XINJIANG HUADIAN TIANSHAN POWER GENERATION CO LTD

Vector tile cutting and publishing system

The invention relates to the field of geographic information system data processing, in particular to a vector tile cutting and publishing system which comprises a heterogeneous data self-adaptive access and normalization subsystem used for receiving multi-mode heterogeneous input; the spatial-temporal feature analysis and index construction middleware is used for generating a metadata index reflecting data spatial distribution features; the vector tile streaming production engine drives a tile generation process based on the metadata index, converts the intermediate state data into a vector tile data packet by adopting a parallel computing architecture, and writes the vector tile data packet into a tile storage warehouse; and the standardized service publishing bus is used for monitoring the change state of the tile storage warehouse in real time and externally providing a vector tile network service interface which accords with a standard protocol. According to the method, the structure barriers of heterogeneous data sources are eliminated, the stability of data production is ensured through topology self-healing, the defect that one set of rules cut all is overcome through a density-based parameter inversion mechanism, and the tile size and the visualization precision are balanced.
Owner:YUNTU ZHIXING (BEIJING) TECHNOLOGY CO LTD

NoC congestion control method based on redundant data merging

The invention discloses an NoC congestion control method based on redundant data merging, and relates to the technical field of general parallel computing and network-on-chip. Aiming at NoC congestion caused by redundant data transmission in a general parallel computing architecture and a bottleneck problem of a memory controller caused by the NoC congestion, an adopted scheme comprises the following steps of: marking cache requests from different SMs and aiming at the same cache block by integrating a grouping and merging unit in the memory controller; and after the memory controller accesses and obtains the cache blocks, injecting the cache blocks into the NoC according to the request marks, combining the cache blocks shared by the plurality of SM into a single data packet by utilizing the NoC characteristics, and sending the single data packet to each SM initiating the request through multicast routing. The method is applied to a general parallel computing architecture, and transmission of redundant data in the NoC can be avoided.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Parallel computing and distributed network access risk management and control device and management and control method

The invention discloses a parallel computing and distributed network access risk management and control device and a management and control method, and relates to the technical field of risk management and control. The system comprises three layers, namely network element equipment, a network controller and a security management and control platform, the security management and control platform integrates a PKI-based password infrastructure, a tool engine and an AI engine and uniformly distributes security management and control related strategies, and the network controller is used as an edge execution component and is connected with various network element equipment of a managed network, so that the security management and control related strategies can be uniformly distributed. The risk management and control device is constructed on a full-switching parallel computing architecture which provides a virtual computing power pool. Under the support of high-performance computing power provided by full-switching parallel computing, a safety management and control platform based on a distributed safety management and control system is constructed, unified state monitoring, risk analysis and strategy issuing are carried out on network access risks, and the safety of network access is improved through distributed network event and performance management and control. And network access risk real-time management and control are carried out.
Owner:SHENZHEN Y& D ELECTRONICS CO LTD

Block-level air quality dynamic traceability method and system based on domestic credit and innovation platform

The invention discloses a block-level air quality dynamic traceability method and system based on a domestic credit and creativity platform, belongs to the technical field of block-level air quality dynamic traceability, integrates 12 types of data sources of an air quality monitoring station, a traffic checkpoint and an industrial internet of things, supports minute-level data updating, carries a three-level nested grid parallel computing architecture, and provides a dynamic traceability system for the block-level air quality. The method supports bidirectional data coupling and distributed calculation of 10km, 3km and 250m grids, comprises a three-dimensional value surface rendering component and a dynamic particle tracking component, realizes pollution cloud picture rendering delay less than or equal to 500ms and particle tracking precision less than or equal to 50m, sets a threshold trigger type early warning mechanism, and automatically generates management and control suggestions in combination with a pollution contribution decoupling result.
Owner:JINAN UNIVERSITY +1

An event-driven and multi-layer parallel architecture-based revenue processing method and system

The present disclosure provides a kind of income processing method and system based on event-driven and multilayer parallel architecture, the method comprises: receiving and processing multi-source trigger signal, generates the standardized trigger context of income calculation task;According to the standardized trigger context, the income calculation task is split into multiple independent subtasks, encapsulated as a standardized income event message and published;Subscription and processing standardized income event message, split each independent subtask into atomic subtask, execute distributed income calculation, generate customer income raw results;Customer income raw results are converged and persisted to unified income accounting table.The present disclosure effectively solves the bottleneck of traditional system in performance and expansibility through event-driven and parallel computing architecture, and further through the innovative initiative audit repair mechanism, builds endogenous, automated data quality guarantee system, realizes the overall upgrade of reliability, accuracy and operation efficiency of income processing system.
Owner:BANK OF HANGZHOU CO LTD

Transformer heavy overload prediction method and device, equipment and storage medium

This invention discloses a method, apparatus, device, and storage medium for predicting transformer overload. By calculating the multi-dimensional dynamically coupled adjacency matrix of the transformer, this adjacency matrix can be used in conjunction with a spatiotemporal attention bidirectional feedback network to effectively improve the accuracy of overload prediction and reduce the false judgment rate under extreme conditions. Furthermore, both the multi-dimensional dynamically coupled adjacency matrix and the spatiotemporal attention bidirectional feedback network are lightweight parallel computing architectures, resulting in a short prediction time for the target transformer's load rate and low resource consumption, which can meet the real-time dispatching requirements of the power grid. The invention also obtains overload and overload benchmark thresholds based on the historical number of overloads in the target transformer's static equipment dataset. These thresholds are used to predict whether the target transformer is overloaded or overloaded, allowing for threshold settings based on the specific conditions of the target transformer to reflect individual differences and further improve prediction accuracy.
Owner:YUNNAN POWER GRID CO LTD ELECTRIC POWER RES INST

Distributed compressor margin calculation method and device

The invention provides a distributed gas compressor margin calculation method and device, and belongs to the field of gas compressor pneumatic optimization. The method provided by the invention comprises the following steps: building a distributed parallel computing architecture; constructing a flow field convergence judgment rule based on an efficiency residual error; and adopting a surge boundary self-search strategy for dynamically adjusting the back pressure, carrying out multi-round full three-dimensional flow field parallel calculation by virtue of the distributed parallel calculation architecture, verifying a calculation result of each round in combination with the flow field convergence judgment rule, and iteratively adjusting the outlet back pressure to approach to a gas compressor near surge point, so as to achieve the purpose of reducing the flow field convergence judgment rule. Until a near surge point of convergence and divergence critical states is determined and performance parameters are acquired; and total pressure ratio and flow performance parameters of the design point of the gas compressor are extracted, and the margin of the gas compressor is calculated by combining the performance parameters corresponding to the near surge point. According to the distributed compressor margin calculation method and device, the margin calculation precision and calculation efficiency are both improved, and meanwhile calculation resources are utilized to the maximum extent.
Owner:JINCHENG NANJING ELECTROMECHANICAL HYDRAULIC PRESSURE ENG RES CENT AVIATION IND OF CHINA

A method and device for multi-objective collaborative optimization of coal blending and blending combustion in a thermal power plant

The application discloses a kind of multi-objective collaborative optimization method and device of coal blending of thermal power plant, belong to energy management and intelligent control technical field.Integrating the coal quality of thermal power plant, operation, emission and scheduling data, construct unified feature library, and based on this, set the initial weight of economic, safety, environmental protection target, establish multi-objective optimization model;Adopt parallel computing architecture to solve model, generate multiple candidate blending scheme;Subsequently, combined with the artificial adjustment result of operator and actual combustion feedback data, carry out secondary prediction and comparative analysis;Finally, through continuous learning adjustment behavior and operation effect, drive model automatically update weight and optimization algorithm combination, form closed loop self-learning mechanism, to realize multi-objective collaborative optimization and adaptive decision of thermal power plant.The application improves the intelligence, economy and environmental protection of system operation.
Owner:XIAN TPRI BOILER ENVIRONMENTAL PROTECTION ENG CO LTD

Satellite internet model security data sharing method based on federated learning

The invention discloses a federated learning-based satellite internet model security data sharing method, which comprises the following steps of: firstly, constructing a self-adaptive ground mobile terminal clustering strategy based on a dynamic graph attention mechanism, and establishing a space-time sensing ring topology by analyzing a node movement rule and communication characteristics to lay a network foundation for decentration calculation; secondly, a modular decomposition confusion encryption scheme is developed, homomorphic encryption of gradient parameters is realized by using a redundant modular component superposition and random number injection technology, the ciphertext size is compressed, and dependence of a third-party center is avoided through distributed private key management. And finally, creating a multi-sub-ring parallel computing architecture, splitting a modular decomposition task into independent computing units, and realizing local model combination based on a proximity communication principle, so that the bottleneck of a traditional server is eliminated, and the aggregation delay is also reduced. According to the whole scheme, through collaborative design of a topology layer, an encryption layer and a calculation layer, feasibility and efficiency balance of cross-domain collaborative training is achieved under the condition that data privacy is guaranteed.
Owner:BEIJING ELECTRONICS SCI & TECH INST

GPU parallel grating generation system and method based on CUDA

The invention discloses a CUDA (Compute Unified Device Architecture)-based GPU (Graphics Processing Unit) parallel grating generation system and a CUDA-based GPU parallel grating generation method, pixel-level mapping calculation in a grating generation process is executed by GPU multithreads by adopting a CUDA-based parallel computing architecture, and better generation efficiency can be kept when high-resolution images and multi-parameter combinations are processed. A pre-verification mechanism of configuration management is matched, calculation interruption and result deviation caused by invalid parameters can be reduced, the generation process is more stable, repeated reproduction is more convenient, meanwhile, input images, output paths, grating parameters and mapping options are presented in a centralized mode through a graphical user interface, progress and state feedback is provided, and the generation efficiency is improved. Therefore, a user can complete configuration, execution and output management in the same process. In combination with compatible processing of Chinese paths and multi-format images and file output modes such as DFF, engineering access cost can be reduced, and convenience of result delivery can be improved.
Owner:ZHEJIANG TRILLION GAME TECH

A graph computation optimization method and system based on multicentrism metric and global optimal ranking

This invention discloses a graph computing optimization method and system based on multi-centrality metrics and globally optimal ranking. The method first calculates the multi-centrality metric of vertices based on an industry knowledge graph, generating candidate ranking sequences adapted to different business operations, and then filters out efficient sequence subsets through performance testing. Subsequently, a parallel computing architecture and a business-region combined partitioning strategy are employed to perform a breadth-first search in parallel, starting from key nodes, generating locally optimal sequences. By incorporating sequence similarity calculations with industry chain weights, the filtered efficient sequence subsets are merged with the locally optimal sequences to form a globally optimal sequence. This effectively overcomes the limitations of traditional single ranking in adaptability across multiple business scenarios and significantly reduces computational complexity. The constructed globally optimal sequence demonstrates higher accuracy and adaptability in businesses such as supply chain risk identification and precise investment attraction, providing an efficient and stable graph computing solution for industrial cloud platforms in the equipment manufacturing field.
Owner:HUNAN GUOZHONG ZHILIAN CONSTR MASCH RES INST CO LTD

Multi-array-element airspace zeroing anti-interference virtual method of MobileNet architecture

The invention relates to a multi-array-element airspace zeroing anti-interference virtual method of a MobileNet architecture, and belongs to the field of Beidou and the field of radio communication and artificial intelligence crossing technologies. The method comprises the steps of radio-frequency signal acquisition and digital preprocessing, MobileNet model improved design, virtual array element extension implementation, airspace zeroing weight calculation and optimization, FPGA hardware acceleration execution, beam forming and anti-interference output. According to the virtual array element expansion technology, the number of effective array elements can be increased by 2-4 times, the spatial freedom degree is increased to 15-31 from 7, 4-8 independent nulls can be formed at the same time within the range of + / -60 degrees, and the multi-interference-source suppression ratio is increased to 35 dB or above from 20 dB of a traditional method; by means of the MobileNet lightweight model, the weight calculation complexity is reduced, and by combining the parallel calculation architecture of the FPGA, the real-time performance is obviously improved, the hardware resource consumption is greatly reduced, and the dynamic adaptive capacity is good.
Owner:BEIDOU APPL DEV RES INST

GPU-based large-breadth SAR image rapid positioning method

The invention discloses a GPU (Graphics Processing Unit)-based large-breadth SAR (Synthetic Aperture Radar) image quick positioning method, which adopts a CPU and GPU heterogeneous parallel computing architecture and utilizes the many-core characteristics of a GPU to ensure that the efficiency is higher and the time consumption is greatly reduced in the large-breadth SAR image quick positioning process and can reach a minute level or even a second level; the parallel advantage of the GPU is fully exerted, the generation of a plurality of parameters is parallelized, the degree of parallelism is improved, and the operation efficiency is greatly improved; the characteristics that SAR images obtained through multi-mode imaging are located in different coordinate systems, and transplantation and development are tedious are considered, the relation between a first-level SAR image and a second-level correction image is directly established through fitting, processing in the different coordinate systems can be completed only by updating the coordinate conversion relation, and the processing efficiency is improved. Therefore, positioning processing of the SAR image obtained by multi-mode imaging is completed, transplantation and development are facilitated, and the method is more universal.
Owner:XIDIAN UNIV

Balanced adjustment method and system for sound console volume controller

The invention discloses a balance adjustment method and system for a sound console volume controller, and belongs to the field of sound console control. The method comprises the steps that audio signals and environment noise are collected and preprocessed; audio state perception and noise identification are carried out by using a deep learning model combining a time sequence convolutional neural network and a long short-term memory network; a volume control and frequency equalization dual-channel parallel computing architecture is adopted, and controlled variables are generated based on a first-order active-disturbance-rejection control algorithm and a second-order active-disturbance-rejection control algorithm; using an improved particle swarm optimization algorithm to carry out collaborative optimization on the dual-channel control parameters; and finally synthesizing and outputting a high-quality audio signal. The system correspondingly comprises an audio acquisition module, a deep learning reasoning module, a dual-channel control module, a parameter optimization module and the like, and can be integrated with a hardware acceleration unit. The method effectively solves the problem of coupling of volume and frequency balance control, and has the advantages of high environmental noise adaptive capacity, intelligent parameter adjustment, high processing real-time performance and the like.
Owner:ENPING JIACHUANG AUDIO TECHNOLOGY CO LTD

Ocean multi-dimensional mixing process global analysis optimization method

The invention relates to the crossing field of marine geochemistry and computational mathematics, in particular to an optimization method for global analysis of a marine multi-dimensional mixing process, which comprises the following steps: S1, data acquisition; s2, carrying out data initialization processing; s3, calculating data; s4, analyzing a result; s5, outputting data; through collaborative design of multi-stage step length optimization and parallel computing architecture, the key problem that calculation efficiency and precision of a traditional OMPA model in a high-dimensional parameter space are difficult to consider at the same time is effectively solved. The core technology breakthrough is embodied in three aspects: firstly, a three-stage step optimization mechanism (0.05-0.01-0.001) with an adaptive characteristic is developed, and the calculation complexity is greatly reduced on the premise of ensuring a global optimal solution; secondly, a parallel computing framework based on MATLAB parFor is constructed, so that the computing efficiency of the 3-6-dimensional end member hybrid model is improved by an order of magnitude; and finally, a standardized data processing flow (z-score method) is matched with a modular architecture design, so that the expandability and applicability of the system are enhanced.
Owner:SECOND INST OF OCEANOGRAPHY MNR

GPU parallel computing driven large-scale unmanned aerial vehicle cluster simulation method

The invention discloses a GPU parallel computing driven large-scale unmanned aerial vehicle cluster simulation method, which comprises the following steps: S1, building an unmanned aerial vehicle numerical simulation model, generating a code and optimizing the code; s2, optimizing a memory access mode of cluster parallel simulation; s3, optimizing a simulation task division strategy according to the characteristics of unmanned aerial vehicle cluster state updating; s4, according to the S1, S2 and S3, cluster parallel simulation is achieved on the code level; s5, realizing acceleration of a cluster formation algorithm by using parallel computing strategies of S2 and S3; and S6, integrating the cluster parallel simulation in S4 and the formation control algorithm in S5 by using a UDP protocol. According to the multi-core parallel computing architecture based on the GPU, efficient simulation of the large-scale unmanned aerial vehicle cluster is achieved, and the technical problems that a traditional CPU simulation method is poor in real-time performance, difficult to achieve and the like in simulation of more than thousands of large-scale unmanned aerial vehicle clusters are solved.
Owner:SICHUAN UNIVERSITY OF SCIENCE AND ENGINEERING

Method and device for multi-modal fault diagnosis of momentum exchange device based on trajectory sensitivity

This invention relates to the field of aerospace technology, and particularly to a method and apparatus for multimodal fault diagnosis of momentum exchange devices based on trajectory sensitivity analysis. The method includes: connecting multiple types of sensors on the momentum exchange device to acquire multimodal monitoring data in real time; running a trajectory sensitivity analysis algorithm to calculate the residuals of each channel and generate normalized multi-channel residual signals; performing fusion analysis on the multi-channel residuals based on a parallel computing architecture, outputting anomaly scores and comparing them with thresholds to determine the fault type and location. This invention uses a sensitivity analysis algorithm to calculate the derivative of the influence of parameter changes on the system's output trajectory in real time, and guides the generation and normalization of residual signals. This amplifies and highlights minute deviations caused by slight increases in friction, bearing wear, etc., thereby achieving timely detection of faults in their early stages and overcoming the shortcomings of traditional methods in terms of insufficient sensitivity to initial weak anomalies.
Owner:CHONGQING UNIV

Method and device for realizing radar data processing

The invention discloses a method and a device for realizing radar data processing, which adopt a parallel computing architecture, perform data throughput with multiple parallelism degrees, optimize a data access mode, further improve throughput rate, improve efficiency and flexibility of data access in a radar sliding window detection process, reduce cache overhead and improve data access efficiency. High-speed, parallel and clear-structure neighborhood data extraction operation is realized, and rapid and accurate data access support is provided for target detection.
Owner:CALTERAH SEMICON TECH (SHANGHAI) CO LTD

Feature-Preserving Mesh Processing Methods and Systems Based on High-Performance Parallel Computing

This invention discloses a feature-preserving mesh processing method and system based on high-performance parallel computing. The input is an unlabeled point cloud dataset, used to train a point cloud feature generator and a prediction head MLP. α The input is a grid dataset with semantic segmentation labels. The grid surface is sampled to generate a sparse point cloud, and noise is added to the grid to represent the redundant structure of the grid geometry clipping task. Then, the trained prediction head MLP is removed. α Train prediction head MLPs separately β With MLP γ The user inputs the mesh to be processed, all vertices are converted into a sparse point cloud, and the point cloud feature generator and prediction head MLP are used. β MLP γ The invention generates semantic segmentation labels and redundant structure category labels for the sparse point cloud, respectively; finally, it performs a culling operation on the points marked as redundant structures to complete the mesh pruning. This invention can effectively understand the features in the mesh data, and the semantic segmentation results are more accurate. The noise addition process of this invention uses a CUDA parallel computing architecture to ensure the high efficiency of the entire computation process.
Owner:SUN YAT SEN UNIV

Methods, apparatus, media, and products for multi-engine asynchronous parallel computing architecture

This invention discloses a method, apparatus, medium, and product for a multi-engine asynchronous parallel computing architecture. The method includes: acquiring interval marker pairs, which indicate the start and end of a target program segment in a preset program; in response to the interval marker pairs, determining the boundary notification corresponding to the target program from among multiple notifications sent to the control core when each asynchronous engine related to the target program segment performs an operation; and determining the actual execution time of the target program segment based on the sending times of all boundary notifications sent by all asynchronous engines related to the target program segment.
Owner:MOXIN ARTIFICIAL INTELLIGENCE TECH (SHENZHEN) CO LTD

Cesium feed-in system DSMC numerical simulation optimization method based on CUDA parallel

The invention relates to the technical field of a cesium feed-in system in a magnetic confinement nuclear fusion negative neutral beam system, in particular to an optimization method for DSMC numerical simulation of a cesium feed-in system based on CUDA parallel. According to the technical scheme, the method comprises the following steps: setting a molecular initial state at a host end, and preparing initial conditions for analog computation; allocating a memory space for the molecular array and the related data structure at the equipment end; molecular motion is calculated in parallel through a GPU at an equipment end, and collision between molecules and a wall surface is processed. According to the method, the simulation efficiency is improved to the minute level through a CUDA parallel computing architecture, a grid-based molecular index and memory access mechanism is optimized, the statistical reliability is guaranteed by adopting a thread-independent random number generator, near-vacuum molecular index is accurately restored by means of a decoupled physical model, and the statistical reliability is improved. Therefore, an efficient and practical analysis tool is provided for cesium feed-in system engineering design on the premise that the numerical precision is guaranteed.
Owner:INST OF ENERGY HEFEI COMPREHENSIVE NAT SCI CENT (ANHUI ENERGY LAB)