Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

23 results about "Massively parallel computation" patented technology

In computing, massively parallel refers to the use of a large number of processors (or separate computers) to perform a set of coordinated computations in parallel (simultaneously). In one approach, e.g., in grid computing the processing power of a large number of computers in distributed, diverse administrative domains, is opportunistically used whenever a computer is available.

Chip system, data transmission method and related equipment

According to the chip system, the data transmission method and the related equipment provided by the embodiment of the invention, the chip system comprises IO core particles and a plurality of calculation core particles, each calculation core particle is configured with a corresponding core particle center, and each calculation core particle is connected with at least two physical interconnection ports in the IO core particles through the core particle center; the core grain center comprises a splitting unit and a merging unit, and the splitting unit is used for distributing data streams sent by the computing core grains to at least one physical interconnection port according to service types; the merging unit is used for recombining the data stream received from the physical interconnection port or directionally forwarding the data stream to the computing core grain, and the data throughput rate and the overall transmission efficiency in a large-scale parallel computing scene are remarkably improved on the premise that the hardware specification of the external IO core grain does not need to be modified.
Owner:SUZHOU YIZHU INTELLIGENT TECH CO LTD

Terahertz communication sensing and calculating integrated system and method thereof

The invention provides a terahertz communication sensing and computing integrated system based on hierarchical resource scheduling and mixed waveforms. The system comprises a unified radio frequency front end which adopts a high-integration-level terahertz phased array chip and supports ultra-large bandwidth transmission and flexible multi-beam forming; the baseband processing unit (BBU) is used for realizing real-time processing and large-scale parallel computing of signals; the OCSE waveform generation module is used for generating an ''orthogonal traffic signal-perception embedded (OCSE)'' waveform; the parallel receiver processing module comprises a communication demodulation path and a sensing parameter extraction path which are arranged in parallel; and the three-layer intelligent resource scheduler constructs an application-function-physics three-layer closed-loop scheduling model based on a deep reinforcement learning (DRL) algorithm, and is used for dynamically adjusting the distribution proportion of physical resources between communication and sensing functions according to an upper-layer application real-time task demand and a bottom-layer physical environment sensing result so as to maximize the comprehensive efficiency of the system.
Owner:WANWEI PERCEPTION (NINGBO) ELECTRONICS CO LTD

Millimeter wave radar boiler heating surface tube crack detection method based on GPU parallelization

The invention provides a millimeter-wave radar boiler heating surface tube crack detection method based on GPU parallelization, and the method comprises the steps: constructing a three-layer parallelization frame integrating data flow, calculation and control, and employing the large-scale parallel calculation capability of a GPU to detect the crack of a heating surface tube of a millimeter-wave radar boiler. The whole-process acceleration processing from preprocessing, range profile reconstruction, feature extraction to intelligent identification is carried out on massive echo signals acquired by the millimeter wave radar, so that high-precision and high-real-time online detection and positioning of the heating surface pipe wall fine cracks are realized in a complex environment in a boiler with strong interference; finally, the detection efficiency is improved, meanwhile, the technical effect of high detection confidence is kept, and powerful technical support is provided for safe operation and preventive maintenance of the power plant boiler.
Owner:XIAN THERMAL POWER RES INST CO LTD +1

Real-time 3D reconstruction method and device based on speckle image, equipment and medium

The invention discloses a real-time 3D reconstruction method, device and equipment based on a speckle image and a medium, and relates to the technical field of computer vision and three-dimensional reconstruction, the method makes full use of the large-scale parallel computing capability of a GPU, key steps such as stereo matching, parallax filtering and point cloud generation are efficiently executed in a video memory, and the real-time 3D reconstruction efficiency is improved. And the second-level or even millisecond-level 3D reconstruction response speed is realized. Compared with traditional CPU serial processing, the method has the advantages that data processing delay can be greatly reduced, continuous three-dimensional reconstruction of dynamic objects or real-time scenes is supported, and the space structure of the current moment can be displayed in real time. The technology effectively improves the timeliness and availability of 3D information acquisition, and can be widely applied to the fields of robot navigation, augmented reality, industrial detection, dynamic modeling and the like.
Owner:ARIEMEDI MEDICAL SCI BEIJING CO LTD

An AI multi-agent and digital twin fusion production process visualization method, medium and system

The application provides an AI multi-agent and digital twin fusion production scheduling process visualization method, medium and system, belonging to the technical field of AI multi-agent production scheduling. Large-scale parallel computing is realized by constructing a GPU three-layer CUDA processing architecture. The first layer of data preprocessing grid performs data cleaning in parallel. The second layer of negotiation analysis grid runs a Transformer-based agent interaction recognition model for parallel pattern recognition. The third layer of visualization calculation grid runs an LSTM-CNN fusion production scheduling process mapping model to generate visualization data. The multi-head attention mechanism dynamically adjusts the allocation of computing resources. The pipeline parallel and data parallel strategies optimize the computing performance. The asynchronous data transmission and computing overlap technology hides the memory access delay, solving the technical problem that multi-agent high-frequency negotiation data cannot be processed in real time.
Owner:BEIJING NANCAL RUIYUAN DIGITAL TECH CO LTD

A bow sonar demonstration device based on GPU parallel architecture

The application discloses a bow sonar demonstration device based on a GPU parallel architecture, which mainly comprises a hardware system, a software system and a demonstration model, wherein the hardware system is used for completing a teaching task of simulating an operating sonar, wherein a GPU is used as a signal processing machine to provide large-scale parallel computing capacity and ensure real-time construction of a complex acoustic scene; the software system comprises signal processing software and display control software and is used for simulating teaching demonstration of basic sonar functions including noise warning, tracking, listening and feature analysis; and the demonstration model comprises one or more of a vector hydrophone model, a scalar hydrophone model, a sound base array model and a submarine model. The device realizes the visualization of abstract concepts of sonar equipment through three-dimensional dynamic visualization and real-time signal feedback.
Owner:SHENYANG LIAOHAI EQUIP

Method and apparatus for accelerating exact probability simulation in combinational equivalence checking

This invention discloses an accelerated method and apparatus for precise probabilistic simulation in combinatorial equivalence checking, relating to the fields of electronic design automation (EDA) and high-performance computing. The method includes: parsing the input circuit combinational logic netlist to form an EPS operator graph and corresponding data distribution configuration, transmitting this data to a GPU; starting the GPU to distribute the EPS operator graph to a processing unit array for EPS calculation based on the data distribution configuration using a data flow-driven approach; and collecting the running status information of each round of calculation and feeding it back to the CPU for dynamic optimization of the next round of calculation. This method can transform circuit calculation tasks into data flow-driven EPS operator graphs and execute them on the GPU, efficiently utilizing the GPU's massively parallel computing capabilities to overcome performance bottlenecks. The pipelined task advancement based on superblocks and the dynamic optimization mechanism based on performance reports can adapt to the computational characteristics of different circuit structures, balance the load, optimize resource allocation, and significantly shorten the verification cycle of combinatorial equivalence checking.
Owner:BEIJING NEW ENERGY VEHICLE TECH INNOVATION CENT CO LTD

Chip system, data transmission method and related device

The chip system, the data transmission method and the related equipment provided in the embodiments of the present application, the chip system comprises: an IO chip and a plurality of computing chips, each computing chip is configured with a corresponding chip hub, and each computing chip is connected with at least two physical interconnection ports in the IO chip through the chip hub; the chip hub comprises a splitting unit and a merging unit, the splitting unit is used for distributing the data stream sent by the computing chip to at least one physical interconnection port according to the service type; the merging unit is used for recombining or directionally forwarding the data stream received from the physical interconnection port to the computing chip, without modifying the external IO chip hardware specification, the data throughput rate and the overall transmission efficiency in the large-scale parallel computing scene are significantly improved.
Owner:SUZHOU YIZHU INTELLIGENT TECH CO LTD

A GPU fluid simulation method for real-time interactive applications

The application provides a GPU fluid simulation method for real-time interactive application, and relates to the technical field of fluid simulation. The method first uses a particle system to emit a large number of particles, calculates fluid particle positions using a position-based fluid particle simulation algorithm, updates the motion state of the particles, then obtains a smooth fluid surface according to the particle positions through screen space processing, and finally uses a realistic rendering method to color the fluid surface to obtain a water body image. Since the state update of the large number of particles and the pixel-level calculation of the screen space are both large-scale parallel calculations, they are suitable for GPU processing, so all the calculation processes involved in the method are run on the GPU end. The method can be widely applied to various demand scenarios in the fields of industrial simulation and electronic games.
Owner:NORTHEASTERN UNIV CHINA

Optimization processing method, system, device and medium of ai accelerator

The application provides an AI accelerator processing optimization method, system, device and medium, and the method is applied to an AI accelerator in communication connection with a main memory. Through the collaborative architecture of the AI accelerator and the 3D DRAM, the application realizes reduction of transmission overhead and power consumption in the data carrying process, improvement of the computing throughput and the energy efficiency ratio, and simultaneously supports large-scale parallel computing tasks relying on the high bandwidth characteristics of the 3D DRAM.
Owner:SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD

A Scalable Hardware Accelerator for Neural Networks Based on FPGA

This invention provides a scalable hardware accelerator for neural networks based on FPGA, comprising a storage module, a computing module, and a control module. The storage module includes a preset parameter storage area and an intermediate feature map storage area. The preset parameter storage area stores convolutional weight parameters, bias parameters, and computation instruction sequences. The intermediate feature map storage area stores feature map data generated by each layer of the neural network, employing a multi-bank parallel feature map storage structure to achieve parallel access to intermediate feature map data through the collaborative work of multiple banks. The computing module includes a large-scale parallel computing unit array for convolutional neural network computation, achieving high-speed execution of convolution operations through the collaborative work of multiple parallel computing units. The control module performs timing control on convolutional kernel loading, bias loading, feature map reading, and computation instruction sequence execution. This invention improves computational efficiency while enhancing the system's flexibility and scalability.
Owner:PEKING UNIV

Three-dimensional to two-dimensional rendering method and system, electronic equipment and storage medium

The invention provides a three-dimensional-to-two-dimensional rendering method and system, electronic equipment and a storage medium, and the method comprises the steps: firstly receiving original three-dimensional scene data, carrying out the asynchronous and parallel preprocessing of the original three-dimensional scene data, and generating standardized data; and then carrying out synchronous and parallel data conversion processing on the standardized data by a plurality of rendering channels to generate a plurality of intermediate result images, so that the large-scale parallel computing capability of image processing equipment is fully utilized, the processing efficiency is improved, and the stability of color and light and shadow effects is improved. And finally, performing weighted mixing and synthesis processing on the plurality of intermediate result images to generate a corresponding two-dimensional image, thereby avoiding the problem that real-time feedback cannot be realized due to the fact that a complete rendering pipeline needs to be executed again during parameter adjustment, and ensuring that the rendering effect achieves the effect of what you see is what you get. Therefore, the effect stability of three-dimensional to two-dimensional rendering can be ensured, resource waste can be avoided, the time cost is effectively saved, and the rendering efficiency is improved.
Owner:BYD CO LTD

Optimization processing method, system and equipment of AI accelerator and medium

The invention provides a processing optimization method, system and device of an AI accelerator and a medium. The method is applied to the AI accelerator in communication connection with a main memory. Through the collaborative architecture of the AI accelerator and the 3D DRAM, the transmission overhead and power consumption in the data carrying process are reduced, the computing throughput and the energy efficiency ratio are improved, and meanwhile large-scale parallel computing tasks are supported by means of the high-bandwidth characteristic of the 3D DRAM.
Owner:SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD

Transient electromagnetic arbitrary waveform three-dimensional forward modeling method, device and equipment

PendingCN122065573ABreaking through sequence constraintsMaintain numerical accuracyDesign optimisation/simulationSpecial data processing applicationsTransient electromagneticsMesh grid
The invention relates to a transient electromagnetic arbitrary waveform three-dimensional forward modeling method and device and computer equipment. The method comprises the following steps: decomposing an emission source into a plurality of independent sources, adaptively constructing a local grid of each receiving point according to skin depth information of each time channel, and forming a minimum observation unit on a single time channel by a single independent source and a single local grid; determining a characteristic step length of a corresponding time channel according to the skin depth information, and calculating an electromagnetic field response result on a minimum observation unit; electromagnetic field response results of the multiple minimum observation units are subjected to parallel calculation, and a complete transient electromagnetic arbitrary waveform three-dimensional forward modeling result is obtained through synthesis. According to the method, a complete parallelization calculation framework is realized, the calculation load is greatly reduced while the numerical precision is maintained, the broadband transient electromagnetic efficient forward modeling simulation is possible in a large-scale parallel calculation environment, and a powerful technical support is provided for three-dimensional electromagnetic detection under a complex geological condition.
Owner:AEROSPACE INFORMATION TECH UNIV

Hierarchical Latch-Tree Memory Architecture for Large-Scale Parallel Computing Systems

A hierarchical latch-tree memory architecture distributes and stores data values for massively parallel computing arrays without conventional SRAM or cache memory. Transparent latch circuits arranged as sequential tree levels propagate data on phase-shifted clock signals, providing local retention and eliminating read-write contention. Separate forward and reverse latch-tree networks broadcast activation data and collect output sums from processing elements. The latch hierarchy maintains balanced timing and extremely low energy per bit, supporting rack-scale computing throughput on the order of one zettaFLOPS (FP4 sparse AI inference) within practical power limits.
Owner:SILVEBROOK KIA

Video subtitle removal method, chip, GPU and electronic equipment

The invention relates to the technical field of computers, in particular to a video subtitle removal method, a chip, a GPU (Graphic Processing Unit) and electronic equipment, the method is executed on the GPU and comprises the following steps: receiving an input video and decoding the input video into a frame sequence; performing subtitle area detection on the current frame in the frame sequence, and generating a corresponding subtitle mask; in response to the subtitles existing in the current frame, predicting the content of the subtitle area according to the subtitle mask by using the spatial-temporal correlation of the previous frame and / or the subsequent frame in the frame sequence; and filling the predicted content into the caption region to generate a caption-removed video frame. According to the embodiment of the invention, a series of tasks such as video decoding, subtitle area detection, content prediction based on spatial-temporal correlation and final filling are deployed on the GPU, the large-scale parallel computing capability of the GPU is utilized to accelerate the processing flow, and the visual naturalness and the spatial-temporal coherence of the repaired video picture are improved.
Owner:MOORE THREADS TECH CO LTD

AI-based large-scale computer room computing power resource dynamic scheduling operation and maintenance system

This invention relates to the field of cloud computing resource management technology and discloses an AI-based dynamic scheduling and operation and maintenance system for large-scale data center computing resources. The system includes a kernel status monitoring module, a scheduling priority quantification module, a resource configuration retrieval module, and a synchronous clock scheduling execution module. It collects the system call trajectory of the task to be scheduled in the kernel state and the synchronous blocking state of the network protocol stack, extracts runtime characteristic data representing the execution phase, and calculates the scheduling urgency index by combining the topological weights of the task's directed acyclic graph. Then, it matches the target control group quota and writes the target control group quota into the resource limit parameter file when the task enters the communication waiting window. This invention locks the resource quota update action within the task's logical rest period, effectively eliminating the kernel state context switching overhead caused by computing power scheduling, ensuring the logical clock consistency of large-scale parallel computing tasks, and improving the global turnover efficiency of the resource pool.
Owner:SHAANXI KERIDI ELECTRONIC TECH CO LTD

Large loop source transient electromagnetic fast forward modeling method based on GPU acceleration

The invention discloses a large-loop source transient electromagnetic fast forward modeling method based on GPU acceleration, and the method comprises the steps: firstly carrying out the parameter setting and preprocessing, then segmenting a large-loop source into a north edge, a south edge, an east edge and a west edge, and carrying out the discretization of each edge into N electric dipoles through employing a Gaussian-Legendre quadrature method; constructing a loop edge-frequency sampling point-wave number-time channel four-level parallel computing architecture to compute each edge; wherein each side in the GPU adopts a frequency-wave number-time parallel computing architecture, and parallel computing is performed on each electric dipole, so that the frequency-time domain response of each electric dipole is obtained; and finally, after calculation of all the electric dipoles is completed, vectorization calculation and summation are carried out on transient responses of all the electric dipoles on the GPU, a large loop source total field response at the current observation point is obtained, and a final time domain electromagnetic field value is output. According to the method, the advantages of large-scale parallel computing of the GPU can be efficiently and fully played, so that the computing efficiency is effectively improved on the premise of ensuring the computing precision.
Owner:YUNLONG LAKE LAB OF DEEP UNDERGROUND SCI & ENG +1

Electric leakage compensation method of 2T0C in-memory calculation array

The invention discloses an electric leakage compensation method for a 2T0C in-memory calculation array, and belongs to the technical field of storage and calculation integration. A 2T0C main array and reference column driving mode is adopted, the compensation circuit is connected between the RBL of the 2T0C main array and the RBL of the reference column, the compensation circuit is used for dynamically adjusting the read voltage, calculation errors caused by the leakage effect can be effectively compensated, and IR-Drop is balanced by exerting symmetrical influence on calculation results of columns in the storage and calculation array. By adopting the method, stable calculation precision can be provided during large-scale parallel calculation, the system performance is remarkably improved, and the influence of a leakage effect on calculation accuracy is avoided.
Owner:SEMICON TECH INNOVATION CENT(BEIJING) CORP +1

GPU-accelerated fast forward method for transient electromagnetic method with large loop source

The application discloses a GPU acceleration-based large-loop-source transient electromagnetic fast forward method, which comprises the following steps: firstly, parameter setting and pretreatment are performed; then, the large-loop-source is divided into four edges of north, south, east and west, and each edge is discretized into N electric dipoles by using a Gauss-Legendre quadrature method; a four-level parallel computing architecture of loop edge-frequency sampling point-wave number-time channel is constructed to calculate each edge; wherein, the parallel computing architecture of frequency-wave number-time is used in each edge in the GPU to perform parallel calculation on each electric dipole, so that the frequency-time domain response of each electric dipole is obtained; finally, after the calculation of all electric dipoles is completed, the vectorization calculation and summation of the transient responses of all electric dipoles are performed on the GPU to obtain the total field response of the large-loop-source at the current observation point, and the final time domain electromagnetic field value is output. The above method can efficiently and fully exert the large-scale parallel computing advantage of the GPU, so that the calculation efficiency is effectively improved under the premise of ensuring the calculation precision.
Owner:YUNLONG LAKE LAB OF DEEP UNDERGROUND SCI & ENG +1

Video subtitle removing method, chip, GPU and electronic device

This disclosure relates to the field of computer technology, and in particular to a video subtitle removal method, chip, GPU, and electronic device. The method, executed on a graphics processing unit (GPU), includes: receiving an input video and decoding the input video into a frame sequence; detecting subtitle regions in the current frame of the frame sequence and generating a corresponding subtitle mask; responding to the presence of subtitles in the current frame, predicting the content of the subtitle region based on the subtitle mask and utilizing the spatiotemporal correlation of preceding and / or subsequent frames in the frame sequence; and filling the subtitle region with the predicted content to generate a video frame with subtitles removed. According to embodiments of this disclosure, by deploying a series of tasks such as video decoding, subtitle region detection, spatiotemporal correlation-based content prediction, and final filling onto the GPU, the massively parallel computing capabilities are utilized to accelerate the processing flow and improve the visual naturalness and spatiotemporal coherence of the repaired video.
Owner:MOORE THREADS TECH CO LTD

Parallel code automatic tuning method based on PPCG

The invention relates to a PPCG-based parallel code automatic tuning method, and belongs to the field of high-performance computing. According to the method, the tuning parameter dictionary is automatically generated and transmitted to the PPCG to serve as parallel code generation parameters, including tile size, block size and grid size, so that the execution performance of the generated parallel codes on a target hardware platform is optimized. According to the method, the computing power of the GPU can be fully utilized, in the application of large-scale parallel computing, the generation efficiency and execution performance of parallel codes can be improved, the workload of manual adjustment and optimization is reduced, and an automatic solution is provided for high-performance computing.
Owner:BEIJING INST OF COMP TECH & APPL

Neuron synchronization and multilayer core particle interconnection technology of brain-like inference chip

PendingCN121658258AInterprogram communicationArtificial lifeConcurrent computationInterconnect technology
The invention discloses a neuron synchronization and multilayer core particle interconnection technology of a brain-like reasoning chip, and belongs to the field of brain-like calculation. In order to solve the problems that in the prior art, neuron synchronization precision is insufficient, multi-layer core grain interconnection data transmission efficiency is low, and system integration optimization is insufficient, an integrated brain-like reasoning chip architecture is constructed by designing a self-adaptive neuron synchronization module, an efficient multi-layer core grain interconnection module, a reasoning acceleration module and a computing resource scheduling module. The framework realizes vertical integration of core grains based on a 2.5 D packaging technology, guarantees fine synchronization of neurons in large-scale parallel computing through a dynamic synchronization mechanism and accurate time sequence control, improves data transmission efficiency among the core grains, specially adapts to hardware requirements of brain-like reasoning tasks, and remarkably improves accuracy, speed and system throughput of neural network reasoning. The method can be widely applied to the technical fields of artificial intelligence reasoning, integrated circuit design, neuromorphic calculation, brain-computer interfaces and the like.
Owner:广西华芯振邦半导体有限公司