Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

109 results about "CUDA" patented technology

CUDA (Compute Unified Device Architecture) is a parallel computing platform and application programming interface (API) model created by Nvidia. It allows software developers and software engineers to use a CUDA-enabled graphics processing unit (GPU) for general purpose processing — an approach termed GPGPU (General-Purpose computing on Graphics Processing Units). The CUDA platform is a software layer that gives direct access to the GPU's virtual instruction set and parallel computational elements, for the execution of compute kernels.

CUDA kernel generation method and device based on reinforcement learning

The invention provides a CUDA kernel generation method and device based on reinforcement learning, and a CUDA kernel generation method and device based on reinforcement learning, which utilize a CUDA kernel generation and optimization model of reinforcement learning training and adopt a multi-round reinforcement learning training scheme. According to the method, the continuous optimization and iteration process of code trigger generation, execution and feedback followed by an engineer during development of the CUDA kernel is simulated, and performance feedback in a real execution environment can be fused into a model training process, so that the model autonomously explores an optimization path under a reinforcement learning mechanism, and automatic and intelligent generation of the high-performance CUDA kernel is realized.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Infrared bidirectional heat effect simulation method based on radiation intensity and GPU acceleration

The invention discloses an infrared bidirectional thermal effect simulation method based on radiation and GPU acceleration, and belongs to the field of thermal radiation simulation and parallel computing. Dispersing the complex geometry into patch units, constructing an inter-patch radiation energy balance equation based on a radiance theory, and iteratively solving the final temperature of the patches through a radiance method; a three-level parallel strategy of a task level, a data level and an instruction level is designed, data intensive tasks such as shape factor calculation and radiance iteration are migrated to a GPU to be executed, data storage is optimized in combination with a structure array (SoA) layout and a sparse matrix compression technology, and efficient data sharing of a CPU end and a GPU end is achieved through a CUDA zero copy technology. According to the method, the dependence of a traditional method on regular grids is broken through, the calculation efficiency and precision of million-level surface patch heat radiation transmission in a complex scene are remarkably improved, and an efficient tool is provided for heat radiation indirect transmission calculation and a global illumination model in the field of three-dimensional scene infrared simulation.
Owner:ZHEJIANG UNIV

Cooperative parallel memory allocation

Apparatuses, systems, and techniques to perform multi-threaded memory allocation in parallel by one or more software programs being performed on a parallel processing unit (PPU), such as a graphics processing unit (GPU), or any other processing unit capable of supporting multi-threaded software execution. In at least one embodiment, one or more software programs expressed in part by code using an application programming interface for parallel computing, such as CUDA, perform allocation, search, and deallocation of memory efficiently and in parallel on a GPU.
Owner:NVIDIA CORP

Space-based full-link photoelectric imaging simulation signal generation method based on CUDA (Compute Unified Device Architecture)

The invention provides a space-based full-link photoelectronic imaging simulation method based on CUDA (Compute Unified Device Architecture), and solves the problems of complicated calculation, insufficient simulation precision and high hardware requirements in the prior art. The method comprises the following steps: dividing the surface of a target three-dimensional model into triangular surface elements and calculating target intrinsic radiation scattering data of each surface element; generating background intrinsic radiation scattering data in real time through CUDA (Compute Unified Device Architecture) parallel calculation according to the space-time spectrum parameters and the earth surface / cloud layer / limb basic model library; calculating characterization radiation before entrance pupil based on target and background intrinsic data; the entrance pupil radiation is converted into a detector voltage signal through CUDA parallel optimization, and a gray infrared simulation image is generated. According to the method, the calculation efficiency is remarkably improved, and space-based full-link photoelectric imaging simulation with high authenticity and real-time performance is realized.
Owner:XIDIAN UNIV

Homomorphic encryption-based spatiotemporal big data distributed privacy computing system and device

The application provides a space-time big data distributed privacy computing system and equipment based on homomorphic encryption, and relates to the technical field of data processing. It comprises a space privacy computing operator, which is used for the computing characteristics of vector data and raster data, and realizes homomorphic encryption computing of vector data and raster data by using TFHE homomorphic encryption algorithm and CKKS homomorphic encryption algorithm respectively; a distributed space privacy computing framework is used, each node of the distributed space privacy computing framework is respectively provided with a space privacy computing operator, and the distributed space privacy computing framework realizes multi-node task parallelization and safe cooperation based on a key management mechanism under a distributed environment and a remote Boolean decryption protocol based on inadvertent transmission; each node in the distributed space privacy computing framework is provided with a GPU acceleration architecture, the GPU acceleration architecture is respectively provided with a corresponding CUDA heterogeneous computing optimization scheme for the Boolean logic operation of vector calculation and the matrix operation characteristics of raster calculation, and the processing efficiency of a single node is improved.
Owner:PEKING UNIV

High-fidelity lightweight world model construction method for end-to-end autonomous driving test

The application relates to a world model construction method, in particular to an end-to-end automatic driving test high-fidelity lightweight world model construction method, which constructs a high-fidelity world model, solves the problems of large world model parameters and low inference efficiency, performs knowledge distillation on the world model, reduces model parameters on the basis of retaining world model generation capacity, and improves inference efficiency; a CUDA operator is self-defined for a world model calculation bottleneck part, memory allocation is optimized, and a single-device multi-thread scheduling and multi-device cooperative calculation method are used to improve the inference efficiency of the high-fidelity world model. The application can construct an end-to-end automatic driving test high-fidelity lightweight world model, effectively solve the problems of low multi-modal information alignment accuracy, poor cross-view and cross-frame consistency, and low inference efficiency of the existing world model, improve the confidence of an end-to-end automatic driving system test process, and greatly accelerate the test efficiency of the end-to-end automatic driving system.
Owner:JILIN UNIVERSITY

Data processing method and system

Embodiments of the present specification provide a data processing method and system, the method comprising: a reasoning framework, in response to a cold start instruction sent by a cloud platform, cold starting on a target graphics processor, sending an opening hijacking command to a hijacking module, and during the process of loading first model weights of a first reasoning model, initiating a video memory application to the target graphics processor; the hijacking module, in response to the opening hijacking command, hijacking the video memory application, and redirecting the video memory application to a first virtual address in a fixed virtual address space; the reasoning framework, in a case where it is determined that the first model weights are completed loading, sending an ending hijacking command to the hijacking module, and based on the first model weights in the first virtual address, constructing a first reasoning execution graph to execute a first reasoning task using the first reasoning execution graph. By hijacking the CUDA video memory allocation and redirecting it to a pre-reserved fixed address, the address of the model weights is ensured to be persistent and stable, so that the CUDA Graph continues to be effective in multiple loadings and hot switching.
Owner:ALIBABA CLOUD COMPUTING CO LTD

Method and system for quickly encoding image in 3D engine

The invention provides a method and a system for quickly encoding an image in a 3D (three-dimensional) engine. The method for quickly encoding the image in the 3D engine comprises the following steps of: rendering the contents of a plurality of Cameras to the same RenderTexture, and distinguishing the positions of different Cameras; obtaining a bottom layer graphics library pointer of a RenderTexture object through a RenderTexture.GetNativeTexturePtr method provided by Unity, and obtaining a bottom layer graphics library pointer of the RenderTexture object according to the bottom layer graphics library pointer of the RenderTexture object and the bottom layer graphics library pointer of the RenderTexture object and the bottom layer graphics library pointer of the RenderTexture object and the bottom layer graphics library pointer of the RenderTexture object and the bottom layer graphics library pointer of the RenderTexture object. A graphic pointer is obtained through RenderTexture, a resource format is converted and mapped into a CUDA memory, a GPU is used for coding into a preset format, and a video memory is reused to reduce resource consumption. The system comprises modules corresponding to the steps of the method.
Owner:HUIZHIAN INFORMATION TECH CO LTD

GPU parallel thread block adjustment and kernel function scheduling optimization method and system

The invention discloses a GPU (Graphics Processing Unit) parallel thread block adjustment and kernel function scheduling optimization method and system. The method comprises the following steps: acquiring GPU equipment attributes and kernel function execution attributes; determining a feasible thread block size solution space of each kernel function based on GPU equipment attributes and kernel function execution attributes, and selecting a thread block size which enables the load balance degree evaluation model value to be maximum as optimal execution configuration of the kernel function; on the basis of the determined optimal thread block size of each kernel function, dividing an execution grid of each kernel function into a plurality of kernel function slices, and establishing a concurrent execution relationship between the kernel functions; a CUDA graph is constructed in a mixed mode according to the established concurrent execution relation between the kernel functions and data dependence; and by taking an actual model of particle simulation software as a test object, respectively comparing GPU occupancy rates before and after thread block adaptive adjustment, multi-stream concurrent scheduling and hybrid CUDA graph scheduling optimization, kernel function execution time and overall simulation efficiency, and outputting an optimized particle simulation result.
Owner:XI AN JIAOTONG UNIV

Dual-talking detection and acoustic echo cancellation co-processing method based on dual-cuda streams

The present application belongs to the technical field of speech processing, and particularly relates to a double-talking detection and acoustic echo cancellation cooperative processing method based on double CUDA streams. The method comprises the following steps: creating a first CUDA stream and a second CUDA stream on the same graphics processing unit; establishing a shared memory control package for cross-stream control between the first CUDA stream and the second CUDA stream; organizing a microphone near-end signal and a far-end reference signal into continuous frames with a preset frame length and frame shift; in the second CUDA stream, recording a synchronization event after writing sub-band control parameters and frame-level mode markers into the shared memory control package; in the first CUDA stream, performing partition blocking frequency domain sub-band adaptive filtering to update filter coefficients and generate echo suppression output. The present application realizes a highly cooperative, low-delay, high-robust real-time processing system for double-talking detection and echo cancellation, and has operation efficiency, control accuracy and engineering realizability, and is suitable for voice communication, remote conference and intelligent voice interaction systems.
Owner:CHINA NUCLEAR IND MAINTENANCE

Time-domain electromagnetic field simulation calculation method based on GPU parallel acceleration and corresponding product

The application relates to the technical field of electromagnetic simulation, and provides a time-domain electromagnetic field simulation calculation method based on GPU parallel acceleration and a corresponding product.The method comprises the following steps: setting simulation parameters; precalculating physical coordinates of all grid points, generating an observation point coordinate array and a source point coordinate array, and transmitting the arrays to GPU device memory; initializing a GPU solver, allocating GPU memory resources, and configuring a CUDA thread block and a grid structure, so that the number of threads matches the number of observation points; for each time step corresponding to a time step, source field data of the current time step is asynchronously updated into a ring buffer; when the simulation time exceeds the time required for electromagnetic waves to propagate from a source area to an observation area, a GPU kernel function without atomic operation is started, the electromagnetic field values of each observation point are calculated in parallel based on the precalculated coordinate array and the source field data in the ring buffer; and the calculation results are transmitted from the GPU device back to the host end for storage and post-processing verification.
Owner:ROCKET FORCE UNIV OF ENG

An image processing method, apparatus and electronic device

The application provides an image processing method, device and electronic equipment, the method comprising: dividing a to-be-processed image into a plurality of sub-blocks; processing the plurality of sub-blocks using a plurality of CUDA (Compute Unified Device Architecture) streams, each CUDA stream in the plurality of CUDA streams comprising: a calculation operation and a memory operation, and a sub-block in the plurality of sub-blocks being processed asynchronously between the calculation operation in a first CUDA stream and the memory operation in a second CUDA stream; and splicing and fusing the processing results of the plurality of sub-blocks to obtain a processing result of the to-be-processed image. In the implementation process of the above scheme, the multi-core resources of the GPU are fully utilized through the asynchronous processing between the calculation operation and the memory operation in different CUDA streams, the idle time of the processing unit is reduced, the resource utilization rate of the CUDA stream is maximized, and therefore the overall processing efficiency of the high-resolution image is improved.
Owner:北京天数智芯半导体科技有限公司

CUDA-based collision detection method and device, electronic equipment and storage medium

This invention relates to the field of autonomous driving technology, providing a CUDA-based collision detection method, apparatus, electronic device, and storage medium. The CUDA-based collision detection method includes: in response to a planned candidate trajectory and collected obstacle information, abstracting the candidate trajectory into a sequence of bounding boxes and the obstacle information into a set of obstacle points; processing the bounding box sequence and the obstacle point set in parallel using CUDA to determine matching pairs with potential collision risks, where each matching pair consists of a bounding box and an obstacle point; and in response to the obtained matching pairs, detecting the collision risk of each matching pair in parallel using CUDA to obtain the collision detection result of the candidate trajectory. This invention utilizes the parallel computing power of CUDA to greatly improve detection efficiency, and further improves the speed and accuracy of collision detection by first coarsely screening matching pairs with potential collision risks and then finely detecting the collision risks of each matching pair, achieving efficient and accurate collision detection of candidate trajectories.
Owner:SHANGHAI WESTWELL INFORMATION & TECH CO LTD

Method and system for large model training acceleration based on cuda shared memory and a pci express expansion card

The application provides a large model training acceleration method and system based on CUDA shared memory and a PCIe expansion card, hardware configuration information is acquired to identify the PCIe expansion card and build an expansion storage space, and a unified management heterogeneous memory space is formed in combination with GPU shared memory; in the training process, according to a preset initialization strategy, model data blocks are loaded to different levels of the heterogeneous memory, data access frequency is monitored in real time, data migration instructions are dynamically generated based on multiple frequency thresholds, and intelligent scheduling of the data blocks among the levels of storage is realized; a transparent address remapping mechanism is used to enable CUDA kernel functions to seamlessly access the heterogeneous memory space; periodic data snapshots and persistent storage are simultaneously supported, and quick recovery after interruption of training is ensured; the application breaks through the GPU display memory capacity limit, improves the large model training efficiency, reduces the dependence on multi-card hardware and cost, and guarantees data security and training reliability.
Owner:SHENZHEN QUANXING TECH CO LTD

GPU resource dynamic allocation method, system and device and storage medium

The invention discloses a GPU resource dynamic allocation method, system and device and a storage medium, and relates to the technical field of GPU resource management, and the method can accurately grasp the use condition of the current resource through collecting the resource state information and the load condition in real time, and provides an accurate basis for resource allocation. And on the basis of a CUDA API detection mechanism, the actual demand of a task on GPU resources can be accurately judged, and invalid locking and waste of the resources are avoided. And resource allocation is carried out according to the preset dynamic allocation strategy, so that the rationality and high efficiency of resource allocation are ensured. And the GPU resources are released in time when the preset release condition is met, so that the utilization rate and the circulation efficiency of the resources are further improved. In this way, the utilization rate of GPU resources can be remarkably increased, and the problems of GPU resource recovery and re-circulation can be effectively solved.
Owner:INNER MONGOLIA ELECTRIC POWER (GRP) CO LTD DIGITAL RES BRANCH

Hybrid expert model multi-lexical element prediction method and device based on dense-sparse parallel computing, medium, program product and terminal

According to the hybrid expert model multi-lexical element prediction method and device based on dense-sparse parallel calculation, the medium, the program product and the terminal provided by the invention, parallel execution is realized in a Tensor Core and a CUDA Core through a dense-sparse collaborative scheduling mechanism, the problem of low hardware utilization rate caused by expert activation fragmentation when MTP and MoE are combined is solved, and the prediction efficiency of the hybrid expert model multi-lexical element prediction method and device based on dense-sparse parallel calculation is improved. The hardware efficiency and throughput performance of MTP and MoE joint reasoning are obviously improved; according to the method, the structure of the MoE model does not need to be modified or additionally trained, high-performance reasoning can be realized only through optimization in the operation process, and the method is suitable for a production scene in which the model cannot be modified; according to the dynamic pruning strategy based on the activation frequency and the gating weight, redundant calculation is effectively eliminated, the model reasoning speed is remarkably increased while the precision of the output result is guaranteed, and the stability and reliability of the model generation quality are ensured.
Owner:SHANGHAI JIAOTONG UNIV

Rule fast lookup method, device and computer equipment based on CUDA

The application relates to a CUDA-based rule fast searching method, device and computer equipment, wherein the CUDA-based rule fast searching method comprises the following steps: obtaining a constant string of each rule, and compiling the constant string through a string matching algorithm to obtain a first Boolean array of single non-repeated characters in the constant string; further, obtaining a to-be-matched string of to-be-matched data, and creating a second Boolean array based on the to-be-matched string; and based on CUDA, performing multi-thread matching processing on the first Boolean array and the second Boolean array; when the to-be-matched string contains all the constant strings, the rule and the data are matched successfully. Through the application, the problem that data cannot be quickly matched with rules is solved, and the efficiency of rule searching is improved.
Owner:HANGZHOU ANHENG INFORMATION SECURITY TECH CO LTD

Gpu-based parallel spectral delay correction method and apparatus

This application provides a GPU-based parallel spectral delay correction method and apparatus, applicable to the field of artificial intelligence technology. The method includes: acquiring multiple parallel units; determining the current parallel unit in chronological order, wherein all time sub-intervals within the current parallel unit simultaneously undergo multiple iterative correction scans through their respective CUDA blocks until a stopping iteration condition is met; after completing multiple iterative correction scans, serially compensating the time sub-intervals within the parallel unit to obtain the accurate solution for the current parallel unit; and sending the accurate solution of the last time sub-interval in the current parallel unit to the next parallel unit. The correction scan process includes: combining the solution calculations of the computing nodes in the current parallel unit into a batch; calling a batch solver to perform batch calculations to obtain an initial solution for each computing node; and performing parallel batch refinement based on the initial solution through the thread group corresponding to each computing node. This application improves GPU utilization.
Owner:北京天数智芯半导体科技有限公司

An AI multi-agent and digital twin fusion production process visualization method, medium and system

The application provides an AI multi-agent and digital twin fusion production scheduling process visualization method, medium and system, belonging to the technical field of AI multi-agent production scheduling. Large-scale parallel computing is realized by constructing a GPU three-layer CUDA processing architecture. The first layer of data preprocessing grid performs data cleaning in parallel. The second layer of negotiation analysis grid runs a Transformer-based agent interaction recognition model for parallel pattern recognition. The third layer of visualization calculation grid runs an LSTM-CNN fusion production scheduling process mapping model to generate visualization data. The multi-head attention mechanism dynamically adjusts the allocation of computing resources. The pipeline parallel and data parallel strategies optimize the computing performance. The asynchronous data transmission and computing overlap technology hides the memory access delay, solving the technical problem that multi-agent high-frequency negotiation data cannot be processed in real time.
Owner:BEIJING NANCAL RUIYUAN DIGITAL TECH CO LTD

AI compiler

The embodiment of the invention provides an AI compiler, and relates to the field of communication. The AI compiler comprises a high-level processing layer which is used for converting a CUDA C Kernel code into an NVVM dialect and converting a non-CUDA C Kernel code into a non-NVVM dialect; the middle-end conversion layer is integrated with a first conversion module and a second conversion module, and is used for converting the NVVM dialect into an RISCV vector dialect or a matrix dialect based on the first conversion module and converting the non-NVVM dialect into an RISCV vector dialect or a matrix dialect based on the second conversion module; and the low-level processing layer is used for converting the RISCV vector dialect or the matrix dialect into a pure machine code file, so that the pure machine code file runs on an RISCV chip. Through the embodiment of the invention, the problem that the executable file translated by an AI compiler in the related technology can only run in a non-open source chip generally is solved, and the effect of enabling the generated executable file to run in an open source chip is further achieved.
Owner:SANECHIPS TECH CO LTD

Sequence parallelization method and device of attention module

The invention discloses a sequence parallelization method and device of an attention module. The method comprises the following steps: receiving a block sparse attention module and hardware parameters; carrying out load division on the block sparse attention module based on the hardware parameters to generate a parallel dependency graph; performing operator-level context tiling processing on the parallel dependency graph to obtain a transformed parallel dependency graph; and carrying out operator modeling on the transformed parallel dependency graph based on preset continuous execution time of each operator to generate a CUDA flow scheduling graph. According to the method, in the calculation process of a large language model, load balanced distribution and operator-level parallel scheduling of the block sparse attention module between cross-calculation nodes and cross-graphics processor equipment can be realized. On the premise of ensuring the correctness of the calculation and communication dependency relationship, the collaborative execution efficiency of calculation and communication operation is improved, and the overall operation performance and the resource utilization rate in the multi-graphics processor cluster environment are remarkably optimized.
Owner:TSINGHUA UNIVERSITY

Shipbuilding gantry crane cart anti-collision protection system and method based on video recognition

The invention discloses a shipbuilding gantry crane cart anti-collision protection system and method based on video recognition, and belongs to the field of industrial automation and safety control. The camera transmits an encoded video stream to the algorithm server through an RTSP protocol, and the FFmpeg pulls the video stream and carries out hard decoding to obtain data as algorithm input; the CUDA kernel function completes preprocessing of video frames; a CUDA parallelization processing algorithm and a TensorRT loading INT8 quantitative model are used for executing rapid reasoning; an alarm area is selected on a video picture through a configuration file, and once foreign matters appear in the area, an algorithm sends alarm information to a PLC through the Ethernet within 200 ms; the PLC judges whether collision danger occurs or not according to the direction of the cart, and the corresponding area is prompted through the voice broadcaster. According to the system, precise recognition and millisecond-level active braking of small obstacles are achieved.
Owner:DALIAN COSCO KHI SHIP ENG

A multi-event face search system and method based on GPU parallel acceleration and milvus

The application discloses a multi-competition face searching system and method based on GPU parallel acceleration and Milvus, and relates to the technical field of face recognition and vector retrieval. The system is divided into an access layer, a service scheduling layer, a GPU inference engine layer and a three-level storage layer; GPU batch parallel face detection and 512-dimensional feature extraction are realized by relying on CUDA; video frame feature batch deduplication is completed by using high-speed operation of the GPU; Milvus adopts competition_id as a partition key to realize multi-competition data partition isolation, and is matched with Milvus-GPU to accelerate partition vector retrieval; the service scheduling layer dynamically divides GPU computing power priority, and preferentially guarantees real-time retrieval resources; the system configures hierarchical API keys and IP white lists to realize safety management and control, and is compatible with watermark storage and extra_data parameter transparent transmission functions. The application overcomes the defects of low efficiency of traditional CPU serial processing, large video storage redundancy and high full-database retrieval delay, has the advantages of fast processing speed, high hardware utilization rate, small storage cost and strong safety, and is suitable for competition personnel verification, park visitor retrieval, security blacklist control and other scenes.
Owner:JIANGSU CAMBRIAN INFORMATION TECHNOLOGY CO LTD

A flood evolution analysis method based on instant compiling technology

The application discloses a flood evolution analysis method based on instant compiling technology, mainly comprising the following steps: setting a hardware acceleration environment, specifying a calculation time and an iteration step length, loading a grid topological structure, an initial water level and an initial flow velocity, a grid roughness field and model boundary conditions, optimizing codes of a finite volume method and a HLLC shallow water equation solving algorithm based on instant compiling technology, and realizing parallel acceleration simulation of flood evolution. The method is written in Python language with strong interpretability to compile shallow water equation solving codes, and machine codes are generated by instant compiling technology during runtime to greatly improve code running efficiency. Different hardware platforms such as CPU, GPU or CUDA can be selected to perform parallel acceleration solving on the shallow water equation, simple and efficient simulation of flood evolution is realized, and the method has the advantages of strong code readability and high performance.
Owner:NANJING INST OF GEOGRAPHY & LIMNOLOGY

A CUDA acceleration-based large three-dimensional image morphing method

The present application relates to a kind of based on CUDA acceleration's super large three-dimensional image morphing method, comprising: the affine transformation matrix of two sets of registration feature points is calculated by affine transformation;Target image is blocked, and the image morphing of each small block image block obtained after blocking is carried out;Utilize affine transformation matrix, realize three-dimensional image voxel traversal parallelization operation by CUDA;Read the 3D image block after morphing, splice into the 2D continuous sequence image of original image size.The present application utilizes the principle of target image blocking, and utilizes the parallel architecture of GPU to accelerate the calculation process of image linear interpolation, while solving can be for super large three-dimensional image morphing, also effectively improves the running efficiency of image registration;Experiment proves that this method can process more than twice the computer memory super large three-dimensional image data under the platform of CPU as i5-10400F, GPU as RTX-3060, memory as 16G, and about 11 times speedup ratio is obtained, greatly shorten the time of morphing registration.
Owner:ANHUI UNIV

DEM sunshine duration method and system based on GPU and terrain entropy optimization

The invention discloses a DEM (Digital Elevation Model) sunshine duration method and system based on GPU (Graphics Processing Unit) and terrain entropy optimization, and belongs to the technical field of geographic information science and space calculation. The method comprises the following steps: firstly, calculating topographic information entropy by using digital elevation model data to quantify topographic complexity of different regions; and the terrain information entropy is used as a self-adaptive driving factor to dynamically adjust the search radius and the search step length of the sun visible range, so that the condition that no shielding exists in a terrain gentle region is quickly judged, and the terrain shielding relationship is finely simulated in a mountain region or a complex terrain region. Besides, the method is realized by adopting a CUDA (Compute Unified Device Architecture) parallel architecture, and is combined with a GPU (Graphics Processing Unit) parallel strategy of partitioning and partitioning, so that the calculation load of high-resolution data is effectively reduced, and the overall operation efficiency and the memory utilization rate are improved.
Owner:ZHEJIANG UNIV

Graphics processor based multi-satellite signal acquisition method

The application provides a multi-satellite signal acquisition method based on a graphics processor. Through the asynchronous pipeline architecture of CPU-GPU cooperation, after completing parameter configuration and data asynchronous uploading at the central processor end, the graphics processor end divides multiple CUDA streams based on the parallel subtasks of the CUDA stream based on the Doppler dimension, combines the pre-bound Fourier transform plan, accelerates the down-conversion through the frequency lookup table, performs frequency domain correlation and non-coherent accumulation in the code phase domain on multiple CUDA streams in parallel, and finally extracts the local peak value through one-level reduction and returns it asynchronously. The central processor end summarizes the results and judges the effectiveness based on the carrier-to-noise ratio. The design realizes the deep integration of the Doppler task-level pipeline and the code phase thread-level parallelism, avoids non-merged access of the GPU device memory, significantly improves the efficiency of multi-constellation signal acquisition, and supports multi-frequency point signal processing of GPS, Beidou and the like through a unified pre-configuration interface, thereby enhancing the architecture universality.
Owner:SUN YAT SEN UNIV

Graphics processor based multi-satellite signal acquisition method

ActiveCN122111684BTerm memoryNon coherent
The application provides a multi-satellite signal acquisition method based on a graphics processor. Through the asynchronous pipeline architecture of CPU-GPU cooperation, after completing parameter configuration and data asynchronous uploading at the central processor end, the graphics processor end divides multiple CUDA streams based on the parallel subtasks of the CUDA stream based on the Doppler dimension, combines the pre-bound Fourier transform plan, accelerates the down-conversion through the frequency lookup table, performs frequency domain correlation and non-coherent accumulation in the code phase domain on multiple CUDA streams in parallel, and finally extracts the local peak value through one-level reduction and returns it asynchronously. The central processor end summarizes the results and judges the effectiveness based on the carrier-to-noise ratio. The design realizes the deep integration of the Doppler task-level pipeline and the code phase thread-level parallelism, avoids non-merged access of the GPU device memory, significantly improves the efficiency of multi-constellation signal acquisition, and supports multi-frequency point signal processing of GPS, Beidou and the like through a unified pre-configuration interface, thereby enhancing the architecture universality.
Owner:SUN YAT SEN UNIV