Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

214 results about "CUDA" patented technology

CUDA (Compute Unified Device Architecture) is a parallel computing platform and application programming interface (API) model created by Nvidia. It allows software developers and software engineers to use a CUDA-enabled graphics processing unit (GPU) for general purpose processing — an approach termed GPGPU (General-Purpose computing on Graphics Processing Units). The CUDA platform is a software layer that gives direct access to the GPU's virtual instruction set and parallel computational elements, for the execution of compute kernels.

GPU computing power scheduling method based on one-cloud multi-core heterogeneous computing power platform

The invention provides a GPU computing power scheduling method based on a one-cloud multi-core heterogeneous computing power platform, and the method comprises the following steps: S1, carrying out heterogeneous resource registration and modeling, accessing hardware equipment containing multiple types of GPUs through a resource registration module, collecting the equipment model, the video memory capacity and performance index data, and carrying out heterogeneous resource modeling; constructing a resource feature database containing a topological relation, and supporting hybrid access of chips; s2, virtualized resource reconstruction: pooling a physical GPU into virtual GPU resources by adopting a hardware abstraction layer technology, realizing video memory isolation and calculation unit division through a containerization technology, and configuring each virtual GPU instance with an independent drive stack and a security sandbox; and S3, submitting a multi-modal task, receiving a CUDA / OpenCL calculation task submitted by a user, analyzing task demand parameters including a calculation core number, a video memory occupation amount and a data throughput threshold, and generating a task descriptor containing a priority label.
Owner:SAISI TECH (XIAN) CO LTD

Method and system for realizing CUDA call tracking based on eBPF

InactiveCN120723587AHardware monitoringCall tracingSoftware engineering
The invention provides a method and system for achieving CUDA call tracking based on an eBPF, and relates to the field of GPU call tracking of the eBPF. The method comprises the following steps: dynamically mounting an eBPF program at an API function entry and an API function exit of a user mode CUDA runtime library through an uprobe technology of the eBPF; the eBPF program captures metadata of the CUDA API calling event in a kernel mode and packages the metadata into a predefined structural body; asynchronously transmitting the structural body data to a user mode through an annular buffer area; and calling a standard function library by a user mode program to load an eBPF program, reading data in the annular buffer area in real time, and completing analysis and visual display of a CUDA calling event. According to the method, the calling parameter, the return value and the nanosecond timestamp are effectively captured, and the observability and the debugging efficiency of the GPU application program are improved.
Owner:JINAN INSPUR DATA TECH CO LTD

Image-fused end-side cloud collaborative intelligent fire-fighting fire monitoring system

The invention discloses an end-side cloud collaborative intelligent fire-fighting fire monitoring system based on image fusion, and relates to the technical field of intelligent fire-fighting, the system is composed of a plurality of functional modules, and the system comprises a multi-modal image fusion module which generates a dynamic scanning priority map based on prior data, distinguishes a natural heat source from an abnormal fire by using a dual-light fusion algorithm, and sends an image fusion result to a cloud server; a scanning area is divided according to the thermal risk grade, and the thermal imaging resolution is dynamically adjusted; the distributed edge computing module is used for carrying out space-time synchronization on cross-modal data through a multi-modal feature alignment network, and carrying out dynamic allocation on a CUDA core and CPU resources through adaptive computing scheduling; an improved artificial bee colony algorithm is adopted, the bandwidth of the multi-sensor data flow is dynamically allocated through a time-sharing multiplexing protocol, and three-dimensional path planning is carried out; and the end-side cloud collaborative decision module constructs a federated learning driven model sharing network, and each edge node trains a lightweight YOLOv5s pruning model based on local data.
Owner:HANGZHOU ZIPENG TECH CO LTD

Method and system for generating and optimizing CUDA (Compute Unified Device Architecture) code based on multi-dimensional feature search and enhancement

The invention provides a CUDA code generation and optimization method and system for multi-dimensional feature search and enhancement, and the method comprises the steps: generating an initial CUDA code through a large language model based on task description and GPU hardware parameters; executing triple verification and a feedback mechanism on the initial CUDA code; according to feedback information of triple verification, optimizing CUDA codes through multi-dimensional feature search; based on error information, adjusting large language model input to regenerate codes; dynamically selecting an optimization strategy based on a performance index, wherein the optimization strategy comprises a thread block size and a memory access mode; and iteratively executing the steps until the CUDA code which passes triple verification and meets the target performance is generated. According to the method, the code features of the CUDA code are analyzed, and a triple verification and feedback mechanism including compilation feasibility, logic correctness and performance during execution is formed, so that an automatic evaluation process of the CUDA code is realized, and the problems that the CUDA code optimization process is difficult to be automated and an optimization target is lacked in the CUDA code optimization process are solved.
Owner:SHANGHAI JIAOTONG UNIV

High-fidelity lightweight world model construction method for end-to-end automatic driving test

The invention relates to a world model construction method, in particular to a high-fidelity lightweight world model construction method for an end-to-end automatic driving test, which is used for constructing a high-fidelity world model and performing knowledge distillation on the world model to solve the problems of huge parameters, low reasoning efficiency and the like of the world model. Model parameters are reduced on the basis of reserving the world model generation capability, and the reasoning efficiency is improved; a CUDA operator is developed in a user-defined mode for the bottleneck part of world model calculation, video memory allocation is optimized, and the reasoning efficiency of a high-fidelity world model is improved based on a single-device multi-thread scheduling and multi-device cooperative calculation method. According to the method, the high-fidelity lightweight world model for the end-to-end automatic driving test can be constructed, the problems that an existing world model is low in multi-modal information alignment precision, poor in cross-view and cross-frame consistency, low in reasoning efficiency and the like are effectively solved, the confidence coefficient of the end-to-end automatic driving system test process is improved, and the test efficiency of the end-to-end automatic driving system is improved. And the testing efficiency of the end-to-end automatic driving system is greatly improved.
Owner:JILIN UNIVERSITY

Protocol-Aware Provisioning of Resource over CXL Fabrics

Dynamic provisioning of resources in a datacenter improves the utilization efficiency of compute, memory, storage, and network resources, while maintaining flexibility to meet changing demands. Embodiments herein disclose protocol-aware provisioning of resources over Compute Express Link (CXL) fabrics, enabling multi-protocol pooling and disaggregation of resources, including memory and workload-specific accelerators such as GPUs and DSAs. In some embodiments, a Resource Provisioning Unit (RPU) facilitates intent-based protocol translations and mappings between address spaces, potentially enabling the creation of large-scale compute-memory fabrics that may utilize both coherent and non-coherent Non-Transparent Bridging (NTB) between multiple protocols in a single system, serving as the underlying infrastructure for executing workloads such as Large-Language Models (LLMs) and Deep Learning Recommendation Models (DLRMs). Some embodiments also optimize low-latency communication between processes running on different nodes by enabling host-to-host memory provisioning for libraries such as OpenMP, Pthreads, or CUDA, which can utilize shared memory.
Owner:HYATT GAYA OPAL MS +1

Multi-view collaborative 3D Gaussian splash optimization method and system

The invention discloses a multi-view collaborative 3D Gaussian splash optimization method and system, and the method comprises the steps: constructing a multi-level heterogeneous video memory pool, and dynamically dividing a video memory in 3D Gaussian reconstruction into a view exclusive memory block and a global shared memory pool; a mixed rendering-gradient pipeline is designed, and hardware-level pipeline parallelism in 3D Gaussian reconstruction is realized through a double-buffer asynchronous switching mechanism based on a CUDA Warp-level parallel primitive fusion forward rendering and back propagation thread group; performing multi-view gradient joint optimization, screening an effective gradient path in 3D Gaussian reconstruction based on the visibility mask matrix, and performing projection error weighted fusion on a multi-view gradient tensor; and implementing a multi-modal densification decision, generating a 3D Gaussian candidate splitting position in 3D Gaussian reconstruction through Monte Carlo sampling, calculating a joint optimization objective function by combining a multi-view projection residual error and a gradient contribution factor, and finally realizing 3D Gaussian reconstruction. According to the invention, high-precision and low-delay large-scale scene real-time rendering and training can be realized.
Owner:ZHEJIANG UNIV

CUDA kernel generation method and device based on reinforcement learning

The invention provides a CUDA kernel generation method and device based on reinforcement learning, and a CUDA kernel generation method and device based on reinforcement learning, which utilize a CUDA kernel generation and optimization model of reinforcement learning training and adopt a multi-round reinforcement learning training scheme. According to the method, the continuous optimization and iteration process of code trigger generation, execution and feedback followed by an engineer during development of the CUDA kernel is simulated, and performance feedback in a real execution environment can be fused into a model training process, so that the model autonomously explores an optimization path under a reinforcement learning mechanism, and automatic and intelligent generation of the high-performance CUDA kernel is realized.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Space-space-based infrared imaging simulation platform based on CUDA (Compute Unified Device Architecture)

The invention relates to an air-space-based infrared imaging simulation platform based on CUDA (Compute Unified Device Architecture). The air-space-based infrared imaging simulation platform comprises a global calculation class, a three-dimensional position calculation module, a camera generation module, a sun calculation module, a background data extraction module, a target calculation module and an imaging effect calculation module. The platform meets the requirement of high-precision signal-level space-space-based infrared detection imaging simulation, the simulation image data generation efficiency is greatly improved under the condition that the signal-level imaging simulation precision in a space scene is guaranteed, and rapid generation of imaging simulation data of a space platform joint observation target in a multi-observation track scene is achieved. And the universality of the platform is improved by using modular data and development.
Owner:XIDIAN UNIV

Strong real-time signal processing method based on multi-stream processing

The invention discloses a strong real-time signal processing method based on multi-stream processing, and the method comprises the steps: carrying out the remainder of a preset stream number of received coherent processing interval data through a CPU according to a CPI number, obtaining a stream number, distributing the coherent processing interval data to each stream distributed calculation program corresponding to the stream number, each stream distributed computing program comprises a CUDA stream and a CPU processing thread which are executed in sequence; the GPU and the CPU execute each stream distributed calculation program to overlap the calculation time and the data copying time of the CUDA stream and the CPU processing thread to obtain each trace point condensation data with different completion time; the CPU obtains the corresponding report to be sent according to the trace point condensation data, and orderly sends the report to be sent to the target program based on a report waiting synchronization mechanism, and the application can complete front-end downloading data receiving, detection processing and report sending in strong real time, thereby improving the development efficiency, shortening the product development period, and reducing the project cost.
Owner:CNGC INST NO 206 OF CHINA ARMS IND GRP

Method and system for selecting optimal network topology structures of different types of large models

The invention discloses an optimal network topology structure selection method and system for different types of large models, and belongs to the technical field of large models. The method comprises the following steps: expanding a search object into a large model, selecting the large model, expanding an Embedding operator for each selected large model, and adding a data processing dimension; determining several kinds of candidate cluster topological structures, and performing training time simulation on each candidate cluster topological structure on a coarse granularity level; performing time measurement on the key operator by using CUDA (Compute Unified Device Architecture), wherein the time measurement is used as an input parameter in a simulation process; the simulation training time under different cluster topological structures is compared, and whether an obvious topology selection tendency exists or not is determined according to a simulation result; and selecting the cluster topological structure with the shortest training time as an optimal solution. The topological structure can provide the fastest training speed under a given large model and a data processing task; and the efficiency of the selection process is greatly improved.
Owner:SHENZHEN INST OF ADVANCED TECH

Homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multi-thread mapping

The invention provides a homomorphic encryption acceleration method based on polynomial multiplication optimization and GPU multi-thread mapping, and relates to the field of computer science and technology. The method comprises the following steps: on the basis of a CUDA-GPU structure model, according to parallel computing characteristics of a GPU, adopting a nested loop optimization method to perform loop optimization on an algorithm structure based on NTT transformation and an algorithm structure based on INTT transformation respectively to obtain a loop optimization algorithm based on NTT transformation and a loop optimization algorithm based on INTT transformation, so as to obtain a loop optimization algorithm based on NTT transformation and a loop optimization algorithm based on INTT transformation; through a BFV primitive root pre-calculation and pre-storage joint optimization algorithm, homomorphic multiplication in GPU parallel calculation environment homomorphic encryption operation is improved, and improved homomorphic multiplication is obtained; inputting two pieces of ciphertext data to be processed; and processing the two pieces of to-be-processed ciphertext data by adopting improved homomorphic multiplication to obtain a processed ciphertext result. By adopting the method, the polynomial multiplication execution efficiency of the GPU parallel computing environment can be improved.
Owner:UNIV OF SCI & TECH BEIJING +1

GPU-CUDA-based SPH pollutant transport simulation parallel computing method

The invention provides an SPH pollutant transport simulation parallel computing method based on GPU-CUDA. The SPH pollutant transport simulation parallel computing method based on GPU-CUDA comprises the following steps that firstly, parallel computing parameters are initialized, and particle information is initialized; 2, copying necessary data from the CPU end to the GPU end; step 3, carrying out parallel computing of the SPH method, wherein the parallel computing specifically comprises the following sub-steps: dividing background grids; particle sorting: calculating a grid where the particles are located according to the position information of the particles, sorting all the particles according to spatial positions, and ensuring that adjacent particles are continuously stored in a memory; a multi-thread particle search strategy: particle attribute updating: updating the position and speed of the particle, and completing simulation of the current time step; processing boundary conditions; iteration is repeated until the simulation end time is reached; 4, after simulation is finished, result data are copied back to the CPU end from the GPU end.
Owner:TIANJIN UNIV

Lightweight 3D medical image real-time reasoning method and system based on edge calculation

The invention relates to the technical field of image data processing, and particularly discloses a lightweight 3D medical image real-time reasoning method and system based on edge computing, and the method comprises the steps: building a C + + frame model, carrying out anisotropy detection and processing, separating low-resolution axis resampling, carrying out 3D resampling, determining the step length of a sliding window, and generating a weight map. And performing prediction, sliding window reasoning and image post-processing by using a lightweight student model, performing processing including size adjustment, voxel communication and integrity filling on a predicted segmentation map obtained through reasoning, and outputting a final global segmentation map. According to the method, the lightweight student model is adopted to replace the traditional nnUNet model reasoning which needs five times, the reasoning frequency is reduced by 80%, and the single reasoning time is remarkably shortened; meanwhile, CUDA parallelization is achieved through C + + language reconstruction, and through data block extraction, mirror image transformation and GPU acceleration of a weighted aggregation operator, the whole process performance is improved by more than 10 times compared with Python implementation.
Owner:XIAOZHI FUTURE (CHENGDU) TECH CO LTD

Intelligent computing center computing power asymmetric collaborative scheduling method oriented to common computing power

The invention provides an asymmetric collaborative scheduling method for computing power of an intelligent computing center oriented to common computing power, and relates to the technical field of intelligent computing centers, intelligent computing centers and computing power infrastructure, the method comprises the following steps: S1, obtaining operation parameters corresponding to each computing power operation task in a plurality of computing power operation tasks in the intelligent computing center, the operation parameters comprise a streaming processor (SM) utilization rate and other parameters, and the other parameters comprise at least one of a video memory utilization rate, a unified computing device architecture (CUDA) core activeness, a memory occupancy rate and a bandwidth occupancy rate; s2, dividing the plurality of computing power operation tasks based on the operation parameters to obtain at least one group; and S3, scheduling the computing power operation tasks included in different groups to different acceleration cards of the intelligent computing center, wherein each acceleration card of the intelligent computing center is used for executing the computing power operation tasks in one group. The resource utilization rate can be greatly improved, the computing power leasing cost is greatly reduced, and wide application of the general computing power is achieved.
Owner:DATACANVAS LTD

River water and sediment transportation process numerical simulation method based on multi-GPU parallel framework

The invention discloses a numerical simulation method for a river water and sediment transportation process based on a multi-GPU parallel framework. The method specifically comprises the following steps: dividing a calculation region into triangular grid calculation regions with the same number as GPUs; the variables are initialized; creating two CUDA flows to control parallel calculation of the two GPUs, and storing data of each river channel into the corresponding GPU; directly carrying out data communication between the GPUs; in each GPU, calculating a source item and flux of each triangular grid to obtain a water depth value, and calculating a water level value of each grid according to the water depth; according to the obtained water level value, water level result data calculated in each GPU are stored in a CPU to complete water level data merging; and calculating to obtain the water level values of each river reach and the outlet section of the river at each moment. According to the river sediment transportation process simulation method based on triangular grid multi-GPU parallel computing, the water and sediment transportation process of the river area can be efficiently simulated, and the water level value of the section can be rapidly predicted.
Owner:INST OF EARTH ENVIRONMENT CHINESE ACAD OF SCI

Infrared bidirectional heat effect simulation method based on radiation intensity and GPU acceleration

The invention discloses an infrared bidirectional thermal effect simulation method based on radiation and GPU acceleration, and belongs to the field of thermal radiation simulation and parallel computing. Dispersing the complex geometry into patch units, constructing an inter-patch radiation energy balance equation based on a radiance theory, and iteratively solving the final temperature of the patches through a radiance method; a three-level parallel strategy of a task level, a data level and an instruction level is designed, data intensive tasks such as shape factor calculation and radiance iteration are migrated to a GPU to be executed, data storage is optimized in combination with a structure array (SoA) layout and a sparse matrix compression technology, and efficient data sharing of a CPU end and a GPU end is achieved through a CUDA zero copy technology. According to the method, the dependence of a traditional method on regular grids is broken through, the calculation efficiency and precision of million-level surface patch heat radiation transmission in a complex scene are remarkably improved, and an efficient tool is provided for heat radiation indirect transmission calculation and a global illumination model in the field of three-dimensional scene infrared simulation.
Owner:ZHEJIANG UNIV

Lightweight CNN (Convolutional Neural Network) image classification method and device for assisting rapid agricultural detection

The invention discloses a lightweight CNN (Convolutional Neural Network) image classification method and a lightweight CNN image classification device for assisting agricultural rapid detection. The classification method comprises the following steps: step 1, designing a crop disease and insect pest detection network; 2, designing a patch merging module; 3, designing a convolution block module; 4, preparing an experimental data set; 5, testing the performance of the model; step 6, model performance index determination; and 7, comparing the model parameter quantity with the calculated quantity. The lightweight CNN image classification device is composed of an edge calculation host, a high-precision image acquisition system and a high-speed communication module, and a complete edge calculation solution is formed. The edge computing host adopts NVIDIA JetsonAGX Orin as a core computing unit, integrates a 12-core ARM Cortex-A78AE CPU (central processing unit) and an Ampere architecture GPU (graphics processing unit) with 2048 CUDA (compute unified device architecture) cores, is equipped with a 64GB LPDDR5 memory, and has the beneficial effects that efficient detection and accurate recognition are realized; edge calculation is adaptive, low in consumption and high in efficiency; a data fusion and model optimization mechanism; and light weight and hardware adaptive design are realized, so that good practicability and adaptability are realized.
Owner:CHANGCHUN INST OF TECH

AI multi-agent and digital twinborn fusion scheduling process visualization method, medium and system

The invention provides an AI multi-agent and digital twinborn fusion production scheduling process visualization method, medium and system, and belongs to the technical field of AI multi-agent production scheduling. Large-scale parallel computing is achieved by constructing a GPU three-layer CUDA processing architecture, a first layer data preprocessing grid executes data cleaning in parallel, a second layer data preprocessing grid executes data cleaning in parallel, and a third layer data preprocessing grid executes data cleaning in parallel; the second layer of negotiation analysis grid operates an agent interaction recognition model based on Transform to perform parallel mode recognition, the third layer of visual calculation grid operates an LSTM-CNN fused production scheduling process mapping model to generate visual data, and calculation resource allocation is dynamically adjusted through a multi-head attention mechanism. The calculation performance is optimized by adopting a pipeline parallel and data parallel strategy, and the memory access delay is hidden by combining an asynchronous data transmission and calculation overlapping technology, so that the technical problem that the real-time parallel processing of the multi-agent high-frequency negotiation data cannot be realized is solved.
Owner:BEIJING NANCAL RUIYUAN DIGITAL TECH CO LTD

Cooperative parallel memory allocation

Apparatuses, systems, and techniques to perform multi-threaded memory allocation in parallel by one or more software programs being performed on a parallel processing unit (PPU), such as a graphics processing unit (GPU), or any other processing unit capable of supporting multi-threaded software execution. In at least one embodiment, one or more software programs expressed in part by code using an application programming interface for parallel computing, such as CUDA, perform allocation, search, and deallocation of memory efficiently and in parallel on a GPU.
Owner:NVIDIA CORP

CUDA (Compute Unified Device Architecture) and PyTorch mixed compiled ground penetrating radar rapid full waveform inversion method

The invention is suitable for the technical field of ground penetrating radar exploration, and provides a CUDA (Compute Unified Device Architecture) and PyTorch mixed compiled ground penetrating radar rapid full waveform inversion method, a PyTorch automatic differential function is self-defined by innovatively carrying out joint compilation on PyTorch and CUDA kernel functions, and a PyTorch platform is combined with CUDA mixed programming to realize automatic updating of two parameters of dielectric constant and conductivity. According to the method, GPU parallel calculation is used, the calculation efficiency of the ground penetrating radar data full-waveform inversion technology is remarkably improved, a two-parameter model can be quickly inverted, and the overall calculation efficiency is improved. The improvement of the efficiency is of great significance to the improvement of the overall effect of ground penetrating radar exploration and the practical application in the fields of urban underground space detection, road maintenance and the like.
Owner:JILIN UNIVERSITY

Lightweight 3D medical image real-time reasoning method and system based on edge computing

The present invention relates to the field of image data processing technology, and specifically discloses a lightweight 3D medical image real-time reasoning method and system based on edge computing, including the construction of a C++ framework model, anisotropy detection and processing, separation of low-resolution axis resampling, 3D resampling, determining the sliding window step size and generating a weight map, using a lightweight student model for prediction, sliding window reasoning, image post-processing, performing size adjustment, voxel connectivity and integrity filling on the predicted segmentation map obtained by reasoning, and outputting the final global segmentation map. The present invention uses a lightweight student model to replace the traditional nnUNet, which requires 5 model reasonings, reducing the number of reasonings by 80% and significantly shortening the single reasoning time; at the same time, C++ language reconstruction is used to achieve CUDA parallelization, and through GPU acceleration of data block extraction, mirror transformation and weighted aggregation operators, the performance of the entire process is improved by more than 10 times compared with Python implementation.
Owner:XIAOZHI FUTURE (CHENGDU) TECH CO LTD

Protocol-aware provisioning of resource over CXL fabrics

Dynamic provisioning of resources in a datacenter improves the utilization efficiency of compute, memory, storage, and network resources, while maintaining flexibility to meet changing demands. Embodiments herein disclose protocol-aware provisioning of resources over Compute Express Link (CXL) fabrics, enabling multi-protocol pooling and disaggregation of resources, including memory and workload-specific accelerators such as GPUs and DSAs. In some embodiments, a Resource Provisioning Unit (RPU) facilitates intent-based protocol translations and mappings between address spaces, potentially enabling the creation of large-scale compute-memory fabrics that may utilize both coherent and non-coherent Non-Transparent Bridging (NTB) between multiple protocols in a single system, serving as the underlying infrastructure for executing workloads such as Large-Language Models (LLMs) and Deep Learning Recommendation Models (DLRMs). Some embodiments also optimize low-latency communication between processes running on different nodes by enabling host-to-host memory provisioning for libraries such as OpenMP, Pthreads, or CUDA, which can utilize shared memory.
Owner:HYATT GAYA OPAL MS +1

Space-based full-link photoelectric imaging simulation signal generation method based on CUDA (Compute Unified Device Architecture)

The invention provides a space-based full-link photoelectronic imaging simulation method based on CUDA (Compute Unified Device Architecture), and solves the problems of complicated calculation, insufficient simulation precision and high hardware requirements in the prior art. The method comprises the following steps: dividing the surface of a target three-dimensional model into triangular surface elements and calculating target intrinsic radiation scattering data of each surface element; generating background intrinsic radiation scattering data in real time through CUDA (Compute Unified Device Architecture) parallel calculation according to the space-time spectrum parameters and the earth surface / cloud layer / limb basic model library; calculating characterization radiation before entrance pupil based on target and background intrinsic data; the entrance pupil radiation is converted into a detector voltage signal through CUDA parallel optimization, and a gray infrared simulation image is generated. According to the method, the calculation efficiency is remarkably improved, and space-based full-link photoelectric imaging simulation with high authenticity and real-time performance is realized.
Owner:XIDIAN UNIV

Homomorphic encryption-based spatiotemporal big data distributed privacy computing system and device

The application provides a space-time big data distributed privacy computing system and equipment based on homomorphic encryption, and relates to the technical field of data processing. It comprises a space privacy computing operator, which is used for the computing characteristics of vector data and raster data, and realizes homomorphic encryption computing of vector data and raster data by using TFHE homomorphic encryption algorithm and CKKS homomorphic encryption algorithm respectively; a distributed space privacy computing framework is used, each node of the distributed space privacy computing framework is respectively provided with a space privacy computing operator, and the distributed space privacy computing framework realizes multi-node task parallelization and safe cooperation based on a key management mechanism under a distributed environment and a remote Boolean decryption protocol based on inadvertent transmission; each node in the distributed space privacy computing framework is provided with a GPU acceleration architecture, the GPU acceleration architecture is respectively provided with a corresponding CUDA heterogeneous computing optimization scheme for the Boolean logic operation of vector calculation and the matrix operation characteristics of raster calculation, and the processing efficiency of a single node is improved.
Owner:PEKING UNIV

High-fidelity lightweight world model construction method for end-to-end autonomous driving test

The application relates to a world model construction method, in particular to an end-to-end automatic driving test high-fidelity lightweight world model construction method, which constructs a high-fidelity world model, solves the problems of large world model parameters and low inference efficiency, performs knowledge distillation on the world model, reduces model parameters on the basis of retaining world model generation capacity, and improves inference efficiency; a CUDA operator is self-defined for a world model calculation bottleneck part, memory allocation is optimized, and a single-device multi-thread scheduling and multi-device cooperative calculation method are used to improve the inference efficiency of the high-fidelity world model. The application can construct an end-to-end automatic driving test high-fidelity lightweight world model, effectively solve the problems of low multi-modal information alignment accuracy, poor cross-view and cross-frame consistency, and low inference efficiency of the existing world model, improve the confidence of an end-to-end automatic driving system test process, and greatly accelerate the test efficiency of the end-to-end automatic driving system.
Owner:JILIN UNIVERSITY

Data processing method and system

Embodiments of the present specification provide a data processing method and system, the method comprising: a reasoning framework, in response to a cold start instruction sent by a cloud platform, cold starting on a target graphics processor, sending an opening hijacking command to a hijacking module, and during the process of loading first model weights of a first reasoning model, initiating a video memory application to the target graphics processor; the hijacking module, in response to the opening hijacking command, hijacking the video memory application, and redirecting the video memory application to a first virtual address in a fixed virtual address space; the reasoning framework, in a case where it is determined that the first model weights are completed loading, sending an ending hijacking command to the hijacking module, and based on the first model weights in the first virtual address, constructing a first reasoning execution graph to execute a first reasoning task using the first reasoning execution graph. By hijacking the CUDA video memory allocation and redirecting it to a pre-reserved fixed address, the address of the model weights is ensured to be persistent and stable, so that the CUDA Graph continues to be effective in multiple loadings and hot switching.
Owner:ALIBABA CLOUD COMPUTING CO LTD

Method and system for quickly encoding image in 3D engine

The invention provides a method and a system for quickly encoding an image in a 3D (three-dimensional) engine. The method for quickly encoding the image in the 3D engine comprises the following steps of: rendering the contents of a plurality of Cameras to the same RenderTexture, and distinguishing the positions of different Cameras; obtaining a bottom layer graphics library pointer of a RenderTexture object through a RenderTexture.GetNativeTexturePtr method provided by Unity, and obtaining a bottom layer graphics library pointer of the RenderTexture object according to the bottom layer graphics library pointer of the RenderTexture object and the bottom layer graphics library pointer of the RenderTexture object and the bottom layer graphics library pointer of the RenderTexture object and the bottom layer graphics library pointer of the RenderTexture object and the bottom layer graphics library pointer of the RenderTexture object. A graphic pointer is obtained through RenderTexture, a resource format is converted and mapped into a CUDA memory, a GPU is used for coding into a preset format, and a video memory is reused to reduce resource consumption. The system comprises modules corresponding to the steps of the method.
Owner:HUIZHIAN INFORMATION TECH CO LTD