Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

606 results about "Graphical processing unit" patented technology

A graphics processing unit ( GPU) is a specialized electronic circuit designed to rapidly manipulate and alter memory to accelerate the creation of images in a frame buffer intended for output to a display device. GPUs are used in embedded systems, mobile phones, personal computers, workstations, and game consoles.

Virtual stylist

An example operation may include at least one of receiving, via a user interface of a device, an activation input from a user to initiate a session, capturing, by a camera of the device, a scan of a body of the user, wherein the capturing comprises recording at least one image and / or at least one video of the user, processing the at least one image and / or video to generate a three- dimensional model of the user comprising measurements and contours of the body, retrieving, from a database, at least one clothing item associated with the user, the at least one clothing item comprising dimensional attributes and texture attributes, rendering, by a graphics processing unit, the at least one clothing item onto the three-dimensional model to generate a visual representation, wherein the rendering simulates draping behavior, movement, and light interaction of the at least one clothing item relative to the three-dimensional model, and displaying, on the user interface, an interactive visualization comprising the visual representation of the three-dimensional model with the at least one clothing item from multiple viewing angles.
Owner:ELGORT PENELOPE

An integrated system for dam failure detection, early warning and evacuation support

An integrated system for detecting dam breaches, as an early warning system and for evacuation assistance, consisting of: A geoinformatics analysis module comprising a modeling processor with a graphics processing unit, configured to: analyze spatial, hydrological, and topographic data of a dam and its downstream area; generate hydrological models using probable maximum rainfall (PMP) data to determine probable maximum flood discharge (PMF); generate hydraulic models to simulate flood propagation and determine flood depth, flood extent, flow velocity, and flood arrival time; create flood vulnerability models, vulnerability models, and flood hotspot models using a multi-criteria decision-making (MCDM) framework; and generate evacuation plans for high-risk zones based on the models. A sensor monitoring subsystem includes: a distributed fiber optic sensor (DFOS) configured to monitor the structural load and deformation of a dam, and a radar level gauge configured to measure the water level in the reservoir; a microcontroller unit configured to receive and process sensor data from the sensor monitoring subsystem, to detect dam breaches based on abnormal structural movements detected by the DFOS or sudden drops in water level below predefined thresholds detected by the radar level gauge; a communication module configured to transmit warning messages when dam breaches are detected; a warning system configured to: send SMS notifications to residents in downstream high-risk zones via a GSM communication module, generate real-time alerts via an IoT platform on user devices, and activate siren systems in villages to issue loud alarms and voice announcements; and an evacuation assistance platform configured to provide real-time evacuation assistance in emergencies due to dam breaks, displaying site maps with nearby emergency shelters, evacuation routes, hospitals, government buildings and aid centers, showing hydrological layers including flood extent, water depth, flow velocity and arrival time of the water, and enabling real-time location tracking in relation to safety zones and flood-prone areas.
Owner:DONGALE TUKARAM DR KOLHAPUR +4

Neural network using dynamically compressed and decompressed weights

A method for training or performing inference using a neural network involves performing per-layer decompression and compression of neural network weights. More particularly, compressed weights are retrieved for a particular layer of the neural network. The weights correspond to neurons in the layer. The compressed weights are decompressed, and input data for that layer is subsequently processed using the decompressed weights. This dynamic decompression and recompression of weights allows memory, and in particular random access memory of graphical processing units, to be efficiently used.
Owner:ROYAL BANK OF CANADA

Methods and systems for dynamically optimizing and modifying allocation of virtual graphical processing units

A method for dynamically optimizing and modifying allocation of at least one disaggregated graphical processing unit (GPU) to at least one process via at least one virtual graphical processing units (vGPU) includes receiving, by a resource manager, from a runtime component executing on a first computing device, a request for allocation, to a process executing on the first computing device, of access to a GPU in the first computing device. The method includes allocating, by the resource manager, to the process, access to at least one vGPU associated with a second GPU on a second computing device. The method includes transmitting, by the resource manager, to the runtime component, an identification of the at least one vGPU identifying the at least one vGPU as located on the first computing device. The method includes instantiating, by the runtime component, the process with access to the at least one vGPU.
Owner:EXOSTELLAR INC

Disaggregated server architecture

Systems, devices, and methods for disaggregating networking components are provided. An example networking chassis includes a first disaggregated server device supported by the networking chassis that includes a first central processing unit (CPU) and a first graphics processing unit (GPU) coupled with the first CPU. The networking chassis further includes a first insertable switch module communicably coupled with the first disaggregated server device that includes first switching chipsets and a first fabric management controller coupled with the first switching chipsets. The first insertable switch module at least partially controls data transmission associated with the first disaggregated server device. The first GPU of the first disaggregated server device is isolated on the first disaggregated server device, supported on the first disaggregated server device in the absence of other GPUs, or is otherwise the only GPU on the first disaggregated server device so as to provide modularity in networking applications.
Owner:NVIDIA CORP

Decentralized uplink detection through expectation propagation

Apparatuses, systems, and techniques to detect uplink data in a multi-user multiple input multiple output wireless communication system. In at least one embodiment, uplink data is detected using one or more graphics processing units to perform parallel computations in order to form a consensus belief about information received by one or more base stations in a wireless network.
Owner:NVIDIA CORP

Graphic processing unit cooperative processing method and system for accelerating three-dimensional generative model

The invention discloses a graphic processing unit cooperative processing method and system for accelerating a three-dimensional generative model. The model comprises a diffusion processing stage and a Gaussian splash rendering stage. In the diffusion processing stage, the time step number and the data precision calculated by the neural network are dynamically adjusted according to the visual angle sensitivity, and a specific graphic processing unit core is scheduled for execution. In the Gaussian splash rendering stage, pixel block tasks are predicted and intelligently scheduled to a stream processor based on historical data; and in the stream processor, the processing sequence of the Gaussian ball by the rendering unit is optimized according to the geometric distance, and the close-range unit is preferentially processed. During back propagation, the built-in logic of the stream processor preaggregates the gradient contributions of the same Gaussian ball, and then updates the gradient contributions through a single write operation. Through collaborative optimization of software and hardware, the efficiency bottleneck of the 3D generation model on a graphic processing unit is effectively solved, the calculation speed, the throughput and the hardware utilization rate are remarkably improved, and the model generation time is shortened.
Owner:SHANGHAI JIAOTONG UNIV

Key value database storage method and system based on CSD

The invention relates to the technical field of data storage, and discloses a CSD-based key value database storage method and system, and the method comprises the steps: responding to a user inserting a key value pair into a key value database, storing the key value pair into a memory table, and converting the memory table into an immutable memory table after a preset triggering condition is satisfied; aligning and coding the key value pairs in the immutable memory table according to a uniform length to generate key value pairs with fixed step lengths; calling a computing storage device to automatically compress the key value pair with the fixed step length and flash the key value pair into an SSTable file; receiving a merging operation triggering instruction, and controlling the computing storage device to obtain the SSTable file and decompress the SSTable file; and calling the graphic processing unit to merge the decompressed files to generate a new SSTable file, and calling the computing storage device to automatically compress the new file and flash the new file to a disk. According to the method, the read-write throughput and merging efficiency of the key value database are remarkably improved, and the problems of write blockage and write pause are effectively relieved.
Owner:JINAN INSPUR DATA TECH CO LTD +1

Orchestration of ai model deployment on multi-GPU systems

Apparatuses, systems, and frameworks for provisioning of efficient pipelines capable of multi-model inference and data processing using multiple processing units, including streaming data applications. The disclosed techniques include, during an initialization stage, assigning a plurality of machine learning models (MLMs) for execution on graphics processing units (GPUs), allocating memory space, on a hub GPU, to the plurality of MLMs, storing input data on the hub GPU before transferring the input data to other GPUs for execution. During an execution stage, output data is initially stored on GPUs that generated the output data before transferring the output data to the hub GPU.
Owner:NVIDIA CORP

Intelligent cooperative high-power proton exchange membrane fuel cell life prediction and fault diagnosis method and system

The invention relates to the technical field of fuel cell monitoring and health management, and discloses an intelligent cooperative high-power proton exchange membrane fuel cell life prediction and fault diagnosis method and system, and the method comprises the following steps: S1, collecting multi-dimensional data of a stack key region in real time through a distributed sensing array; s2, removing noise and abnormal values of the collected data, and carrying out feature extraction; s3, constructing a dual-model collaborative prediction framework; s4, inputting the real-time features into the double models; s5, generating a maintenance strategy and early warning; according to the invention, special hardware accelerators such as a graphic processing unit or a field-programmable gate array are integrated on the edge computing module, so that the processing speed and efficiency of high-frequency and multi-source sensor data at a local end are remarkably improved; a distributed computing framework is adopted in a cloud analysis module to process mass feature data transmitted by an edge end, so that the capability and scale of the system for performing deep analysis on the health state of the fuel cell, accurately predicting the remaining service life and quickly tracing the fault are enhanced.
Owner:INNER MONGOLIA YIPAI HYDROGEN ENERGY TECH CO LTD

Managing GPU resources on a container orchestration platform

A technique manages computing resources on a container orchestration platform. Such a technique involves establishing a pool of computing resources on the container orchestration platform. Such a technique further involves, after the pool of computing resources is established, receiving graphics processing unit (GPU) provisioning requests (GPRs) which identify workspaces. Such a technique further involves allocating computing resources from the pool to the workspaces identified by the GPRs based on a set of GPR prioritization policies.
Owner:AVESHA INC

Remote desktop resolution optimization method and system in wayland environment

The invention provides a remote desktop resolution optimization method and system in a wayland environment, and the method comprises the steps: S1, obtaining a wayland protocol of screen display data, and adding a mode event in the wayland protocol, so as to achieve the communication between a remote desktop client and a local synthesizer; s2, the synthesizer calculates to obtain the optimal resolution, and sends the optimal resolution to a server side of the remote desktop through a mode event; and S3, after receiving the optimal resolution, the server issues the optimal resolution to the graphic processing unit and updates the modified display effect in the next frame of picture, and the step S2 is executed repeatedly until a stable state is reached. The method has the advantages of high flexibility, high efficiency and the like, and can allow the user to adjust the resolution of the remote desktop in real time under the condition of keeping the remote session uninterrupted.
Owner:KYLIN CORP

Fault detection method and device, electronic equipment and storage medium

The invention discloses a fault detection method and device, electronic equipment and a storage medium, and relates to the technical field of computers.The method comprises the steps that enhanced error report register data and machine inspection architecture register data from a server are received, the enhanced error report register data and the machine inspection architecture register data are sent when the server detects that the graphics processing unit has uncorrectable errors and the number of times of correctable errors reaches a preset threshold value; under the condition that the enhanced error report register data and the machine check architecture register data are valid, determining whether the graphics processing unit is in a healthy state according to the machine check architecture register data; if not, the fault type and the fault position of the graphic processing unit are determined according to at least one of the processing core information and the memory information included in the enhanced error report register data, so that specific internal equipment with faults of the graphic processing unit can be accurately positioned.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Processing unit fault detection

Various example embodiments of a processing unit fault detection capability are presented. The processing unit fault detection capability may be configured to support detection of faults in processor cores of a processing unit based on arrangement of the processor cores to form a data processing pipeline and monitoring of the processor cores of the data processing pipeline based on monitoring of the data processing pipeline (e.g., based on propagation of heartbeat messages via the data processing pipeline). The processing unit fault detection capability may be configured to support detection of faults in processor cores of various types of processing units, such as central processing units (CPUs), graphics processing units (GPUs), network processing units (NPUs), or the like.
Owner:NOKIA SOLUTIONS & NETWORKS OY

Hardware apparatus for isolated virtual environments

A hardware apparatus for isolated virtual environments includes graphics processing unit comprising a first dedicated memory and a first plurality of processing cores, a central processing unit comprising a second dedicated memory and a second plurality of processing cores, a field programmable gate array comprising a third dedicated memory, a control and data bus assembly connecting the field programmable gate array, the central processing unit, and the graphics processing unit, and a hypervisor located on a non-volatile memory of the hardware apparatus, the hypervisor configured to create one or more virtual machines by isolating the graphics processing unit, the central processing unit and the field programmable gate array.
Owner:PARRY LABS LLC

Systems and methods for heterogeneous large language model encoder and decoder processing

Systems and methods are disclosed for efficient memory allocation for processing large language model encoders and decoders based on an attention model. The system can utilize a plurality of two types of processors suitable for different types of LLM processing. These include Neural Processor Units and Graphic Processing Units. Each NPU processor has dedicated DDR memory coupled to each NPU. The DDR memory caches the neural network weights used in the generation of neural network activations. A plurality of GPUs provides KVQ token processing. LLM tokens processing can be performed in parallel batches or sub-batches to utilize idle NPU processors within the neural network. In some embodiments, the NPUs are structured in a matrix with a bus between adjacent processors. In another embodiment, NPUs provide both KVQ processing and neural network processing. The system can be integrated on a silicon substrate using chiplets in a 2.5 or 3-D architecture.
Owner:EXPEDERA INC

Multi-speech synthesis model bearing method and device based on virtual GPU

The invention provides a multi-speech synthesis model bearing method and device based on a virtual GPU, and relates to the technical field of graphics processing units, and the method comprises the steps: carrying out the virtualization processing of a physical graphics processing unit, dividing the physical graphics processing unit into a plurality of virtual processing units with independent video memories and calculation quotas, and combining a resource scheduling mechanism, and deploying the speech synthesis language model instances in a plurality of service containers, and constructing a plurality of speech synthesis model bearing units. After a voice synthesis request is accessed, the scheduling module carries out load balancing according to the request connection number of each bearing unit, the request is distributed to a target bearing unit with the minimum connection number, and a voice generation task is completed by a virtual processing unit bound with the target bearing unit. According to the invention, resource division can be carried out on the physical graphic processing unit, and efficient operation of the multi-speech synthesis model is realized.
Owner:ZHEJIANG RONGQI MANUFACTURING TECHNOLOGY CO LTD

Techniques for implementing multiple large language models on a single physical computing device

Techniques for executing multiple large language models (LLMs) on a single physical computing device, including receiving a task to be performed by at least a first LLM and a second LLM, creating a plurality of virtual compute nodes on the single physical computing device for processing at least some of first set of sub-tasks to be performed by the first LLM and second set of sub-tasks to be performed by the second LLM, wherein each one of the plurality of virtual compute nodes includes one or more graphics processing unit (GPU) cores allocated for performing a sub-task of the first set of sub-tasks or the second set of sub-tasks, and executing at least some of the first set of sub-tasks and the second set of sub-tasks substantially in parallel across the plurality of virtual compute nodes using shared memory. Inferences between the first and second LLMs may be shared.
Owner:DESAI AMBARISH AJIT +2

Performance test method and device for virtual graphic processing unit

The invention discloses a performance testing method and device for a virtual graphics processing unit. The method comprises the following steps: acquiring a business load type and a test target of a to-be-tested virtual graphic processing unit; the service load type and the test target are analyzed by using a configuration decision model, a target test configuration strategy of the virtual graphic processing unit is determined, the configuration decision model is obtained by training based on a dual-depth Q network algorithm, and the target test configuration strategy is determined; the target test configuration strategy is used for reflecting respective execution sequences and resource configuration information of multiple performance tests to be executed by the virtual graphics processing unit; and performing multiple performance tests on the virtual graphic processing unit based on the target test configuration strategy to obtain a plurality of corresponding performance test results. According to the method and the device, the technical problems of relatively low efficiency and subjective deviation of a configuration result caused by manually configuring various performance tests of the virtual graphic processing unit in related technologies are solved.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Intelligent chart visualization method based on model control protocol

The invention discloses an intelligent chart visualization method based on a model control protocol. The method comprises the following steps: step 1, establishing connection between a visual rendering client and an intelligent analysis server; 2, receiving a natural language visualization instruction of a user, and obtaining source data; 3, intention recognition is conducted on the natural language visualization instruction, and feature extraction is conducted on source data; 4, determining an adaptive visual chart type; 5, generating a model control protocol instruction packet; 6, receiving and analyzing a model control protocol instruction packet; and step 7, constructing a three-dimensional grid model, and submitting the three-dimensional grid model to a graphic processing unit for rendering display. According to the method, an artificial intelligence natural language understanding technology and a three-dimensional graph programming visualization technology are combined, and the method has the advantages of chart self-adaption, visualization and scene fusion, real-time and efficient rendering and the like.
Owner:ZHEJIANG UNIV

Real-time object detection from decompressed images

A system and method for deploying machine learning models with consistent fixed-point arithmetic processing. A computing device in a vehicle receives a trained machine learning model that was trained using training images processed with fixed-point arithmetic operations. The computing device receives images from a camera and processes the received images using fixed-point arithmetic operations that are consistent with those used during training. The machine learning model is executed using the processed images to detect objects such as vehicles, pedestrians, persons on bikes, stop lights, or stop signs. The computing device may comprise specialized hardware including a graphics processing unit (GPU), digital signal processing hardware decoder, or single-purpose hardware decoder such as an application-specific integrated circuit (ASIC) or field-programmable gate array (FPGA). The system may generate alerts or determine vehicle operation changes based on detected objects, and may operate in surveillance mode when the vehicle is parked.
Owner:NETRADYNE INC

Dynamic sharing of graphics processing unit (GPU) computational capabilities based on processing density

Approaches presented herein provide systems and methods for dynamic allocation of processing units to increase computational density. Idle times between sequential processing tasks may be computed and, if the idle time exceeds a threshold capacity, additional sequential processing tasks may be allocated to a common processing unit. As a request, a first portion of a first sequential processing task may be executed, then a second portion of a second sequential processing task may be executed prior to executing a subsequent portion of the second sequential processing task. By using the idle time between portions of sequential processing tasks, output perform may be maintained while using additional processing capabilities that would otherwise remain idle.
Owner:NVIDIA CORP

Providing native support for generic pointers in a graphics processing unit

Systems and methods for supporting generic pointers in hardware of a graphics processing unit (GPU) are provided. In various examples, a GPU includes multiple sub-cores each having a processing resource and a load / store pipeline. The processing resource is operable to receive a memory access message including a generic pointer. The processing resource is further operable to output a load or store operation to the load / store pipeline based on the memory access message, in which an address for the load or store operation is determined based on a base address of a named memory type of a plurality of named memory types referenced by the generic pointer and an offset into a memory of the named memory type. The load / store pipeline is operable to, responsive to receipt of the load or store operation, access the memory at the address.
Owner:INTEL PRODUCTS IP LLC

Methods and apparatuses for executing GPU task in confidential compute architecture

A graphics processing unit (GPU) task is executed in a confidential compute architecture. GPU software in a non-secure world configures, based on task code and a cache description of a GPU task, a stub data structure including cache areas allocated based on the cache description and metadata indicating each cache area. In a realm segment in a memory, a root world root monitor configures a real data structure corresponding to the stub data structure, and stores to-be-processed confidential data. The root monitor updates a granule protection table (GPT) table so that based on an updated GPT table, a target segment storing the metadata and the task code is accessible to a GPU and has realm world permission for all other objects. The root monitor modifies a target mapping relationship so that the GPU executes the GPU task by using the target segment and the real data structure.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD +1

Quantum-based dispatch of workgroups

Quantum-based dispatch of workgroups is described. An example of an apparatus includes a computer memory to store data for processing, including data for an application; and one or more processors including a graphical processing unit (GPU), the GPU including multiple chiplets, each of the multiple chiplets including compute containers and a cache, each compute container including a plurality of processing resources, and a dispatcher for dispatching workgroups to the processing resources of the GPU, wherein dispatching workgroups includes dispatching workgroups for the application according to a selected workgroup quantum, the selected workgroup quantum having a certain size and shape.
Owner:INTEL CORP

System and method for executing fused neural-network layer architectures

The present disclosure provides a system and method for executing fused neural network layers using a graphics processing unit (GPU). The fused neural network layer combines multiple neural network operations into a single GPU kernel function, for efficient utilization of a GPU shared memory to reduce global memory transactions. The fused neural-network-layer system configures GPU thread blocks to iterate tiles in the GPU shared memory across portions of input tensors, and to perform a sequence of neural network layer operations using the tiles, before storing the results in an output tensor. The neural network layer operations can include element-wise, normalization, and pooling operations. The system also supports fused layers with nested traversals of input tensors, such as matrix multiplication, convolution, and attention mechanisms. The fused neural-network-layer system improves performance and reduces memory overhead compared to executing each layer as a separate GPU kernel, thereby enabling faster training and inference times.
Owner:BAYRIS INC

Information retrieval method, device and equipment and computer readable storage medium

The invention discloses an information retrieval method, device and equipment and a computer readable storage medium, and relates to the technical field of natural language process.The method includes the steps that after to-be-retrieved text data is received, vectorization processing is conducted on the to-be-retrieved text data, the storage space and the calculation overhead are greatly reduced, a hierarchical index structure is constructed in advance, and the retrieval efficiency is improved. An adjacency list corresponding to each vector node in a hierarchical index structure is set to contain a fixed preset number of similar vectors, a parallel computing architecture of the graphic processing unit is adapted, parallel information retrieval of the graphic processing unit is facilitated, rough information retrieval from a coarse-grained index end is transited to fine information retrieval layer by layer, and the information retrieval efficiency of the graphic processing unit is improved. The information retrieval process is accelerated, the technical problem that the parallel computing capacity of the image processing unit cannot be fully utilized is solved, and the technical effects that the parallel processing capacity of the image processing unit is fully utilized, the performance is greatly improved, and the information retrieval efficiency is improved are achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Systems and methods for enhancing execution of interpreted computer languages

Systems and methods are provided that incorporate a compiler configured to convert interpreted language code (e.g., Python) into native machine code. According to some embodiments, the system generates the native machine code into a format that is consistent with known infrastructure. The native machine code can be converted into a format based on a low level virtual machine “LLVM” infrastructure. In various embodiments, the system enables a compiler framework that improves execution of code for interpreted languages. According to one embodiment, the system can be tailored for execution on specific processors, for example, a graphics processing unit (“GPU”) that is optimized for highly parallel computations.
Owner:EXALOOP INC

Performance test method, device and equipment of graphic processing unit and storage medium

The invention discloses a performance testing method, device and equipment of a graphic processing unit and a storage medium, and relates to the technical field of servers. The configuration file comprises category weights of a plurality of test categories corresponding to the preset scene, test item weights of a plurality of test items and a reference running duration corresponding to each test item; running the test case corresponding to each test item on the GPU to be tested, and determining the actual running time length corresponding to each test item; according to the test item weight of each test item, the category weight of the test category to which the test item belongs, the reference operation duration and the actual operation duration, determining a comprehensive performance index of the to-be-tested graphic processing unit; and determining whether the performance of the to-be-tested graphic processing unit meets the performance requirement of the preset scene or not according to the comprehensive performance index. According to the method, the weight of each test index for evaluating the performance of the to-be-tested GPU can be adjusted according to the preset scene, and the performance of the to-be-tested GPU can be accurately evaluated.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Cooperative parallel memory allocation

Apparatuses, systems, and techniques to perform multi-threaded memory allocation in parallel by one or more software programs being performed on a parallel processing unit (PPU), such as a graphics processing unit (GPU), or any other processing unit capable of supporting multi-threaded software execution. In at least one embodiment, one or more software programs expressed in part by code using an application programming interface for parallel computing, such as CUDA, perform allocation, search, and deallocation of memory efficiently and in parallel on a GPU.
Owner:NVIDIA CORP