Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

254 results about "Graphical processing unit" patented technology

A graphics processing unit ( GPU) is a specialized electronic circuit designed to rapidly manipulate and alter memory to accelerate the creation of images in a frame buffer intended for output to a display device. GPUs are used in embedded systems, mobile phones, personal computers, workstations, and game consoles.

Managing GPU resources on a container orchestration platform

A technique manages computing resources on a container orchestration platform. Such a technique involves establishing a pool of computing resources on the container orchestration platform. Such a technique further involves, after the pool of computing resources is established, receiving graphics processing unit (GPU) provisioning requests (GPRs) which identify workspaces. Such a technique further involves allocating computing resources from the pool to the workspaces identified by the GPRs based on a set of GPR prioritization policies.
Owner:AVESHA INC

Intelligent chart visualization method based on model control protocol

The invention discloses an intelligent chart visualization method based on a model control protocol. The method comprises the following steps: step 1, establishing connection between a visual rendering client and an intelligent analysis server; 2, receiving a natural language visualization instruction of a user, and obtaining source data; 3, intention recognition is conducted on the natural language visualization instruction, and feature extraction is conducted on source data; 4, determining an adaptive visual chart type; 5, generating a model control protocol instruction packet; 6, receiving and analyzing a model control protocol instruction packet; and step 7, constructing a three-dimensional grid model, and submitting the three-dimensional grid model to a graphic processing unit for rendering display. According to the method, an artificial intelligence natural language understanding technology and a three-dimensional graph programming visualization technology are combined, and the method has the advantages of chart self-adaption, visualization and scene fusion, real-time and efficient rendering and the like.
Owner:ZHEJIANG UNIV

System and method for executing fused neural-network layer architectures

The present disclosure provides a system and method for executing fused neural network layers using a graphics processing unit (GPU). The fused neural network layer combines multiple neural network operations into a single GPU kernel function, for efficient utilization of a GPU shared memory to reduce global memory transactions. The fused neural-network-layer system configures GPU thread blocks to iterate tiles in the GPU shared memory across portions of input tensors, and to perform a sequence of neural network layer operations using the tiles, before storing the results in an output tensor. The neural network layer operations can include element-wise, normalization, and pooling operations. The system also supports fused layers with nested traversals of input tensors, such as matrix multiplication, convolution, and attention mechanisms. The fused neural-network-layer system improves performance and reduces memory overhead compared to executing each layer as a separate GPU kernel, thereby enabling faster training and inference times.
Owner:BAYRIS INC

Cooperative parallel memory allocation

Apparatuses, systems, and techniques to perform multi-threaded memory allocation in parallel by one or more software programs being performed on a parallel processing unit (PPU), such as a graphics processing unit (GPU), or any other processing unit capable of supporting multi-threaded software execution. In at least one embodiment, one or more software programs expressed in part by code using an application programming interface for parallel computing, such as CUDA, perform allocation, search, and deallocation of memory efficiently and in parallel on a GPU.
Owner:NVIDIA CORP

Artificial intelligence-based business management system for optimizing real-time decisions

An artificial intelligence-based business management system for real-time decision optimization, consisting of: a data acquisition unit configured to ingest structured, semi-structured and unstructured data streams from enterprise resource planning (ERP) systems, customer relationship management (CRM) platforms, Internet of Things (IoT) devices, external market feeds and financial transaction systems, encrypting, timestamping and verifying the data prior to further processing; a graph processing unit that is communicatively connected to the said acquisition layer, wherein the unit is configured to encode heterogeneous data into dynamic graph structures comprising nodes representing business entities and edges representing transaction or relationship dependencies, wherein the unit is further configured to perform deduplication, metadata tagging and real-time data synchronization; a decision optimization unit operationally linked to the graph processing unit, wherein the decision optimization unit includes modules for reinforcement learning, modules for Bayesian optimization and multi-objective solvers configured to simulate multiple alternative decision paths and select an optimal path based on performance indicators such as cost efficiency, resource utilization, customer satisfaction and risk minimization; a real-time inference control unit with specialized hardware cores, including at least one graphics processing unit (GPU), a field-programmable gate array (FPGA) and an application-specific integrated circuit (ASIC), wherein the accelerator performs inference tasks of the decision optimization unit with a latency in the millisecond range; a diagnostic processing unit configured to generate causal diagrams, feature mapping maps, and interpretable result summaries according to the optimization outputs; and a control interface unit configured to transmit optimized decisions to process controls within the enterprise, robot actuators, planning systems or interactive dashboards, with the interface supporting bidirectional communication for higher-level interventions, error feedback and triggers for re-optimization.
Owner:ABUELENAIN EMAD EDDIN AHMED +4

Image processing method and device and medium

The embodiment of the invention relates to an image processing method and device and a medium. The method proposed herein includes: generating, by a graphics processing unit, a first reference frame, a second reference frame, and motion vector information corresponding to the first reference frame; transmitting the motion vector information to a digital signal processor by a graphic processing unit; the digital signal processor divides the motion vector information into a plurality of groups; processing the plurality of groups in parallel by using a plurality of processing cores of the digital signal processor, each processing core being configured to: write a group of motion vectors indicated by a corresponding group to a group of projection positions of the group of motion vectors in the interpolated frame by using a scatter instruction; respectively reading source pixel information from the first reference frame and the second reference frame based on the written group of motion vectors by utilizing a collection instruction; the source pixel information is fused by a digital signal processor to generate an interpolated frame. In this way, according to the embodiment of the invention, the generation efficiency of the interpolation frame can be effectively improved.
Owner:VASTAI TECH (SHANGHAI) INC

System and method for fine-tuning rotated outlier-free large language models for effective weight-activation quantization

A computing device includes at least one processor, one or more non-transitory computer-readable storage media, a system for fine-tuning a large language model under low-bit weight-activation quantization. The computing device further comprises a graphics processing unit (GPU), a neural processing unit (NPU), or a tensor processing unit (TPU). The hardware interface module of the system is configured to load a low-bit model representation from the memory module and transmit the model representation to the GPU, NPU, or TPU for inference execution.
Owner:THE HONG KONG UNIV OF SCI & TECH

COMMUNICATION OPTIMIZATION FOR MoE BY OFFLOADING EXPERTS TO NICs

Embodiments herein describe a system including a plurality of hardware accelerators including at least one mixture-of-experts (MoE) layer having multiple experts and a plurality of network interface cards (NICs) coupled to the plurality of hardware accelerators, wherein at least one expert of the multiple experts is offloaded from the plurality of hardware accelerators to the plurality of NICs. The plurality of hardware accelerators may be graphics processing units (GPUs). In one example, a subset of the multiple experts are selectively offloaded from the plurality of GPUs to the plurality of NICs based on memory and computational capacity available on the plurality of NICs. In another example, the multiple experts are designated as either hot experts or cold experts. The cold experts are offloaded from the plurality of GPUs to the plurality of NICs and the hot experts are duplicated for each of the plurality of GPUs.
Owner:ADVANCED MICRO DEVICES INC +1

Large model deployment and service method and device for realizing intelligent customer service based on container arrangement, equipment and medium

The invention discloses a large model deployment and service method and device for realizing intelligent customer service based on container arrangement, equipment and a medium, and relates to the field of artificial intelligence, and the method comprises the steps: constructing a standardized container mirror image which comprises an inference engine, a model acceleration library and a model encryption and decryption assembly and is adaptive to an intelligent customer service scene; establishing a graphic processing unit resource management system in the container arrangement cluster; storing model data of the intelligent customer service model to a target storage system, and determining a mounting mechanism corresponding to the model data; generating a container arrangement and deployment file based on a graphic processing unit resource management system, the operation demand of the standardized container mirror image and a mounting mechanism corresponding to the model data; and submitting the container arrangement deployment file to a container arrangement platform, so that the container arrangement platform carries out intelligent customer service model deployment, and the deployed intelligent customer service model is utilized to receive a user request and return a corresponding target response. According to the method, the deployment complexity is reduced, and the model deployment efficiency and stability are improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Three-dimensional visual synthesis system based on heterogeneous collaboration and screen alignment fusion

The invention relates to a three-dimensional visual synthesis system based on heterogeneous collaboration and screen alignment fusion, which is characterized in that a master control system generates complete viewpoint configuration information of a current scene acquired by an image acquisition device, and the complete viewpoint configuration information is used as a single data source and is simultaneously distributed to a graphic processing unit and a neural rendering coprocessor of the master control system; synchronous starting and parallel execution of the two heterogeneous rendering pipelines are achieved. And in combination with a subsequent screen alignment fusion mechanism, an end-to-end collaboration framework with low delay, low coupling and high coordination is constructed, and the key problem of real-time integration of high-fidelity contents is effectively solved. On the premise of not depending on a traditional deep buffer area or geometric alignment, high-fidelity visual content generated by special hardware acceleration and a dynamic environment of a main control system are efficiently fused, and complex effects such as lightweight asynchronous communication and real-time shadow are supported. Therefore, multiple bottlenecks in the aspects of real-time performance, compatibility and engineering deployment in the prior art are broken through.
Owner:SHANGHAI TECH UNIV

Methods and apparatus to perform instruction-level graphics processing unit (GPU) profiling based on binary instrumentation

Disclosed examples include generating instrumented code by inserting profiling instructions at insertion points in code; outputting the instrumented code for execution by second programmable circuitry; and accessing profiling data generated by the second programmable circuitry based on the instrumented code.
Owner:INTEL CORP

Techniques to support transformer models in analog compute-in-memory hardware

The present disclosure provides a method for implementing transformer models in analog compute-in-memory hardware. The method comprises training a target neural network using one or more operators on one or more graphics processing units, generating one or more datasets from full network traces to capture input-output relationships of non-vector-matrix multiplication operations, training one or more multi-layer perceptrons to approximate the non-vector-matrix multiplication operations using the one or more datasets, replacing the original non-vector-matrix multiplication operations with the trained one or more multi-layer perceptrons, and mapping the resulting multi-layer perceptron-only neural network to an analog compute-in-memory architecture. The non-vector-matrix multiplication operations comprise layer normalization operations, softmax operations, and GELU activation operations. The analog compute-in-memory architecture comprises crossbar arrays of memory elements that store weight values as analog quantities using conductance or capacitance properties.
Owner:GEORGIA TECH RES CORP

Systems and methods for high-bandwidth minimally invasive brain-computer interfaces

Systems and methods for high-bandwidth, minimally invasive brain-computer interfaces (BCIs) are disclosed. The BCIs are configured for deployment and operation in conjunction with a comprehensive interventional electrophysiology procedural suite. Three primary methods of minimally invasive electrode array delivery are disclosed: (1) cortical surface delivery, (2) ventricular delivery, and (3) endovascular delivery. Additionally, systems and methods for interacting with such high-bandwidth electrode arrays are discussed, including real-time imaging, signal processing, and neural decoding. Systems and methods for architectures for accelerating the underlying computational processes (such as graphics processing units or tensor processing units) are also discussed. Multiple applications of BCIs are discussed, with emphasis on restoration, rehabilitation, and augmentation of neurologic function.
Owner:PRECISION NEUROSCIENCE CORP

Parallel slice encoding across gpus with predicated multi-reference image

A processing system employs at least two graphics processing units (GPUs) to encode video. The GPUs employ sets of predicated values that indicate when reconstructed slices have been transferred between the GPUs. Furthermore, each GPU maintains a set of previous reference images, and encodes video slices based on the previous reference images having an expected predicated value. This allows each GPU to identify which reference images to use for encoding. This in turn allows the processing system to encode video frames without synchronization of the GPUs, while maintaining the quality of the encoded video.
Owner:ATI TECHNOLOGIES ULC

Computing system with graphics processing unit (GPU) overlay with quantum processing unit (QPU)

The present patent document provides a design of an efficient hybrid quantum classical computing system capable of information processing based on both quantum computing using different quantum states of qubits and classical digital computing using a digital processor including one or more graphics processing unit (GPU) processors.
Owner:SEEQC INC

server

This application provides a server applicable to the field of server technology. The server includes: a first graphics processing unit configured to write first data to be transmitted to a first memory expansion device according to a first address; the first memory expansion device configured to, upon detecting that the first data to be transmitted has been written to the first memory expansion device, write the first address to a central processing unit; and the central processing unit configured to, upon detecting that the first address has been written to the central processing unit, query a second address based on the first address, and transmit the first data to be transmitted to another server according to the second address via a fast computation link switch, so that the other server, upon detecting that the first data to be transmitted has been written to a second memory expansion device in the other server, invokes a second graphics processing unit in the other server to read the first data to be transmitted from the second memory expansion device.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Multi-tile graphics processing unit

An apparatus to facilitate processing in a multi-tile device is disclosed. The apparatus comprises a plurality of processing tiles, each including a memory device and a plurality of processing resources, coupled to the device memory, and a memory management unit to manage the memory devices in each of the plurality of tiles to perform allocation of memory resources among the memory devices for execution by the plurality of processing resources.
Owner:INTEL CORP

Display driving apparatus, and display apparatus having the same

PendingUS20260046472A1Television system detailsCathode-ray tube indicatorsGraphicsMain processing unit
A display apparatus includes: a video processing unit processing a video signal; a graphics processing unit processing a graphics signal; a mixing unit mixing video corresponding to the processed video signal and graphics corresponding to the processed graphics signal; a display unit outputting the mixed video and graphics; and a main processing unit configured to control the video processing unit, the graphics processing unit, the mixing unit and the display unit.
Owner:SAMSUNG ELECTRONICS CO LTD

Task scheduling method, electronic device, storage medium, and chip

This application provides a task scheduling method, an electronic device, a storage medium, and a chip, and relates to the field of terminal technologies. The method includes: acquiring a plurality of graphics processing unit (GPU) tasks, and sorting the plurality of GPU tasks in descending order of priorities of GPU task types; allocating resources to the plurality of GPU tasks in a plurality of ring buffers, so that resource proportions of GPU tasks sorted higher are higher at positions closer to heads of the plurality of ring buffers; and sequentially executing the plurality of GPU tasks. By using the method, in a scenario with heavy GPU load, freezing of a currently displayed picture of an electronic device can be reduced, thereby improving the user experience. This solution requires no GPU frequency increase, so that no GPU power consumption increase is caused.
Owner:HONOR DEVICE CO LTD

Dual-talking detection and acoustic echo cancellation co-processing method based on dual-cuda streams

The present application belongs to the technical field of speech processing, and particularly relates to a double-talking detection and acoustic echo cancellation cooperative processing method based on double CUDA streams. The method comprises the following steps: creating a first CUDA stream and a second CUDA stream on the same graphics processing unit; establishing a shared memory control package for cross-stream control between the first CUDA stream and the second CUDA stream; organizing a microphone near-end signal and a far-end reference signal into continuous frames with a preset frame length and frame shift; in the second CUDA stream, recording a synchronization event after writing sub-band control parameters and frame-level mode markers into the shared memory control package; in the first CUDA stream, performing partition blocking frequency domain sub-band adaptive filtering to update filter coefficients and generate echo suppression output. The present application realizes a highly cooperative, low-delay, high-robust real-time processing system for double-talking detection and echo cancellation, and has operation efficiency, control accuracy and engineering realizability, and is suitable for voice communication, remote conference and intelligent voice interaction systems.
Owner:CHINA NUCLEAR IND MAINTENANCE

Graphics processing unit (GPU) cluster system and server

PCT designated stageWO2026138119A1GraphicsComputer architecture
Embodiments of the present disclosure provide a graphics processing unit (GPU) cluster system and a server. The system comprises: a computing node and a switching node; a first board where the computing node is located and a second board where the switching node is located are orthogonally arranged in a cabinet, and each GPU of the computing node is interconnected with each switching chip of the switching node respectively by means of an orthogonal OD connector.
Owner:ZTE CORP

Static data center power balancing and configuration

A system to select graphics processing units (GPUs) to execute a task is disclosed. In at least one embodiment, GPUs are selected based on one or more task parameters and one or more fused parameters.
Owner:NVIDIA CORP

Flow cytometry waveform processing

PendingUS20260202304A1GraphicsEngineering
Systems and methods for analyzing flow cytometry particles. A flow cytometry system is configured to direct a fluid stream of particles through an interrogation location, and includes a laser configured to emit light toward the interrogation location to produce light signals from the particles, and one or more detectors configured convert the light signals to waveform data. The flow cytometry system includes a waveform acquisition device to continuously digitize the waveform data, and a graphics processing unit configured to apply one or more adjustable threshold voltages to the digitized waveform data to extract event data for the particles.
Owner:BECKMAN COULTER INC

Realtime structured illumination super-resolution high-throughput imaging method

Systems and methods for image reconstruction are provided. The method may include: implementing a first thread to acquire image data of an object by imaging the object in at least one structured illumination direction using an imaging device, and storing the image data into a buffer; implementing a second thread to obtain one or more image sequences from the buffer; and / or implementing a third thread to extract at least one set of raw images from the one or more image sequences, and transmit the at least one set of raw images to a graphics processing unit (GPU) of the one or more processors, to facilitate the GPU to generate a target image based on the at least one set of raw images, each set of the at least one set of raw images corresponding to one of the at least one structured illumination direction.
Owner:GUANGZHOU COMPUTATIONAL SUPER RESOLUTION BIOTECH CO LTD

Computing node with a chassis featuring a front-mounted GPU compartment

A compute node comprising: a chassis (102; 202) with a base (108; 208), a pair of walls (110; 210) each connected to a perimeter side (108A, 108B; 208A, 208B) of the base (108; 208), and a first upper cover section (112A; 212A) connected to the pair of walls (110; 210) to cover part of the chassis (102; 202); a locking mechanism (104, 104A1, 204, 204A1) coupled to a rear inner end face (138; 238) of the first upper cover section (112A; 212A); and a shelf (106; 206) which is slidable from a front of the chassis (102; 202) and attached to the locking mechanism (104, 104A1, 204, 204A1), and wherein each shelf (106; 206) comprises: a front cover (142A); a base (144A) coupled to the front cover (142A); a pair of brackets (146A) connected to the floor (144A); a pair of risers (148A) comprising a first riser (148A1) containing a first plug socket (178A1) and a second riser (148A2) containing a second plug socket (178A2), each riser (148A1, 148A2) of the pair of risers (148A) being connected to a corresponding bracket of the pair of brackets (146A); and A pair of graphics processing unit (GPU) card assemblies (150; 250) comprising a first GPU card assembly (150A1) with a first connector and a second GPU card assembly (150A2) with a second connector, wherein the first GPU card assembly (150A1) and the first riser (148A1) are arranged along one lateral side of the pair of brackets (146A) and the second GPU card assembly (150A2) and the second riser (148A2) are arranged along another lateral side of the pair of brackets (146A), wherein the first connector is opposite the second connector socket (178A2) and the second connector is opposite the first connector socket (178A1), each GPU card assembly (150A1, 150A2) of the pair of GPU card assemblies (150; 250) being connected to a respective riser. (148A1) of the pair of risers (148A) is connected.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

Method and device for acquiring screen data, storage medium and program product

The invention relates to a method and device for obtaining screen data, a storage medium and a program product. Obtaining drawing image data after the terminal performs application drawing on the to-be-displayed interface, wherein the drawing image data comprises image data corresponding to at least one image layer of the to-be-displayed interface; under the condition that it is determined that the at least one image layer meets the preset synthesis condition, according to the drawn image data, synthesizing the physical screen on-screen content of the to-be-displayed interface through the graphic processing unit to obtain synthesis data; the preset synthesis condition represents that the synthesis demand of the at least one layer exceeds the processing capability of a terminal display processing unit; and acquiring screen recording data of the to-be-displayed interface according to preset screen recording parameters and the synthetic data. Thus, in the process of obtaining the screen recording data, if it is determined that the layers participating in synthesis meet the preset synthesis conditions, the GPU is directly used for one-time synthesis, two-time GPU synthesis is not needed, synthesis time consumption is reduced, and then the situation of frame dropping during screen recording is reduced.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD

Method and apparatus for detecting picked object, computer device, readable storage medium, and computer program product

This application relates to a method for detecting a picked object performed by a computer device. The method includes: detecting trigger location information of a trigger operation; in a process of rendering a second image frame by using a graphics processing unit, obtaining location information of a pixel associated with each three-dimensional model in the second image frame, and matching the location information of the pixel associated with each three-dimensional model with the trigger location information; storing, by using the graphics processing unit, a model identifier of a three-dimensional model associated with a pixel with location information successfully matched, to a storage location in a color buffer that corresponds to the successfully matched pixel; and obtaining the model identifier in the color buffer by using a central processing unit, and determining, based on the obtained model identifier, the picked object specified by the trigger operation.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Image data processing method, system and device, storage medium and program product

The invention provides an image data processing method, system and device, a storage medium and a program product. The image data processing method is applied to a graphic processing unit, and comprises the following steps: reading image texture data in a graphic memory, and rendering to obtain rendered data; wherein the GPU and the image coding unit share at least part of cache space in the graphic memory; writing the rendering data into the cache space of the graphic memory; wherein the rendering data written into the cache space is used for the image coding unit to carry out coding processing.
Owner:MOORE THREADS TECHNOLOGY (SHANGHAI) CO LTD

On-demand GPU enablement

An information handling system may include at least one processor, a memory, and a plurality of graphics processing units (GPUS). The information handling system may be configured to: monitor usage of the plurality of GPUS; in response to the usage of the plurality of GPUs being below a first threshold, execute a command to power off a first number of the plurality of GPUs; and in response to the usage of the plurality of GPUs being above a second threshold, execute a command to power on a second number of the plurality of GPUs.
Owner:DELL PROD LP

Cross GPU barriers using memory semantics protocol

PendingUS20260195292A1GraphicsComputer architecture
A system and method for synchronizing multiple graphics processing units (GPUs) connected via an interconnect link. The system implements a synchronization protocol on top of memory semantics provided by the interconnect link. Each GPU maps memory apertures for other GPUs. The protocol executes a broadcast into all aperture ranges from one GPU with appropriate memory semantics and synchronizes other GPUs by waiting for the memory write to become visible. A command processor performs synchronization logic for each GPU. The method establishes communication between GPUs, implements the protocol, maps memory apertures, and executes the synchronization process. This approach leverages existing infrastructure for efficient, decentralized multi-GPU synchronization using N messages for N GPUs.
Owner:INTEL CORP