Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

426 results about "Graphical processing unit" patented technology

A graphics processing unit ( GPU) is a specialized electronic circuit designed to rapidly manipulate and alter memory to accelerate the creation of images in a frame buffer intended for output to a display device. GPUs are used in embedded systems, mobile phones, personal computers, workstations, and game consoles.

Virtual stylist

An example operation may include at least one of receiving, via a user interface of a device, an activation input from a user to initiate a session, capturing, by a camera of the device, a scan of a body of the user, wherein the capturing comprises recording at least one image and / or at least one video of the user, processing the at least one image and / or video to generate a three- dimensional model of the user comprising measurements and contours of the body, retrieving, from a database, at least one clothing item associated with the user, the at least one clothing item comprising dimensional attributes and texture attributes, rendering, by a graphics processing unit, the at least one clothing item onto the three-dimensional model to generate a visual representation, wherein the rendering simulates draping behavior, movement, and light interaction of the at least one clothing item relative to the three-dimensional model, and displaying, on the user interface, an interactive visualization comprising the visual representation of the three-dimensional model with the at least one clothing item from multiple viewing angles.
Owner:ELGORT PENELOPE

An integrated system for dam failure detection, early warning and evacuation support

An integrated system for detecting dam breaches, as an early warning system and for evacuation assistance, consisting of: A geoinformatics analysis module comprising a modeling processor with a graphics processing unit, configured to: analyze spatial, hydrological, and topographic data of a dam and its downstream area; generate hydrological models using probable maximum rainfall (PMP) data to determine probable maximum flood discharge (PMF); generate hydraulic models to simulate flood propagation and determine flood depth, flood extent, flow velocity, and flood arrival time; create flood vulnerability models, vulnerability models, and flood hotspot models using a multi-criteria decision-making (MCDM) framework; and generate evacuation plans for high-risk zones based on the models. A sensor monitoring subsystem includes: a distributed fiber optic sensor (DFOS) configured to monitor the structural load and deformation of a dam, and a radar level gauge configured to measure the water level in the reservoir; a microcontroller unit configured to receive and process sensor data from the sensor monitoring subsystem, to detect dam breaches based on abnormal structural movements detected by the DFOS or sudden drops in water level below predefined thresholds detected by the radar level gauge; a communication module configured to transmit warning messages when dam breaches are detected; a warning system configured to: send SMS notifications to residents in downstream high-risk zones via a GSM communication module, generate real-time alerts via an IoT platform on user devices, and activate siren systems in villages to issue loud alarms and voice announcements; and an evacuation assistance platform configured to provide real-time evacuation assistance in emergencies due to dam breaks, displaying site maps with nearby emergency shelters, evacuation routes, hospitals, government buildings and aid centers, showing hydrological layers including flood extent, water depth, flow velocity and arrival time of the water, and enabling real-time location tracking in relation to safety zones and flood-prone areas.
Owner:DONGALE TUKARAM DR KOLHAPUR +4

Neural network using dynamically compressed and decompressed weights

A method for training or performing inference using a neural network involves performing per-layer decompression and compression of neural network weights. More particularly, compressed weights are retrieved for a particular layer of the neural network. The weights correspond to neurons in the layer. The compressed weights are decompressed, and input data for that layer is subsequently processed using the decompressed weights. This dynamic decompression and recompression of weights allows memory, and in particular random access memory of graphical processing units, to be efficiently used.
Owner:ROYAL BANK OF CANADA

Methods and systems for dynamically optimizing and modifying allocation of virtual graphical processing units

A method for dynamically optimizing and modifying allocation of at least one disaggregated graphical processing unit (GPU) to at least one process via at least one virtual graphical processing units (vGPU) includes receiving, by a resource manager, from a runtime component executing on a first computing device, a request for allocation, to a process executing on the first computing device, of access to a GPU in the first computing device. The method includes allocating, by the resource manager, to the process, access to at least one vGPU associated with a second GPU on a second computing device. The method includes transmitting, by the resource manager, to the runtime component, an identification of the at least one vGPU identifying the at least one vGPU as located on the first computing device. The method includes instantiating, by the runtime component, the process with access to the at least one vGPU.
Owner:EXOSTELLAR INC

Decentralized uplink detection through expectation propagation

Apparatuses, systems, and techniques to detect uplink data in a multi-user multiple input multiple output wireless communication system. In at least one embodiment, uplink data is detected using one or more graphics processing units to perform parallel computations in order to form a consensus belief about information received by one or more base stations in a wireless network.
Owner:NVIDIA CORP

Graphic processing unit cooperative processing method and system for accelerating three-dimensional generative model

The invention discloses a graphic processing unit cooperative processing method and system for accelerating a three-dimensional generative model. The model comprises a diffusion processing stage and a Gaussian splash rendering stage. In the diffusion processing stage, the time step number and the data precision calculated by the neural network are dynamically adjusted according to the visual angle sensitivity, and a specific graphic processing unit core is scheduled for execution. In the Gaussian splash rendering stage, pixel block tasks are predicted and intelligently scheduled to a stream processor based on historical data; and in the stream processor, the processing sequence of the Gaussian ball by the rendering unit is optimized according to the geometric distance, and the close-range unit is preferentially processed. During back propagation, the built-in logic of the stream processor preaggregates the gradient contributions of the same Gaussian ball, and then updates the gradient contributions through a single write operation. Through collaborative optimization of software and hardware, the efficiency bottleneck of the 3D generation model on a graphic processing unit is effectively solved, the calculation speed, the throughput and the hardware utilization rate are remarkably improved, and the model generation time is shortened.
Owner:SHANGHAI JIAOTONG UNIV

Orchestration of ai model deployment on multi-GPU systems

Apparatuses, systems, and frameworks for provisioning of efficient pipelines capable of multi-model inference and data processing using multiple processing units, including streaming data applications. The disclosed techniques include, during an initialization stage, assigning a plurality of machine learning models (MLMs) for execution on graphics processing units (GPUs), allocating memory space, on a hub GPU, to the plurality of MLMs, storing input data on the hub GPU before transferring the input data to other GPUs for execution. During an execution stage, output data is initially stored on GPUs that generated the output data before transferring the output data to the hub GPU.
Owner:NVIDIA CORP

Intelligent cooperative high-power proton exchange membrane fuel cell life prediction and fault diagnosis method and system

The invention relates to the technical field of fuel cell monitoring and health management, and discloses an intelligent cooperative high-power proton exchange membrane fuel cell life prediction and fault diagnosis method and system, and the method comprises the following steps: S1, collecting multi-dimensional data of a stack key region in real time through a distributed sensing array; s2, removing noise and abnormal values of the collected data, and carrying out feature extraction; s3, constructing a dual-model collaborative prediction framework; s4, inputting the real-time features into the double models; s5, generating a maintenance strategy and early warning; according to the invention, special hardware accelerators such as a graphic processing unit or a field-programmable gate array are integrated on the edge computing module, so that the processing speed and efficiency of high-frequency and multi-source sensor data at a local end are remarkably improved; a distributed computing framework is adopted in a cloud analysis module to process mass feature data transmitted by an edge end, so that the capability and scale of the system for performing deep analysis on the health state of the fuel cell, accurately predicting the remaining service life and quickly tracing the fault are enhanced.
Owner:INNER MONGOLIA YIPAI HYDROGEN ENERGY TECH CO LTD

Managing GPU resources on a container orchestration platform

A technique manages computing resources on a container orchestration platform. Such a technique involves establishing a pool of computing resources on the container orchestration platform. Such a technique further involves, after the pool of computing resources is established, receiving graphics processing unit (GPU) provisioning requests (GPRs) which identify workspaces. Such a technique further involves allocating computing resources from the pool to the workspaces identified by the GPRs based on a set of GPR prioritization policies.
Owner:AVESHA INC

Processing unit fault detection

Various example embodiments of a processing unit fault detection capability are presented. The processing unit fault detection capability may be configured to support detection of faults in processor cores of a processing unit based on arrangement of the processor cores to form a data processing pipeline and monitoring of the processor cores of the data processing pipeline based on monitoring of the data processing pipeline (e.g., based on propagation of heartbeat messages via the data processing pipeline). The processing unit fault detection capability may be configured to support detection of faults in processor cores of various types of processing units, such as central processing units (CPUs), graphics processing units (GPUs), network processing units (NPUs), or the like.
Owner:NOKIA SOLUTIONS & NETWORKS OY

Multi-speech synthesis model bearing method and device based on virtual GPU

The invention provides a multi-speech synthesis model bearing method and device based on a virtual GPU, and relates to the technical field of graphics processing units, and the method comprises the steps: carrying out the virtualization processing of a physical graphics processing unit, dividing the physical graphics processing unit into a plurality of virtual processing units with independent video memories and calculation quotas, and combining a resource scheduling mechanism, and deploying the speech synthesis language model instances in a plurality of service containers, and constructing a plurality of speech synthesis model bearing units. After a voice synthesis request is accessed, the scheduling module carries out load balancing according to the request connection number of each bearing unit, the request is distributed to a target bearing unit with the minimum connection number, and a voice generation task is completed by a virtual processing unit bound with the target bearing unit. According to the invention, resource division can be carried out on the physical graphic processing unit, and efficient operation of the multi-speech synthesis model is realized.
Owner:ZHEJIANG RONGQI MANUFACTURING TECHNOLOGY CO LTD

Techniques for implementing multiple large language models on a single physical computing device

Techniques for executing multiple large language models (LLMs) on a single physical computing device, including receiving a task to be performed by at least a first LLM and a second LLM, creating a plurality of virtual compute nodes on the single physical computing device for processing at least some of first set of sub-tasks to be performed by the first LLM and second set of sub-tasks to be performed by the second LLM, wherein each one of the plurality of virtual compute nodes includes one or more graphics processing unit (GPU) cores allocated for performing a sub-task of the first set of sub-tasks or the second set of sub-tasks, and executing at least some of the first set of sub-tasks and the second set of sub-tasks substantially in parallel across the plurality of virtual compute nodes using shared memory. Inferences between the first and second LLMs may be shared.
Owner:DESAI AMBARISH AJIT +2

Intelligent chart visualization method based on model control protocol

The invention discloses an intelligent chart visualization method based on a model control protocol. The method comprises the following steps: step 1, establishing connection between a visual rendering client and an intelligent analysis server; 2, receiving a natural language visualization instruction of a user, and obtaining source data; 3, intention recognition is conducted on the natural language visualization instruction, and feature extraction is conducted on source data; 4, determining an adaptive visual chart type; 5, generating a model control protocol instruction packet; 6, receiving and analyzing a model control protocol instruction packet; and step 7, constructing a three-dimensional grid model, and submitting the three-dimensional grid model to a graphic processing unit for rendering display. According to the method, an artificial intelligence natural language understanding technology and a three-dimensional graph programming visualization technology are combined, and the method has the advantages of chart self-adaption, visualization and scene fusion, real-time and efficient rendering and the like.
Owner:ZHEJIANG UNIV

Real-time object detection from decompressed images

A system and method for deploying machine learning models with consistent fixed-point arithmetic processing. A computing device in a vehicle receives a trained machine learning model that was trained using training images processed with fixed-point arithmetic operations. The computing device receives images from a camera and processes the received images using fixed-point arithmetic operations that are consistent with those used during training. The machine learning model is executed using the processed images to detect objects such as vehicles, pedestrians, persons on bikes, stop lights, or stop signs. The computing device may comprise specialized hardware including a graphics processing unit (GPU), digital signal processing hardware decoder, or single-purpose hardware decoder such as an application-specific integrated circuit (ASIC) or field-programmable gate array (FPGA). The system may generate alerts or determine vehicle operation changes based on detected objects, and may operate in surveillance mode when the vehicle is parked.
Owner:NETRADYNE INC

Quantum-based dispatch of workgroups

Quantum-based dispatch of workgroups is described. An example of an apparatus includes a computer memory to store data for processing, including data for an application; and one or more processors including a graphical processing unit (GPU), the GPU including multiple chiplets, each of the multiple chiplets including compute containers and a cache, each compute container including a plurality of processing resources, and a dispatcher for dispatching workgroups to the processing resources of the GPU, wherein dispatching workgroups includes dispatching workgroups for the application according to a selected workgroup quantum, the selected workgroup quantum having a certain size and shape.
Owner:INTEL CORP

System and method for executing fused neural-network layer architectures

The present disclosure provides a system and method for executing fused neural network layers using a graphics processing unit (GPU). The fused neural network layer combines multiple neural network operations into a single GPU kernel function, for efficient utilization of a GPU shared memory to reduce global memory transactions. The fused neural-network-layer system configures GPU thread blocks to iterate tiles in the GPU shared memory across portions of input tensors, and to perform a sequence of neural network layer operations using the tiles, before storing the results in an output tensor. The neural network layer operations can include element-wise, normalization, and pooling operations. The system also supports fused layers with nested traversals of input tensors, such as matrix multiplication, convolution, and attention mechanisms. The fused neural-network-layer system improves performance and reduces memory overhead compared to executing each layer as a separate GPU kernel, thereby enabling faster training and inference times.
Owner:BAYRIS INC

Systems and methods for enhancing execution of interpreted computer languages

Systems and methods are provided that incorporate a compiler configured to convert interpreted language code (e.g., Python) into native machine code. According to some embodiments, the system generates the native machine code into a format that is consistent with known infrastructure. The native machine code can be converted into a format based on a low level virtual machine “LLVM” infrastructure. In various embodiments, the system enables a compiler framework that improves execution of code for interpreted languages. According to one embodiment, the system can be tailored for execution on specific processors, for example, a graphics processing unit (“GPU”) that is optimized for highly parallel computations.
Owner:EXALOOP INC

Cooperative parallel memory allocation

Apparatuses, systems, and techniques to perform multi-threaded memory allocation in parallel by one or more software programs being performed on a parallel processing unit (PPU), such as a graphics processing unit (GPU), or any other processing unit capable of supporting multi-threaded software execution. In at least one embodiment, one or more software programs expressed in part by code using an application programming interface for parallel computing, such as CUDA, perform allocation, search, and deallocation of memory efficiently and in parallel on a GPU.
Owner:NVIDIA CORP

Artificial intelligence-based business management system for optimizing real-time decisions

An artificial intelligence-based business management system for real-time decision optimization, consisting of: a data acquisition unit configured to ingest structured, semi-structured and unstructured data streams from enterprise resource planning (ERP) systems, customer relationship management (CRM) platforms, Internet of Things (IoT) devices, external market feeds and financial transaction systems, encrypting, timestamping and verifying the data prior to further processing; a graph processing unit that is communicatively connected to the said acquisition layer, wherein the unit is configured to encode heterogeneous data into dynamic graph structures comprising nodes representing business entities and edges representing transaction or relationship dependencies, wherein the unit is further configured to perform deduplication, metadata tagging and real-time data synchronization; a decision optimization unit operationally linked to the graph processing unit, wherein the decision optimization unit includes modules for reinforcement learning, modules for Bayesian optimization and multi-objective solvers configured to simulate multiple alternative decision paths and select an optimal path based on performance indicators such as cost efficiency, resource utilization, customer satisfaction and risk minimization; a real-time inference control unit with specialized hardware cores, including at least one graphics processing unit (GPU), a field-programmable gate array (FPGA) and an application-specific integrated circuit (ASIC), wherein the accelerator performs inference tasks of the decision optimization unit with a latency in the millisecond range; a diagnostic processing unit configured to generate causal diagrams, feature mapping maps, and interpretable result summaries according to the optimization outputs; and a control interface unit configured to transmit optimized decisions to process controls within the enterprise, robot actuators, planning systems or interactive dashboards, with the interface supporting bidirectional communication for higher-level interventions, error feedback and triggers for re-optimization.
Owner:ABUELENAIN EMAD EDDIN AHMED +4

Image processing method and device and medium

The embodiment of the invention relates to an image processing method and device and a medium. The method proposed herein includes: generating, by a graphics processing unit, a first reference frame, a second reference frame, and motion vector information corresponding to the first reference frame; transmitting the motion vector information to a digital signal processor by a graphic processing unit; the digital signal processor divides the motion vector information into a plurality of groups; processing the plurality of groups in parallel by using a plurality of processing cores of the digital signal processor, each processing core being configured to: write a group of motion vectors indicated by a corresponding group to a group of projection positions of the group of motion vectors in the interpolated frame by using a scatter instruction; respectively reading source pixel information from the first reference frame and the second reference frame based on the written group of motion vectors by utilizing a collection instruction; the source pixel information is fused by a digital signal processor to generate an interpolated frame. In this way, according to the embodiment of the invention, the generation efficiency of the interpolation frame can be effectively improved.
Owner:VASTAI TECH (SHANGHAI) INC

Graphics processing unit for processing primitives using rendering space subdivided into plurality of tiles

A graphics processing unit (GPU) and method for processing primitives using a rendering space subdivided into a plurality of tiles are described. The GPU includes a partitioning module that includes a plurality of tile pipelines and a tile arbiter. The tile arbiter receives tile-primitive indications, wherein each of the tile-primitive indications indicates an association between a tile and a set of one or more primitives. The tile arbiter determines to which of the tile pipelines each of the tile-primitive indications is to be sent, according to a dynamic allocation scheme that does not have a fixed mapping between a tile and tile pipeline. The tile arbiter transmits each of the tile-primitive indications to the tile pipeline that is determined for the tile-primitive indication.
Owner:IMAGINATION TECH LTD

Boot survivability for graphics processing unit

A system that includes a graphics processing unit (GPU) that includes: at least one processor and circuitry to: based on failure of the GPU to load boot firmware, operate as a survivability agent to allow for the GPU to boot to a configuration wherein a host system is to communicate with the GPU to determine the failure of the GPU to load boot firmware and to load second boot firmware for access by the GPU. In some examples, the GPU includes an input output (IO) subsystem and to boot to the configuration, the circuitry is to provide the host system with access to an indicator of failure of the GPU and access to the host system to load the second boot firmware into a boot storage accessible to the GPU.
Owner:INTEL CORP

System for integrated data provenance and reconciliation across heterogeneous financial systems

ActiveDE202025104886U1FinanceData streamGate array
A system for integrated data provenance and reconciliation across heterogeneous financial systems, the system includes: a data ingestion module implemented as a hardware interface and configured with high-throughput adapters for structured and unstructured data ingestion from multiple heterogeneous financial subsystems, including bank ledgers, securities trading platforms, risk management databases, and regulatory reporting systems; a schema harmonization processing unit consisting of a dedicated FPGA (Field Programmable Gate Array) structure configured to perform real-time schema alignment, metadata normalization, and semantic mapping across different data models; a lineage tracking processor array implemented on application-specific integrated circuits (ASICs) and configured to encode dataflow transitions into a graph-encoded structure in which each node represents a transformation and each edge represents a dependency; a reconciliation computation cluster unit embodied as a set of graphics processing units (GPUs) configured to execute parallelized reconciliation techniques, anomaly detection routines, and balance verification processes between transformed financial records and reference records; a cryptographic verification unit comprising secure key storage hardware, quantum-resistant encryption modules, and tamper-resistant enclaves configured to generate, anchor, and verify cryptographic signatures for provenance and reconciliation events; a secure storage subsystem configured in a WORM (write-once, read-many) configuration with hardware-based integrity locks for storing immutable provenance and reconciliation logs; a controller with a dedicated visualization processor configured to generate interactive dashboards and drill-down analyses of lineage charts and voting results; and a modular chassis structure consisting of a high-speed backplane interconnect, secure power distribution modules, and tamper-evident enclosures, enabling the system to achieve real-time scalability, hardware-level security, and auditability of financial data flows across heterogeneous infrastructures.
Owner:SUDHANSHU JAIN LITHIA

Graphics processing unit for processing primitives using rendering space subdivided into plurality of tiles

A graphics processing unit (GPU) and method for processing primitives using a rendering space subdivided into a plurality of tiles are described herein. The GPU includes geometry processing logic that includes a plurality of geometry pipelines and a block back-end module. The geometry pipeline is configured to receive a batch of primitives of the sequence of primitives. Each of the geometric pipelines includes: one or more geometric processing modules configured to perform one or more geometric processing functions on the primitives in a batch of primitives received at the geometric pipeline; and a tile front-end module configured to determine, for each tile in a set of one or more tiles, one or more tile-primitive indications, the one or more tile-primitive indications indicate which of the primitives of the batch of primitives received at the geometric pipeline are present within the tile.
Owner:IMAGINATION TECH LTD

Vehicle display instrument control system

The invention relates to the technical field of intelligent cabins and functional safety, in particular to a vehicle display instrument control system, which comprises a prospective sensing unit, a central processing unit, a central processing unit and a central processing unit, and is characterized in that the prospective sensing unit is used for generating multi-dimensional cabin system state data; the risk prediction unit is used for calculating a system risk potential index; the scheduling decision-making unit is used for comparing and analyzing the system risk potential index and a preset risk threshold value; when the system risk potential index exceeds a risk threshold value, determining that the system enters a collaborative mode and calculating a collaborative distribution factor; when the system risk potential index does not exceed the risk threshold value, determining that the system maintains an isolation mode; the rendering adjustment unit is used for generating an adjusted target rendering load index of the infotainment domain; the closed-loop optimization unit is used for carrying out self-adaptive correction on the weight of the risk state quantitative prediction model; the hidden danger of safety function display delay caused by resource competition of the graphic processing unit is effectively solved, and the function safety and reliability of the intelligent cabin are greatly improved.
Owner:NINGBO GUORUI NEW ENERGY TECH CO LTD

System and method for fine-tuning rotated outlier-free large language models for effective weight-activation quantization

A computing device includes at least one processor, one or more non-transitory computer-readable storage media, a system for fine-tuning a large language model under low-bit weight-activation quantization. The computing device further comprises a graphics processing unit (GPU), a neural processing unit (NPU), or a tensor processing unit (TPU). The hardware interface module of the system is configured to load a low-bit model representation from the memory module and transmit the model representation to the GPU, NPU, or TPU for inference execution.
Owner:THE HONG KONG UNIV OF SCI & TECH

COMMUNICATION OPTIMIZATION FOR MoE BY OFFLOADING EXPERTS TO NICs

Embodiments herein describe a system including a plurality of hardware accelerators including at least one mixture-of-experts (MoE) layer having multiple experts and a plurality of network interface cards (NICs) coupled to the plurality of hardware accelerators, wherein at least one expert of the multiple experts is offloaded from the plurality of hardware accelerators to the plurality of NICs. The plurality of hardware accelerators may be graphics processing units (GPUs). In one example, a subset of the multiple experts are selectively offloaded from the plurality of GPUs to the plurality of NICs based on memory and computational capacity available on the plurality of NICs. In another example, the multiple experts are designated as either hot experts or cold experts. The cold experts are offloaded from the plurality of GPUs to the plurality of NICs and the hot experts are duplicated for each of the plurality of GPUs.
Owner:ADVANCED MICRO DEVICES INC +1

Large model deployment and service method and device for realizing intelligent customer service based on container arrangement, equipment and medium

The invention discloses a large model deployment and service method and device for realizing intelligent customer service based on container arrangement, equipment and a medium, and relates to the field of artificial intelligence, and the method comprises the steps: constructing a standardized container mirror image which comprises an inference engine, a model acceleration library and a model encryption and decryption assembly and is adaptive to an intelligent customer service scene; establishing a graphic processing unit resource management system in the container arrangement cluster; storing model data of the intelligent customer service model to a target storage system, and determining a mounting mechanism corresponding to the model data; generating a container arrangement and deployment file based on a graphic processing unit resource management system, the operation demand of the standardized container mirror image and a mounting mechanism corresponding to the model data; and submitting the container arrangement deployment file to a container arrangement platform, so that the container arrangement platform carries out intelligent customer service model deployment, and the deployed intelligent customer service model is utilized to receive a user request and return a corresponding target response. According to the method, the deployment complexity is reduced, and the model deployment efficiency and stability are improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

GPU memory pool manager for virtual shared GPU memory pooling

Disclosed methods provide a virtualized shared graphics processing unit (GPU) memory pool (VSGMP) to virtual machines running on an information handling system. The VSGMP may be implemented with a GPU memory pool (GMP) manager, featuring logic for abstracting the GPU memory pool from the physical memory resources of two or more GPUs. The GMP manager logic may be supported by a lightweight secure operating system (LSOS) capable of enabling functionality for virtualization and other use cases. Disclosed methods manage GPUs in an information handling system featuring two or more GPUs running virtual machines (VMs). When a GPU is assigned to a VM, disclosed methods perform one or more GPU resource allocation operations that support virtual shared pooling of the physical memory resources of two or more GPUs. In at least some embodiments, the GPU allocation operations include allocating at least some non-memory resources of the GPU exclusively to the applicable VM.
Owner:DELL PROD LP

AIGC large model training method and system based on domestic GPU graphics card

The invention provides an AI GC large model training method and system based on a domestic GPU graphics card, and belongs to the technical field of AI GC large model system.The AI GC large model training method comprises the steps that a hardware feature data set of a domestic graphics processing unit is obtained, the data set comprises key indexes of instruction execution efficiency and a data transmission mode, and calculation boundary conditions under hardware constraints are determined; according to the quantitative result of the semantic variation, constructing a triggering condition of parameter adjustment, if the semantic variation exceeds a preset threshold value, starting a parameter dynamic updating mechanism, and determining updated parameter configuration; adjusting a data batch processing sequence in a training process according to the updated parameter configuration in combination with an incremental training efficiency requirement, obtaining a dependency relationship between batches, and judging a completion condition of a training task; the parallel computing logic of the training process is optimized through the dependency relationship among batches, the parameters are iteratively updated by adopting a gradient descent method, an intermediate result after each iteration is obtained, and the direction of training convergence is determined. According to the method, the calculation efficiency of a domestic graphic processing unit can be effectively improved, efficient processing of dynamic data and self-adaptive updating of model parameters are achieved, and a new solution is provided for optimal execution of complex calculation tasks.
Owner:SHENZHEN YUANLIU COMPUTING TECHNOLOGY CO LTD