Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

602 results about "Graphics processing unit" patented technology

A graphics processing unit (GPU) is a specialized electronic circuit designed to rapidly manipulate and alter memory to accelerate the creation of images in a frame buffer intended for output to a display device. GPUs are used in embedded systems, mobile phones, personal computers, workstations, and game consoles. Modern GPUs are very efficient at manipulating computer graphics and image processing. Their highly parallel structure makes them more efficient than general-purpose central processing units (CPUs) for algorithms that process large blocks of data in parallel. In a personal computer, a GPU can be present on a video card or embedded on the motherboard. In certain CPUs, they are embedded on the CPU die.

Layout design rule checking method, system, equipment, medium and product

The invention relates to an application method of a GPU (Graphics Processing Unit) in design rule inspection in the field of integrated circuits, in particular to a layout design rule inspection method, system and equipment, a medium and a product. The layout design rule checking method comprises the following steps: inputting initial layout data, analyzing a design rule checking task to be executed, and marking the priority and complexity level of the task; the method comprises the following steps: dividing layout data into a plurality of data blocks according to regions on the basis of regional characteristics of the layout data, distributing tasks and corresponding data blocks to a GPU computing core according to task priorities and complexity levels, and distributing new data blocks according to the overall load condition of the GPU computing core after the GPU computing core completes data block processing, and repeating the steps until all tasks are executed, and outputting a layout design rule check result. The system, the computer equipment, the computer readable storage medium and the computer program product have the same beneficial effects as the layout design rule checking method.
Owner:HUAXIN GIANTS (HANGZHOU) MICROELECTRONICS CO LTD

Data processing method, product, electronic equipment and computer readable storage medium

The invention discloses a data processing method, a product, electronic equipment and a computer readable storage medium, relates to the technical field of computers, and aims to solve the problems of large GPU (Graphic Processing Unit) space occupation and high access delay in data processing in related technologies. Reading corresponding historical key value cache data from a key value cache of the memory extension equipment according to a historical key value cache data acquisition request sent by a host end; returning the historical key value cache data to the host end, and enabling the host end to process the current input data based on the historical key value cache data by adopting a local model to obtain key value data; and obtaining key value cache data corresponding to the key value data sent by the host end, and sending the key value cache data to a key value cache of the memory extension equipment for storage. The key value cache data is stored in the key value cache of the memory extension equipment, and the historical key value cache data is read from the key value cache, so that the occupation of the memory space of the GPU and the consumption of the memory bandwidth can be reduced, and the access delay is reduced.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Method and apparatus for supporting distributed graphics and compute engines and synchronization in multi-dielet parallel processor architectures

This disclosure describes supporting distributed graphics and compute engines in a multi-dielet processor, such as, for example, a multi-dielet graphics processing unit (GPU), architectures and synchronization in such architectures. Each multi-dielet processor includes a hardware-implemented remapping capability and / or a hardware-implemented memory barrier capability.
Owner:NVIDIA CORP

Rendering processing method and electronic device

PCT designated stageWO2026000316A13D-image renderingComputer hardwareGraphics
The embodiments of the present application relate to the technical field of terminals. Provided are a rendering processing method and an electronic device. The method comprises: determining a plurality of rendering instructions corresponding to first graphic data to be rendered of a first application program, wherein the plurality of rendering instructions include a first rendering instruction; and then reading, from a graphics processing unit (GPU) memory, a first rendering resource corresponding to the first rendering instruction, and executing the plurality of rendering instructions on the basis of rendering resources respectively corresponding to the plurality of rendering instructions, so as to render the first graphic data. In this way, a rendering resource is read from a GPU memory, so as to reduce data interaction between a GPU and a main memory, thereby reducing the bandwidth consumption of a terminal device.
Owner:HONOR DEVICE CO LTD

Thread block scheduling module, general purpose computing graphics processing unit, device and product

The invention provides a thread block scheduling module, a general-purpose computing graphic processing unit, equipment and a product, and relates to the technical field of data process.The method comprises the steps that an operation data analysis unit conducts data analysis on all first thread blocks of a computing task issued by a host, and second thread blocks sharing data with all the first thread blocks are determined; the mapping relation between the first thread block and the second thread block is written into a thread block data mapping table of the queue management unit; the resource management unit is used for recording available resources and used resources of each streaming multiprocessor and information of a thread block currently allocated to each streaming multiprocessor; and the thread block allocation unit is used for allocating the thread blocks of the first target thread block to the corresponding stream multiprocessors according to the thread block data mapping table of the first target thread block under the condition of detecting that any stream multiprocessor recorded by the resource management unit has idle available resources.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Model performance test method and device, electronic equipment and storage medium

The invention discloses a model performance test method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence. The theoretical maximum lexical throughput of a target large language model is calculated based on the video memory bandwidth of a graphics processor, the model parameter quantity, the byte number corresponding to the quantization precision and the video memory bandwidth utilization rate; meanwhile, the benchmark performance throughput is obtained, a theoretical corresponding first concurrency number is calculated in combination with the theoretical maximum lexical unit throughput and the concurrency competition loss coefficient, then the model test is executed based on the first concurrency number to obtain the actual maximum lexical unit throughput and a corresponding second concurrency number, and a model performance test result is generated. The problems that in the prior art, due to the fact that manual testing is conducted depending on manual intervention, a continuous approaching attempt mode is adopted, a reasonable test starting point is not deduced in combination with hardware core bottlenecks and key parameters, evaluation is time-consuming and labor-consuming, the result is prone to being affected by artificial factors, and accuracy and consistency are poor can be solved.
Owner:JINAN INSPUR DATA TECH CO LTD

High-performance CPU-GPU (Central Processing Unit-Graphics Processing Unit) coprocessing architecture of audio frequency integrated signal processor

The invention discloses a high-performance CPU-GPU (Central Processing Unit-Graphics Processing Unit) cooperative processing architecture of an audio frequency integrated signal processor, belonging to the technical field of audio frequency integrated signal processing, the architecture is based on a dynamic priority scheduling model and realizes efficient collaboration of a CPU and a GPU through hierarchical resource management and protocol level optimization, a hardware layer adopts a multi-GPU cluster and a distributed storage node, and the CPU-GPU cooperative processing architecture has the advantages that the multi-GPU cluster and the distributed storage node are integrated; high-concurrency task processing is supported; the transmission layer fuses RapidIO and an Ethernet protocol, and adapts to a short frame control signal and a long packet data stream respectively; and the application layer calculates task resource demands through dynamic priority weights, and ensures low time delay of key tasks in combination with a normalized allocation algorithm. According to the architecture, in an audio signal processing scene, the task scheduling efficiency is improved by 40%, the short frame transmission delay is as low as 0.5, the throughput of a long data stream reaches 100 Gbps, and the requirements for real-time performance and calculation precision in a complex acoustic environment can be met.
Owner:CHINA SHIP DEV & DESIGN CENT +1

Color setting method, graphics processing unit and system on chip

The embodiment of the invention provides a color setting method, a graphic processing unit and a system on chip. The method applied to GPU user mode driving comprises the steps that under the condition that a rendering preparation instruction is received, a first bit field of a sampling descriptor is set as a first identifier, the first identifier is used for pointing to a preset table located in a memory, and the table is used for storing boundary colors with the color format exceeding the bit number of the sampling descriptor; writing a target color carried in the rendering preparation instruction into the preset table; acquiring a second identifier of the target color in the preset table; and writing a second identifier of the target color in the preset table into a second bit field of the sampling descriptor, and sending the sampling descriptor to GPU hardware, so that the GPU hardware realizes texture boundary color setting according to the sampling descriptor. According to the embodiment of the invention, the storage format and precision of the target color can be expanded, so that the setting of the texture boundary color can be improved, and the user experience is improved.
Owner:LOONGSON TECH CORP

Thread group scheduling method and device for GPU (Graphics Processing Unit), graphics processing unit and equipment

The invention discloses a thread group scheduling method and device for a GPU (Graphics Processing Unit), the GPU and equipment. The method comprises the following steps of: 1) receiving a scheduling request of a thread group; 2) task type priority scheduling; according to a thread group weight value set by a user, obtaining execution priorities of the vertex thread group and the fragment thread group in the current scheduling period; 3) instruction type priority scheduling; pre-analyzing to-be-executed instructions of the thread group, and determining a priority sequence of schedulable instruction types; and 4) priority scheduling of the thread groups: selecting the thread group with the highest priority for scheduling according to the specified task type priority and instruction type priority in combination with the thread group generation time. The invention provides a thread group three-level scheduling strategy so as to improve the instruction throughput rate and the key task response speed of the GPU under the complex load.
Owner:WUHAN LINGJIU MICROELECTRONICS CO LTD

AI-powered cybersecurity system for regulatory compliance in energy distribution

A system for AI-supported cybersecurity and regulatory compliance in energy distribution networks, consisting of: a hardware-embedded data acquisition module configured to intercept, capture, and time-stamp operational data streams and to control data traffic from SCADA (Supervisory Control and Data Acquisition) systems, AMI (Advanced Metering Infrastructure) systems, and energy management systems (EMS) via multiple communication protocols without operational latency; an FPGA-based deep packet inspection unit coupled with the data acquisition module, wherein the FPGA firmware is configured to perform line rate filtering, protocol decomposition and metadata extraction of the acquired data and forwards preprocessed packet data to an AI processing unit; an AI processing unit consisting of a multi-core central processing unit (CPU), a dedicated AI accelerator selected from a graphics processing unit (GPU) or a tensor processing unit (TPU), and a volatile memory buffer; a response orchestration module that is communicatively coupled with network management devices and operations controllers, wherein the response orchestration module is configured to perform automated security and compliance remediation measures, including network isolation of compromised segments, enforcement of protocol encryption, and privilege revocation; and an immutable audit logging subsystem configured to record all detected events, compliance assessments, and corrective actions in a blockchain-based distributed ledger, with each log entry cryptographically anchored with a secure hash value and digitally signed with keys stored in a secure hardware enclave.
Owner:ALIF MUHAMMAD +11

GPU check point storage method in large model distributed training

The invention relates to a GPU (Graphics Processing Unit) check point storage method in large model distributed training, which comprises the following steps of: 1, loading check point configuration information when training is started; 2, monitoring a back propagation completion signal of each layer in the training process, and immediately triggering check point fragment storage operation of a certain layer after parameter updating of the layer is completed; 3, the model state of the layer is asynchronously copied to a CPU memory from a GPU memory, and consistency verification is carried out; 4, asynchronously storing the check point fragments from the CPU memory into a persistent storage; 5, when training needs to be recovered, the check point file is loaded from the persistent storage, integrity verification is carried out, all layers of fragmented data are recombined into a complete model state, and training is recovered. And the synchronous blocking and I / O bottleneck of the check points in the training process are reduced, so that the training efficiency is improved, and the fault recovery time is shortened.
Owner:NANJING UNIV OF POSTS & TELECOMM

Real-time semantic segmentation method based on hierarchical feature fusion and channel attention enhancement

The invention relates to a real-time semantic segmentation method based on hierarchical feature fusion and channel attention enhancement, and belongs to the technical field of computer vision. The method comprises a training stage and a reasoning stage. In a training stage, firstly, data enhancement is performed on an input image through an image enhancement strategy, then a semantic segmentation model taking hierarchical feature fusion and channel attention enhancement as a core is constructed, and training is performed by adopting a deep supervision technology to obtain an optimal weight. In the reasoning stage, the optimal weight and the model are loaded, and real-time semantic segmentation is carried out on the image by utilizing a GPU (Graphics Processing Unit). The semantic segmentation model can effectively solve the problems of local detail information loss and intra-class semantic tag inconsistency, and has higher reasoning speed and higher segmentation precision compared with other similar methods.
Owner:KUNMING UNIV OF SCI & TECH

Self-adaptive compressed data direct query method, system and equipment based on GPU (Graphics Processing Unit) and medium

The invention relates to a GPU-based adaptive compressed data direct query method, system and device and a medium, and the method comprises the steps: carrying out the data compression of input column data through employing an adaptive compression strategy, and obtaining a block-level structure suitable for GPU storage and calculation; loading the compressed data into a GPU memory, and performing Tile-level memory management by taking Tile as a basic scheduling unit; and according to the received query statement, executing direct query of the compressed data on the Tile level through the GPU, and outputting a query result. By designing a Tile-level direct query framework, a hardware-aware memory management and control flow coordination mechanism and a self-adaptive compression strategy, high-performance query execution is realized without decompression, and the computing potential of the GPU is fully released. The method can be widely applied to the technical field of big data processing.
Owner:RENMIN UNIVERSITY OF CHINA

Resource configuration methods, resource configuration system, electronic device and storage medium

Disclosed in the present application are resource configuration methods, a resource configuration system, an electronic device and a storage medium. A resource configuration method comprises: acquiring configuration requirement information corresponding to a computing power service; on the basis of the configuration requirement information, performing resource configuration on a general-purpose server and a graphics card resource pool to obtain a configuration result, wherein the general-purpose server is obtained on the basis of processor resource configuration, and the graphics card resource pool comprises a plurality of graphics processing units that support a hot-swapping function; and on the basis of the configuration result, providing a first instance for the computing power service, wherein the first instance comprises a processor resource allocated by the general-purpose server and a processor resource allocated by the graphics card resource pool. The present application solves the technical problems in the related art of low resource utilization rate and poor flexibility of the entire unit caused by heterogeneous models being constrained by the fixed configuration of processor resources of the entire unit.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Architectural integration of a system on a chip design

According to some aspects, there is provided a dynamic reconfigurable system-on-a-chip (SoC). The SoC comprises a central processing unit (CPU), a graphics processing unit (GPU), and an integrated circuit component. The integrated circuit component is programmed according to a first set of instructions to operate the SoC at least in part by controlling the CPU and the GPU to perform a first task. The integrated circuit component is further configured to, upon receiving a second set of instructions for performing a second task, reprogram the integrated circuit component according to the second set of instructions and control the CPU and the GPU to perform the second task.
Owner:FICGM LLC +4

Video memory control method, device, equipment and system and computer storage medium

The embodiment of the invention provides a video memory control method, device, equipment and system and a computer storage medium. The method comprises the steps that in response to an obtained drive upgrading request of an image processing unit (GPU), a before-upgrading page table corresponding to the GPU is obtained, and the before-upgrading page table comprises a mapping relation between a physical address and a virtual address of a GPU video memory; determining a data cleaning queue corresponding to the GPU, wherein the data cleaning queue is used for performing global cleaning on data in a GPU video memory; under the condition that the data cleaning queue is locked, driving upgrading operation is carried out on the GPU based on the driving upgrading request, an upgraded page table corresponding to the GPU is obtained, and the upgraded page table comprises a mapping relation between a physical address and a virtual address of a GPU video memory; and regulating and controlling data in the GPU video memory based on the pre-upgrading page table and the post-upgrading page table. In the embodiment of the invention, the driver upgrading operation of the GPU can be realized without operations such as video memory backup, shutdown, restart and the like.
Owner:ALIBABA CLOUD COMPUTING CO LTD

Link detection system and method, and electronic equipment

The invention discloses a link detection system and method and electronic equipment, and relates to the technical field of graphic processing units, and the system comprises main equipment which is used for receiving a test instruction sent by a processor; randomly generating a plurality of test data packets; and generating and sending a sending timestamp of the first test flow to the slave device based on the first identifier of the target end device, the second identifier of the master device and the plurality of test data packets. And the slave device is used for generating response information and feeding back the response information to the master device according to the received receiving timestamp of the second test flow, the second test flow and the sending timestamp. And the master device is used for recording a return timestamp of the response information when receiving the response information, and determining whether a communication link between the master device and the slave device is abnormal or not according to at least one of an error rate, a link rate, a receiving timestamp, a sending timestamp and the return timestamp included in the response information. The problem that the stability detection effect of the exclusive link between the GPUs is poor can be solved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Key-value cache reuse method, and related apparatus

Disclosed in the embodiments of the present application are a key-value cache reuse method, and a related apparatus, which can be applied to scenarios such as cloud technology, artificial intelligence, intelligent transportation, assisted driving and the Internet of Things. The method comprises: after acquiring an ith-round prompt, on the basis of historical token groups respectively generated by a plurality of tokens included in the ith-round prompt with the previous (i-1)th-round prompt, and according to token arrangement orders, matching tokens having the same token arrangement order, so as to obtain a prefix token group and remaining token groups; acquiring from a graphics processing unit a key-value cache of the prefix token group, and if the acquisition fails, acquiring from a central processing unit a key-value cache corresponding to the prefix token group; and sending the key-value cache of the prefix token group and the remaining token groups to an inference engine, such that the inference engine performs inference on the basis of the key-value cache of the prefix token group and the remaining token groups to obtain response content for the ith-round prompt. Thus, an inference engine directly reuses a key-value cache of a prefix token group without requiring recalculation, thereby improving the speed of inference computation.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Thermal power plant operation digital twin

The invention belongs to the field of digital twinning and energy industry intellectualization, relates to a thermal power plant three-dimensional operation state digital twinning system, and solves the technical problems of traditional operation and maintenance data islands, visualization deficiency and prediction lag. The system constructs a'multi-source data fusion-dynamic physical deduction-virtual-real interaction control 'system: a multi-source data fusion engine gets through a DCS / SIS barrier to realize 200,000 + data point / millisecond-level synchronization; a GPU (Graphics Processing Unit) drives a VFX Graph particle system and an LOD rendering technology to achieve 90fps (at) 4K flue gas flow simulation on RTX 4080; an AR instruction is directly connected with a PLC to construct a virtual-real bidirectional control chain, and a'monitoring-deduction-intervention 'closed loop is formed. Key indexes are as follows: fault early warning is greater than or equal to 72 h, fault positioning is 0.5 h / time, data delay is less than or equal to 10 ms, and model precision is + / -2 cm. And an edge-cloud collaborative architecture is adopted, and lightweight deployment of a Unity model is realized. According to actual measurement of a Jingfeng-Dalberg power plant, the emergency response is shortened from 26 minutes to 4.3 minutes, and a predictable intervention driven innovative paradigm foundation is laid for thermal power digitization.
Owner:INNER MONGOLIA HUIBO TECH ENG CO LTD

Distributed cache coherence protocol based on Ethernet, implementation method, device and system

The invention discloses an Ethernet-based distributed cache coherence protocol, an implementation method, an implementation device and an implementation system. A plurality of computing nodes are connected through a packet switching network. Each computing node comprises a CPU / GPU (Central Processing Unit / Graphics Processing Unit) and a local cache thereof, and is provided with a cache agent. The far-end memory is organized in home nodes, and each home node manages a part of physical address space and is equipped with a directory controller. When the CPU of the computing node accesses a far-end memory address and does not hit in the local cache, the CA of the computing node replaces the far-end memory address and communicates with the DC managing the address through the network so as to maintain the cache consistency of the data among all the nodes. Based on a cache consistency protocol of a directory, the CXL.cache consistency of a plurality of independent computing nodes can be maintained in a low-overhead and high-reliability mode on a high-delay and lossy packet switching network, and broadcast storm caused by a monitoring protocol is avoided.
Owner:SHENZHEN UNIVERSITY OF ADVANCED TECHNOLOGY

Techniques for accelerating queries using multiple graphics processing units

Described are examples for using multiple graphics processing units (GPUs) to accelerate a database query. Data for a database query can be loaded from the database into memories of multiple GPUs for parallel processing by the multiple GPUs. At least a portion of the data loaded into a memory for one of the multiple GPUs can be moved to a memory for a different one of the multiple GPUs. A compute process can be executed, via parallel processing on the multiple GPUs, for the query to perform data processing related to the database query.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Artificial intelligence model training method and device, electronic equipment and storage medium

The embodiment of the invention provides an artificial intelligence model training method and device, electronic equipment and a storage medium, and aims to reduce the cost of applying a quantum technology to AI model training and improve the training efficiency of an AI model. The method comprises the following steps: a central processing unit performs data preprocessing on original training data acquired from a memory; the original training data comprises a training set; executing the following iterative training steps on a to-be-trained artificial intelligence AI model until a preset iterative stop condition is met: inputting the training set into the to-be-trained AI model by the graphics processor, and obtaining a loss function corresponding to a current parameter of the to-be-trained AI model; the quantum processor encodes the loss function into a quantum state; the quantum processor constructs a quantum circuit based on the loss function of the quantum state, and measures the quantum gradient of the loss function through the quantum circuit; the quantum processor updates the current parameter based on the quantum gradient and a preset regularization gradient.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Radiator for high-power GPU (Graphics Processing Unit)

The invention relates to a radiator for a high-power GPU (Graphics Processing Unit), which comprises a cover plate, a splitter plate, a radiating cold plate and a GPU mainboard, and is characterized in that a pair of working medium flowing ports are formed in one side of the cover plate and are respectively used as a working medium inflow port and a working medium outflow port of the radiator; the flow dividing plate is provided with a flow dividing concave cavity communicated with the working medium flow inlet, a flow collecting concave cavity communicated with the working medium flow outlet and a heat exchange concave cavity communicated with the flow dividing concave cavity and the flow collecting concave cavity, and a flow dividing channel with an opening facing the flow dividing concave cavity and a flow collecting channel with an opening facing the flow collecting concave cavity are arranged in the heat exchange concave cavity; micro-structure fins which correspond to the flow dividing concave cavity, the heat exchange concave cavity and the flow collecting concave cavity and are used for enhancing heat exchange are arranged on the heat dissipation cold plate, a GPU core corresponding to the micro-structure fins is arranged on the GPU main board, and the GPU main board and the heat dissipation cold plate are tightly attached by smearing a heat conduction material. According to the invention, the flow stability of the micro-channel radiator under a high-pressure-difference working condition can be improved, and the junction temperature fluctuation of a chip can be reduced.
Owner:SHANGHAI INST OF TECH

Method and host machine for executing computing task through GPU (Graphics Processing Unit)

The invention provides a method for executing a computing task through a GPU and a host machine, the GPU is configured on a target host machine, a first confidential virtual machine containing a first client and a second confidential virtual machine containing a first agent program run on the target host machine, and the second confidential virtual machine loads an encryption and decryption kernel into the GPU in advance. The method comprises the following steps: a first client transmits a calculation task to a first agent program through an encryption channel pre-established with the first agent program; the first agent program encrypts the calculation task through a session key negotiated in advance with the encryption and decryption kernel, and sends an obtained first encryption result to the GPU; the GPU decrypts the first encryption result according to the session key, and restores and executes the calculation task to obtain a calculation result; encrypting the calculation result according to the session key, and sending an obtained second encryption result to the first agent program; and the first agent program decrypts the second encryption result according to the session key, and sends the restored calculation result to the first client through the encryption channel.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

GPU resource intelligent dynamic optimization method and system based on BIOS and BMC

The invention relates to the technical field of computers, and belongs to a BIOS (Basic Input / Output System) and BMC (Baseboard Management Controller)-based GPU (Graphics Processing Unit) resource intelligent dynamic optimization method and system. A four-layer collaborative architecture of a firmware layer (BIOS), a management controller layer (BMC), a system layer (operating system kernel) and a hardware layer (GPU equipment) is adopted, and the architecture is a cross-layer closed-loop optimization architecture. An optimization decision function is decoupled from an operating system kernel and is deployed in an independent hardware management unit, namely a baseboard management controller (BMC). The system collects GPU operation data in real time through an operation system kernel space and sends the GPU operation data to a BMC; an intelligent optimization engine deployed in the BMC analyzes and calculates the data to generate an optimal GPU resource scheduling strategy; the strategy is returned to an operating system kernel and finally executed by a hardware layer, and intelligent dynamic optimization of GPU resources is achieved through cross-level interaction. Therefore, a continuous intelligent closed-loop optimization link of'deployment-data acquisition-intelligent decision-making-strategy execution 'is formed.
Owner:HANGZHOU JINQUN TECHNOLOGY CO LTD

GPU shader rendering computer implementation method of finite element result

The invention discloses a GPU shader rendering computer implementation method of a finite element result. The method comprises the steps that a static three-dimensional grid is adopted as a rendering geometry; the 32-bit floating point type scalar data corresponding to the vertexes are subjected to lossless coding at a CPU end to obtain four-channel 8-bit RGBA vertex color attributes; during data updating, lightweight vertex color data are only transmitted to the GPU; hardware interpolation is carried out on the encoded color attributes by using a GPU rasterizer; and finally, in the fragment shader, decoding the interpolation result of each pixel to reconstruct a scalar value, and mapping the scalar value into a final color. According to the method, the bottleneck of topological calculation of a CPU end and transmission of massive geometric data to the GPU is avoided, the calculation load is transferred to the GPU in a large scale for parallel processing, the rendering efficiency is remarkably improved, and high-frame-rate dynamic visualization of a finite element result is realized.
Owner:CHANGJIANG SPATIAL INFORMATION TECH ENG CO LTD (WUHAN) +1

Large model reasoning method and device, medium and electronic equipment

The invention provides a large model reasoning method and device, a medium and electronic equipment. In the method, a central processing unit can divide a batch processing data set of a to-be-executed reasoning task into sub-batch processing data sets and distribute the sub-batch processing data sets to a plurality of graphics processors, so that the graphics processors execute reasoning calculation based on the distributed sub-batch processing data sets respectively, and an execution result of the to-be-executed reasoning task is obtained. In the process, each graphic processor can preload the weight parameter of the next reasoning calculation layer from the storage space corresponding to the central processor and store the weight parameter into a video memory in the process of executing the reasoning calculation task corresponding to any reasoning calculation layer; therefore, when the calculation task of the current reasoning calculation layer is not completed, the graphics processing unit can load the weight parameters required by the subsequent reasoning calculation layer into the video memory in advance, so that storage is replaced by passage, and the bandwidth utilization rate of the video memory and the overall reasoning throughput are improved.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Data transmission method and device, computer, storage medium and program product

The embodiment of the invention discloses a data transmission method and device, a computer, a storage medium and a program product. The method comprises the steps that a data sending request is responded, and first service data, a first graphic processing unit for sending the first service data and a second graphic processing unit for receiving the first service data are acquired from the data sending request; the first graphic processing unit is one graphic processing unit in the first service equipment; determining a first queue pair used for sending the first service data from S queue pairs associated with the first graphic processing unit; s is a positive integer; and obtaining a first hardware port corresponding to the first queue pair, and sending the first business data to a second graphic processing unit in the second service equipment through the first hardware port and the first queue pair. By adopting the application, the data transmission efficiency can be improved, and the hardware utilization rate and communication performance of the service equipment can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method, device, apparatus and readable storage medium for selecting anti-aliasing algorithm

The embodiments of the present application disclose a method, apparatus, device and readable storage medium for selecting an anti-aliasing algorithm, which is used to dynamically select an anti-aliasing algorithm suitable for the current image scene to improve the anti-aliasing effect and image quality. The method of the embodiment of the present application includes: a CPU obtains at least one instruction for calling an application program interface (API), the at least one instruction carries rendering information of each of M models, and the M models belong to the same frame image, and M is a positive integer; then, based on the rendering information of each of the M models, an anti-aliasing algorithm is selected from a plurality of anti-aliasing algorithms as a target anti-aliasing algorithm; the CPU sends instruction information to a graphics processing unit (GPU), so that the GPU renders at least one frame of image based on the target anti-aliasing algorithm.
Owner:HUAWEI TECH CO LTD

Method and device for managing server GPU (Graphics Processing Unit) equipment

The invention relates to a management method and device for server GPU (Graphic Processing Unit) equipment, and the method comprises the following steps: after a server is started, a basic input / output system BIOS writes GPU static asset information through an H2B PCIE (Peripheral Component Interface Express) shared memory; the baseboard management controller BMC obtains GPU static asset information and extended static asset information of the GPU devices through the H2B PCIE shared memory, and polls all the GPU devices; and if the remote client initiates an information acquisition request of the GPU equipment, the northbound interface of the server acquires and feeds back the GPU equipment information through the callback function. According to the method, the basic input and output system BIOS directly and preferentially transmits the GPU equipment information to the shared memory of the GPU equipment information area predefined by the BMC in the starting stage, and the time from starting of the basic input and output system BIOS to obtaining of the basic information of the GPU equipment by the BMC is shortened.
Owner:POWERLEADER COMPUTER SYST CO LTD