Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

4500results about "Processor architectures/configuration" patented technology

Method for automatically drawing OpenGL program by using Vulkan

The invention discloses a method for automatically drawing an OpenGL (Open Graphics Library) program by using Vulkan. The method comprises the following steps of: creating a context used by the Vulkan, initializing each module, processing an OpenGL instruction related to texture and data buffering, and managing storage of texture and data buffering resources in a video memory; a shader program used by the OpenGL is preprocessed into a format acceptable to Vulkan, and an OpenGL shader program instruction is created and destroyed; processing an OpenGL (Open Graphics Library) instruction related to frame buffering to generate structural body information required by Vulkan dynamic rendering; an OpenGL instruction of the sampler is also created; processing an OpenGL (Open Graphics Library) instruction for creating a vertex input format and managing a vertex data buffer area, and maintaining vertex input information, a vertex buffer area and an index buffer area required by Vulkan; and finally, drawing or calculating, distributing and calling Vulkan on the basis of all the instructions.
Owner:ZHEJIANG UNIV +1

Digital twin three-dimensional scene modeling method based on webGPU

The invention discloses a digital twin three-dimensional scene modeling method based on a webGPU, and relates to the technical field of three-dimensional scene modeling and graphic computing, multi-source sensing data are asynchronously sampled from sparse point cloud, a video texture sequence, a scene semantic tag graph and a structure boundary tuple, and a six-dimensional structure unit group is generated through a normalization operator; constructing a structural unit atlas with nodes representing component entities and edges representing constraint relations, and introducing a tension balance mechanism and multi-scale constraints to generate a modeling path prior model; constructing a graph calculation and graph rendering dual-channel assembly line in the WebGPU, and executing parallel texture mapping and boundary fitting operation; a dynamic sensing module is used for capturing scene disturbance and driving atlas response, and incremental reconstruction of the model is achieved; and finally, mapping the model to a Web terminal, and supporting microscopic semantic query and multi-layer data linkage. According to the method, the response speed, semantic consistency and structural adaptability of three-dimensional modeling in a complex environment are improved.
Owner:ZHEJIANG ZHEFENG YUNZHI TECH CO LTD

Direct3D memory model compatible method based on adaptive occupied resources

The invention discloses a Direct3D memory model compatible method based on adaptive placeholder resources, which comprises the following steps: establishing resource metadata, accessing a scene library, checking an instruction template and double sandboxes when a DXVK is started, describing related resources by the DXVK through the metadata after a D3D application is started, completing resource mapping and metadata dynamic updating, and allocating the resources to the corresponding sandboxes; the DXVK compiles an application shader code, identifies a resource access instruction, matches an access scene and a check template by combining a pipeline stage and binding slot query metadata, instantiates the check instruction and adds the check instruction to the front of a target instruction; according to the unbound resources, adaptive generation of corresponding occupied resources is carried out, the authority and the state are configured, the unbound resources are bound to a Vulkan descriptor set to replace VKNULLHANDLE, and adaptation of the shader interface and the pipeline state parameters is completed to create PSO (Particle Swarm Optimization); and intercepting an application rendering instruction analysis parameter to construct a command buffer area, submitting the command buffer area to a Vulkan queue to trigger a GPU (Graphic Processing Unit) to execute rendering, and finally completing complete conversion and adaptation from the D3D rendering logic to the Vulkan.
Owner:北京麟卓信息科技有限公司

Lossless and lossy automatic hardware compression in graphics-to-graphics network links

An apparatus to facilitate lossless and lossy automatic hardware compression in graphics-to-graphics network links is disclosed. The apparatus includes compressor / decompressor circuitry (CDC) integrated with physical layer (PHY) intellectual property (IP) hardware circuitry for a graphics processor unit (GPU)-to-GPU communication link communicably coupling a first GPU to one or more other GPUs, the CDC to: receive a data message from the first GPU, wherein the data message is in an uncompressed format; determine that a compression process is to be applied to the data message; apply the compression process to the data message to generate a compressed data message; and cause a GPU link IP hardware circuitry that comprises the PHY IP hardware circuitry to transmit the compressed data message over the GPU-to-GPU communication link.
Owner:INTEL CORP

Tensor core matrix multiplication and accumulation with hardware-based statistics collection and outlier suppression

An apparatus providing tensor core matrix multiplication and accumulation (MMA) with hardware-based statistics collection and outlier suppression is disclosed. The apparatus includes processor circuitry comprising at least one processor core comprising matrix multiplication circuitry to: execute a matrix multiplication operation on first input data from a first set of registers and on second input data from a second set of registers; collect, as part of executing the matrix multiplication operation via statistics collection hardware circuitry of the matrix multiplication circuitry, output statistics data corresponding to the matrix multiplication operation; and output the output statistics data along with a result of the matrix multiplication operation; and output statistics storage to store the output statistics data.
Owner:INTEL CORP

Intelligent dynamic management method for GPU (Graphics Processing Unit) computing power and cloud platform

The invention is suitable for the field of GPU management, and provides an intelligent dynamic management method for GPU computing power and a cloud platform, and the method comprises the following steps: collecting task data and GPU state data, and generating a task queue and a resource allocation strategy in combination with a scheduling plug-in; based on the task queue and the GPU real-time load, dynamically adjusting a resource allocation proportion through reinforcement learning to analyze a task dependency relationship and generate a migration plan; optimizing a communication path and adjusting asynchronous transmission delay according to the task dependency relationship and the GPU communication topology, and outputting a synchronous state mark; and monitoring abnormity in combination with the synchronization state and the GPU hardware state, executing thermal migration according to the migration plan, performing video memory recovery, and updating the resource idle list. According to the invention, through algorithm innovation and hardware collaborative optimization, intelligent, dynamic and efficient resource scheduling in the multi-GPU system is realized.
Owner:ZHEJIANG XIANGONG CLOUD TECH CO LTD

Data processor, data processing method, electronic device and storage medium

Provided in the present disclosure are a data processor, a data processing method, an electronic device and a non-transitory computer-readable storage medium. The data processor comprises a tensor operation unit and N compute units, wherein the tensor operation unit is configured to execute tensor computation on input data, so as to obtain tensor computation results; and the N compute units are configured to execute at least one of a vector operation of the tensor computation results and the generation of input data, wherein first data transmission channels are provided between the tensor operation unit and at least some compute units among the N compute units, and the first data transmission channels are used for directly providing the tensor computation results to the compute units, and directly providing the input data from the compute units to the tensor operation unit. By means of the data processor, access to a global memory can be reduced, the waste of resources is reduced, the data transmission delay is shortened, the overall efficiency of operators is significantly improved, the strong computing power of a tensor operation unit itself is effectively used, and the computation efficiency of the data processor is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Game scene optimization method and system based on dynamic rendering

The invention provides a game scene optimization method and system based on dynamic rendering, and belongs to the technical field of game development.The game scene optimization method comprises the steps that firstly, a real-time rendering state data set of a current game scene, rendering resource occupation information containing scene elements and picture frame rendering time consumption parameters are obtained; then, a pre-trained rendering demand analysis model is called for analysis, a rendering demand feature set of the scene elements is obtained, the rendering demand feature set comprises element visual importance features and rendering resource sensitivity features, and a scene element dynamic optimization strategy is generated based on the rendering demand feature set; the scene element dynamic optimization strategy comprises an element detail level adjustment rule and a rendering resource allocation priority parameter, executing a rendering parameter adjustment operation on a target scene element according to the scene element dynamic optimization strategy, generating adjusted scene rendering configuration data, and finally inputting the scene rendering configuration data into a game rendering pipeline for real-time rendering processing. And outputting the optimized game scene picture frame sequence. Therefore, the game picture quality and the operation performance can be improved.
Owner:CHENGDU FUQIAN TECH CO LTD +1

Tensor Memory Accelerator Enhancements

One embodiment provides a graphics processor comprising a memory interface and a graphics core cluster including a plurality of graphics cores and tensor processing circuitry. The tensor processing circuitry includes a local memory, a tensor accelerator coupled with the local memory, the tensor accelerator configured to perform a matrix multiply and accumulate operation, and a tensor data movement accelerator configured to asynchronously transfer tensor data between a global memory coupled to the memory interface and the local memory. The tensor data movement accelerator includes circuitry configured to translate the tensor data from a first tensor format to a second tensor format.
Owner:INTEL CORP

Automatic lecturer video generation method based on AI speech synthesis and animation driving

The invention discloses a lecturer video automatic generation method based on AI speech synthesis and animation driving. The method comprises the following steps: performing structured analysis on a PPT or a text script through an improved interior point method and an incremental shortest path algorithm; performing semantic grouping by applying a full-dynamic parallel single-link clustering algorithm and generating an enhanced script with an expressive mark; a CosyVoice technology is combined with a low-rank approximation method to generate a high-quality voice data stream; establishing a mapping relation between contents and action expressions through semantic analysis, and generating a complete action expression instruction set; and driving the digital human model by using the msueTalk technology, and generating a final lecturer teaching video through a parallel rendering algorithm. According to the invention, the method achieves the efficient and automatic generation of the education video, remarkably improves the content production efficiency, reduces the production cost, and guarantees the specialty and expressive force of the teaching video.
Owner:SHENZHEN XUEYOU TECHNOLOGY CO LTD

Three-dimensional scene optimization method and system for collaborative rendering of dynamic LOD and view cone elimination based on space-time prediction

The invention relates to the field of scene rendering, and particularly discloses a three-dimensional scene optimization method and system for collaborative rendering of dynamic LOD and view cone rejection based on space-time prediction, and the method comprises the steps: dynamically calculating the visibility frequency of an object, and dynamically adjusting the loading and rendering modes of objects with different priorities; predicting a view cone range of multiple frames in the future by using an LSTM network architecture; dividing the scene into uniform grid blocks, and then performing grading elimination; optimizing a heterogeneous computing pipeline; rendering the execution process; the system comprises a dynamic LOD and visual cone rejection collaborative optimization module, an LSTM visual cone prediction module, a block-level mixed rejection module, a heterogeneous calculation pipeline module, an optimization CPU-GPU task allocation and data transmission module and a rendering execution control module. According to the method, the technologies of collaborative optimization of dynamic LOD and view cone removal, block-level mixed removal, heterogeneous calculation assembly line optimization and the like are adopted, so that unnecessary rendering calculation and resource loading are reduced, and the rendering efficiency is improved.
Owner:YANTAI JIERUI NETWORK TRADING

Wafer defect detection method and related equipment

The invention discloses a wafer defect detection method and related equipment, and relates to the technical field of data processing, and the method comprises the steps: receiving a preset first resolution image of a corresponding wafer collected by a camera in response to a wafer defect detection instruction, and carrying out the parallel processing of the preset first resolution image based on a CPU-GPU cooperative calculation strategy, determining a binary mask corresponding to the preset first resolution image according to the preset first resolution image, and judging whether the wafer has defects or not based on a preset wafer defect judgment rule and the binary mask so as to obtain a wafer defect detection result. According to the method and the device, the preset first resolution image is processed in parallel through a CPU-GPU cooperative computing strategy, so that the image processing speed is improved, and the real-time performance of wafer defect detection is improved.
Owner:ZHONGKE SHANHAIWEI (HANGZHOU) SEMICONDUCTOR TECHNOLOGY CO LTD

Virtual actor based on 4D Gaussian splashing and XR and on-site immersive real-time presentation system and method thereof

The invention belongs to the technical field of augmented reality (XR) and computer vision crossing, relates to fusion application in immersive digital performance, and provides a virtual actor reconstruction and immersive presentation system based on 4D Gaussian splash modeling and XR space positioning. The system comprises a set of spherical multi-camera-position high-synchronization camera shooting matrix used for capturing dynamic images of actors; performing dynamic modeling on the multi-angle image through a 4D Gaussian splashing technology, and outputting a virtual actor point cloud model which can be used by XR equipment; vPS visual positioning and an SLAM tracking module are combined, and precise mapping positioning of a performance space is achieved in AR / MR equipment. The method supports the immersive watching of the actor image at the audience end at a 360-degree free visual angle, and realizes the natural presentation of the virtual actor without dead angles and wearing in cooperation with shielding judgment and a real-time rendering engine. The system is widely applicable to on-site entertainment scenes such as immersive theaters, text travel performances, brand activities, concerts and television programs.
Owner:SHANGHAI SHICHEN CULTURAL COMMUNICATION CO LTD

Apparatus and method for block-friendly ray traversal

Apparatus and method for efficient storage of BVH nodes in blocks. For example, one embodiment of an apparatus comprises: bounding volume hierarchy (BVH) construction circuitry to construct a BVH based on primitives of a graphics scene; and block allocation hardware logic coupled to or integral to the BVH construction circuitry, the block allocation hardware logic to allocate a plurality of nodes of the BVH into a plurality of blocks for storage in a cache or memory subsystem, the block allocation hardware logic to maximize a number of blocks which include a leading parent node and one or more corresponding child nodes of the plurality of nodes.
Owner:INTEL CORP

Digital twin platform GPU rendering resource dynamic scheduling method and system

The invention relates to a digital twin platform GPU rendering resource dynamic scheduling method and system, and the method comprises the steps: S1, collecting a rendering task feature signal of a current frame in real time, and collecting a GPU multi-dimensional resource consumption signal; s2, based on the rendering task feature signal, the historical resource consumption signal and the scene dynamic change signal, generating a GPU resource demand prediction signal and a load fluctuation trend signal of a future frame; s3, predicting a signal, a task dependency relationship signal, a user priority signal and a real-time resource bottleneck type signal according to the GPU resource demand; s4, executing the dynamic resource allocation instruction signal; and S5, generating a parameter self-optimization signal for updating the prediction and scheduling logic of a subsequent frame according to the actual resource consumption signal of the current frame and the scheduling effect evaluation signal. According to the dynamic scheduling method and system for the GPU rendering resources of the digital twin platform, the problem that the GPU resource utilization rate is low and the real-time performance is difficult to consider at the same time under the dynamic load can be solved.
Owner:ZHONGKE HUIZHI (BEIJING) TECH CO LTD

Direct3D rendering model compatible method based on dynamic template pool

The invention discloses a Direct3D rendering model compatible method based on a dynamic template pool, which comprises the following steps: establishing three types of mapping tables between D3D and Vulkan when compiling DXVK, constructing resource metadata, and creating a core, extended and temporary three-level template pool according to the mapping tables after starting; when the D3D application creates resources, metadata is initialized, parameters are verified, physical memories are allocated and grouped, resource groups are pre-verified, a batch binding command is generated, memory binding is completed, and the metadata is updated; when a resource view is created, a template is matched from a template pool, a handle is generated after instantiation, a resource handle is associated, and a descriptor is bound; when a rendering state is set and a rendering instruction is executed, the rendering state and the rendering instruction are respectively converted into a Vulkan related state and a Vulkan related instruction through a mapping table, a handle is bound after a PSO cache is inquired, a command buffer area is submitted to a GPU queue to execute drawing, and compatible operation of the D3D application on a platform supporting a Vulkan operating system is realized under the condition that the GPU does not support VKKHRmaintence5 and VKKHRmaintence6 extension.
Owner:北京麟卓信息科技有限公司

Model performance estimation method and device and computer equipment

The invention relates to the technical field of artificial intelligence chips, and discloses a model performance estimation method and device and computer equipment, and the method comprises the steps: determining model configuration data and candidate distributed strategies of a target model; converting the original calculation graph corresponding to the single-card deployment state based on the model configuration data and the strategy configuration data of the candidate distributed strategies to obtain a distributed overhead calculation graph corresponding to the multi-card deployment state; and performing performance estimation on the candidate distributed strategy according to the basic overhead and the additional overhead in the distributed overhead calculation graph to obtain strategy performance data of the target model under the candidate distributed strategy, thereby realizing conversion of a single-card model into a multi-card model. And based on the multi-card model, simulation calculation of the multi-card interconnection mode is realized on the premise of limited hardware resources, so that the influence of a communication operator and a topological structure corresponding to the multi-card interconnection mode in the overall operation of the model is reflected, and the upper limit of the model performance can be accurately evaluated in a simulator verification stage before silicon is applied.
Owner:SHANGHAI BIREN TECH CO LTD

High-resolution machine vision detection method and system for precise chromatic aberration detection and medium

The invention provides a high-resolution machine vision detection method and system for precise chromatic aberration detection and a medium, and the method comprises the steps: carrying out the calibration of a camera based on a camera calibration tool, and synchronously calibrating a light source parameter and a color parameter; setting shooting parameters based on the calibrated camera, obtaining a workpiece image in real time, preprocessing the workpiece image, mapping the preprocessed image to a standard color space from an original RGB space, and extracting color features of the preprocessed image; comparing the color feature with the color feature of the standard sample, calculating a color difference index, and comparing with the color difference index based on a set color difference threshold to obtain a detection result; by calibrating various parameters of the camera, the visual detection precision is ensured, and by performing color space mapping on the workpiece image and analyzing the difference between the color difference index of the color feature and the color difference threshold value, the color abnormal area is accurately analyzed, and the detection precision of the abnormal color is improved.
Owner:HANGZHOU HUICUI INTELLIGENT TECH CO LTD

Heterogeneous GPU resource management scheduling method, computer device, medium and product

The invention discloses a heterogeneous GPU resource management scheduling method, a computer device, a medium and a product. The method comprises the following steps: acquiring computing power resources of each node, including a GPU model, a GPU video memory, a GPU number, a computing power segmentation scheme and a GPU use condition; automatically segmenting the heterogeneous GPU of each node according to the computing power segmentation scheme of each node to obtain resource segmentation information; obtaining an expected computing power and an expected video memory of a to-be-executed task; screening out target nodes meeting the expected computing power and the expected video memory according to a node analysis strategy; and according to the sub-resource analysis strategy, screening out sub-resources meeting the expected computing power and the expected video memory, recording the sub-resources as target sub-resources, and allocating the to-be-executed task to the target sub-resources. According to the method, GPU resource fragmentation is effectively reduced through global and node resource conjoint analysis scheduling, meanwhile, pooling management, dynamic configuration and intelligent scheduling of heterogeneous computing equipment can be achieved, and the resource utilization rate is effectively increased.
Owner:BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD

Memory writing method and device, storage medium and program product

The invention discloses a memory writing method and device, a storage medium and a program product, and relates to the technical field of memory access, and the method comprises the following steps: determining a memory step length of an initial image tensor corresponding to a target image, determining a target vectorization width according to the memory step length and a memory data width of a target graphics processor, and determining a target data dimension of the target graphics processor, determining a target thread grid corresponding to the initial image tensor according to the target data dimension, performing image tensor slicing by using the target thread grid to obtain a target image tensor, and writing the target image tensor into a memory. According to the method, the memory layout (step length) of the input tensor can be dynamically analyzed, so that the optimal vectorization width is adaptively determined, the memory access efficiency is maximized, slicing is performed according to the data dimension when the target graphics processor outputs the data, the method has adaptability and high efficiency, the GPU video memory bandwidth is fully utilized, and the memory bandwidth utilization rate is improved.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Automated hardware-aware deployment of machine learning pipelines on chipsets

A method or system for implementing a machine learning pipeline on a chipset comprising a plurality of hardware compute elements. The system accesses a hardware-agnostic functional description of the machine learning pipeline, wherein the description specifies a plurality of functional modules, including at least one machine learning model. Hardware specifications of the chipset are accessed to identify the available hardware compute elements. Based on the hardware specifications, the functional modules are synthesized into a plurality of interconnected executable components configured to execute on at least two different hardware compute elements. An implementation package is generated, comprising the executable components and metadata describing interconnections between them. The implementation package is then deployed to the chipset, where the executable components are executed by the identified hardware compute elements.
Owner:SIMA TECHNOLOGIES INC

Layout design rule checking method, system, equipment, medium and product

The invention relates to an application method of a GPU (Graphics Processing Unit) in design rule inspection in the field of integrated circuits, in particular to a layout design rule inspection method, system and equipment, a medium and a product. The layout design rule checking method comprises the following steps: inputting initial layout data, analyzing a design rule checking task to be executed, and marking the priority and complexity level of the task; the method comprises the following steps: dividing layout data into a plurality of data blocks according to regions on the basis of regional characteristics of the layout data, distributing tasks and corresponding data blocks to a GPU computing core according to task priorities and complexity levels, and distributing new data blocks according to the overall load condition of the GPU computing core after the GPU computing core completes data block processing, and repeating the steps until all tasks are executed, and outputting a layout design rule check result. The system, the computer equipment, the computer readable storage medium and the computer program product have the same beneficial effects as the layout design rule checking method.
Owner:HUAXIN GIANTS (HANGZHOU) MICROELECTRONICS CO LTD

Infrared image enhancement method and system based on local phase correlation

The invention relates to the technical field of image processing, and discloses an infrared image enhancement method and system based on local phase correlation, and the method comprises the steps: obtaining a plurality of continuous frames of infrared images, carrying out the intelligent partitioning of a reference frame, and calculating the variance feature, the method comprises the following steps: selecting regions of interest with rich information, independently executing phase correlation operation in each region to extract a local translation vector, obtaining global displacement estimation through weighted fusion, adopting an abnormal value detection algorithm to improve robustness, and finally realizing sub-pixel-level image alignment and intelligent weighted fusion. The method is suitable for real-time enhancement processing of satellite-borne infrared remote sensing images, the resource constraint requirement of an embedded platform is met while the processing quality is guaranteed, and an efficient and reliable technical scheme is provided for space remote sensing image processing.
Owner:SHANGHAI WEIXING DATA TECH CO LTD

Hierarchy of neural network scaling factors

Embodiments described herein provide techniques to facilitate hierarchical scaling when quantizing neural network data to a reduced-bit representation. The techniques includes operations to load a hierarchical scaling map for a tensor associated with a neural network, partition the tensor into a plurality of regions that respectively include one or more subregions based on the hierarchical scaling map, hierarchically scale numerical values of the tensor based on a first scale factor and second scale factor via the matrix accelerator circuitry, the first scale factor based on a statistical measure of a subregion of numerical values of within a region of the plurality of regions and the second scale factor based on a statistical measure of the region that includes the subregion, and generate a quantized representation of the tensor via quantization of hierarchically scaled numerical values.
Owner:INTEL CORP

Programmatic Work Assignment For Dynamically Load-Balanced Persistent Execution

In a GPU design, “launching a worker” is de-coupled from “assigning a work item” in a work distributor, and new handshake mechanisms between a worker and the work-distributor is provided for work assignment, in order to provide persistent kernel functionality. In example embodiments, software specifies the work that has to be done, hardware selects a variable number of workers based on available resources, and a hardware scheduler handshaking with the executing workers assigns more work as previously assigned work is completed and / or more resources become available.
Owner:NVIDIA CORP

Artificial intelligence chip and operation method thereof

The invention provides an artificial intelligence chip and an operation method thereof. The artificial intelligence chip comprises a tensor core, a tensor transmitting unit, a general computing core, a general transmitting unit and a thread block segmentation unit. In response to operation of the thread block segmentation unit in the first operation mode, the thread block segmentation unit distributes a thread bundle to the general-purpose transmission unit, and the general-purpose transmission unit transmits an instruction to the tensor core and the general-purpose computing core. In response to operation of the thread block segmentation unit in the second operation mode, the thread block segmentation unit distributes a tensor thread bundle involving tensor calculation to the tensor transmitting unit, the tensor transmitting unit transmits a tensor calculation instruction to the tensor core, the thread block segmentation unit distributes a non-tensor thread bundle not involving tensor calculation to the general transmitting unit, and the general transmitting unit transmits a non-tensor thread bundle involving tensor calculation to the general transmitting unit. And the general transmitting unit transmits the non-tensor calculation instruction to the general calculation core.
Owner:SHANGHAI BIREN TECH CO LTD

Mass data particle system rendering method and device based on WebGL

The invention discloses a WebGL-based mass data particle system rendering method and device, and belongs to the technical field of computer graphic processing. The method comprises the following steps: firstly, selecting the type of a particle emitter according to scene requirements, then constructing a primitive object based on points or quadrangles, and configuring three-dimensional coordinates and texture parameters; aggregation rendering is carried out on a single-frame primitive array through a WebGL vertex shader and a fragment shader, and the drawing calling frequency is reduced; the method comprises the steps of dynamically adjusting particle attributes, including position compensation, color gradient and explicit-implicit control, determining a rendering form and playing time of an animation according to requirements of an actual application scene, and supporting fine configuration of a single particle, including texture animation, dynamic scaling and rotation parameters. When a large amount of data is processed, on the premise of ensuring the rendering quality, the rendering efficiency can be remarkably improved, the consumption of a memory and GPU resources can be reduced, the response time of the system can be optimized, and meanwhile, the interaction experience of a user can be improved.
Owner:RES INST OF CHEM DEFENSE PLA ACAD OF MILITARY SCI

Combined MX and sparsity representation

One embodiment provides a graphics processor comprising a memory interface and a processing cluster array including a plurality of processing resources interconnected via a switched interconnect network, at least one of the plurality of processing resources including a matrix accelerator configured to execute an instruction to perform a multi-dimensional sparse matrix multiply and accumulate operation having input in a sparse microscaling format including merged sparsity and scaling metadata.
Owner:INTEL CORP