Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

736 results about "Video memory" patented technology

Method for automatically drawing OpenGL program by using Vulkan

The invention discloses a method for automatically drawing an OpenGL (Open Graphics Library) program by using Vulkan. The method comprises the following steps of: creating a context used by the Vulkan, initializing each module, processing an OpenGL instruction related to texture and data buffering, and managing storage of texture and data buffering resources in a video memory; a shader program used by the OpenGL is preprocessed into a format acceptable to Vulkan, and an OpenGL shader program instruction is created and destroyed; processing an OpenGL (Open Graphics Library) instruction related to frame buffering to generate structural body information required by Vulkan dynamic rendering; an OpenGL instruction of the sampler is also created; processing an OpenGL (Open Graphics Library) instruction for creating a vertex input format and managing a vertex data buffer area, and maintaining vertex input information, a vertex buffer area and an index buffer area required by Vulkan; and finally, drawing or calculating, distributing and calling Vulkan on the basis of all the instructions.
Owner:ZHEJIANG UNIV +1

Multi-modal large model dynamic compression and reasoning optimization method based on MoE architecture

The invention relates to a multi-modal large model dynamic compression and reasoning optimization method based on a MoE architecture. The method comprises the following steps: establishing an edge computing system conforming to medical equipment specifications, constructing a medical image analysis network based on an improved hybrid expert MoE architecture, and adopting a three-layer cascade structure of a feature coding layer, a dynamic routing layer and an expert execution layer; executing expert module dynamic loading and video memory optimization; executing knowledge graph compensation and domain knowledge injection; executing hardware instruction level optimization and calculation acceleration; executing multi-expert feature fusion and decision weighting; performing diagnosis result generation and confidence evaluation; performing real-time data return and model iterative optimization; executing multi-device cooperation and load balancing; executing system security monitoring and exception handling; and generating a structured diagnostic report. The problem that the precision loss of a multi-modal large model is difficult to meet actual requirements is solved, and medical feature adaptive dynamic compression, medical hardware collaborative energy efficiency optimization and cross-modal compensation of medical knowledge enhancement are realized.
Owner:SUZHOU WUDING NETWORK TECHNOLOGY CO LTD

GPU computing power resource scheduling method and system

The invention relates to the technical field of data analysis, and discloses a GPU computing power resource scheduling method and system, and the method comprises the steps: collecting node hardware parameters and dynamic load indexes of a GPU cluster to construct a multi-dimensional resource feature vector of the GPU cluster, and constructing a resource portrait of the GPU cluster; establishing a node health degree scoring model of the GPU cluster, and generating a health degree score of a cluster node corresponding to the GPU cluster; analyzing a video memory demand of the GPU task request, and calculating an intensive identifier and a communication dependency relationship; determining the SLA weight of the GPU task request, calculating the resource shortage sensitivity of the GPU task request based on the video memory demand, and calculating the target task priority of the GPU task request in combination with the SLA weight; and determining a resource scheduling node group requested by the GPU task in the resource portrait, generating resource scheduling parameters of the resource scheduling node group, and executing scheduling of computing power resources of the GPU cluster based on the resource scheduling parameters. According to the method, the scheduling efficiency of the GPU computing power resources can be improved.
Owner:SHENZHEN DIXI YUNLIAN TECH CO LTD

Game animation character display method based on virtual reality technology

The invention provides a game cartoon character display method based on a virtual reality technology. The method comprises the following steps: acquiring three-dimensional model data of a game cartoon character in a target display area through a user interaction terminal; the virtual reality content management platform determines a dynamic rendering precision level according to the model data, and generates a real-time rendering strategy including a model patch reduction coefficient, a texture compression rate and a skeleton animation updating frequency in combination with terminal performance parameters; the strategy is sent to a virtual reality supervision platform and a user interaction terminal, and a virtual reality rendering engine platform is instructed to execute real-time rendering; and the virtual reality supervision platform monitors the frame rate fluctuation data of the head-mounted display device, calculates the scene rendering stability, sends a rendering optimization instruction if the scene rendering stability is lower than a threshold value, and adjusts the model data acquisition frequency and the video memory cleaning period to optimize the performance. The rendering efficiency and the system stability can be improved, and the hardware load and the frame rate fluctuation are reduced.
Owner:JIANGSU JIUQU INTERACTIVE ENTERTAINMENT NETWORK TECHNOLOGY CO LTD

Deep learning training and reasoning task dynamic cooperation system based on GPU space-time resource sharing

The invention provides a deep learning training and reasoning task dynamic cooperation system based on GPU space-time resource sharing, and the system comprises a GPU resource state perceptron which is used for monitoring a GPU kernel function call sequence and a video memory distribution state of a distributed training task in real time, dynamically capturing a calculation gap and video memory fragments generated by the training task, and transmitting the calculation gap and the video memory fragments to the GPU resource state perceptron; generating a two-dimensional resource spatial-temporal characteristic spectrum, and predicting a GPU calculation idle period caused by communication synchronization based on an LSTM model; the kernel function dynamic scheduling decision maker is used for performing priority division and dynamic resource quota allocation on an online reasoning task and an offline reasoning task by adopting a self-adaptive allocation strategy on the basis of a resource spatial-temporal characteristic spectrum so as to realize spatial-temporal resource decoupling of the training task and the reasoning task; and the kernel function execution arbiter is used for dynamically controlling submission and blockage of the reasoning task kernel function according to the scheduling decision through video memory space multiplexing and a calculation instruction arbitration mechanism. According to the method, the GPU resource utilization rate is remarkably improved, on the premise that stable training task performance is guaranteed, fragmented resources are effectively utilized to support parallel execution of multiple types of reasoning tasks, and efficient resource collaboration of a deep learning task cluster is achieved.
Owner:NANJING INFORMATION HIGH-SPEED RAILWAY RES INST OF SCI AND TECH

End-side multi-mode large model accelerated reasoning method and system

The invention provides an end-side multi-modal large model accelerated reasoning method and system, and the method comprises the steps: carrying out the two-stage screening and rearrangement of visual tokens based on the CLS attention and text-to-visual attention in a visual encoder and pre-filling stage, and constructing a sparse attention and sparse key value cache; in a decoding stage, an important neuron set is judged according to activation gating or historical statistics, only a corresponding feedforward network weight is pulled and calculated, missed weights are loaded on demand through asynchronous I / O, and hot neurons are maintained in a high-speed memory to utilize model sparsity, so that video memory / memory occupancy and calculation overhead are remarkably reduced on an end side; throughput and time delay performance are improved. According to the method, the internal memory and computing resources required by reasoning of the multi-modal large language model are reduced from two dimensions by utilizing the endogenous sparsity of the end-side large language model in input and the model, so that a higher reasoning speed is achieved by utilizing fewer resources on the premise of keeping the size of the model unchanged, and the performance of the whole system is improved.
Owner:SHANGHAI JIAOTONG UNIV

Cache management method and device, storage medium and electronic equipment

The invention provides a cache management method, a cache management device, a computer storage medium and electronic equipment, and relates to the technical field of computers. The method comprises the steps of receiving a reasoning task request and distributing the reasoning task request to a target storage page; key value cache information of the first round of reasoning task is stored in a hard disk cache, and when the second round of reasoning task is executed, key value cache information generated before the second round of reasoning task is preloaded layer by layer from the hard disk cache; when the last round of reasoning task is received, storing first target key value cache information correspondingly generated by the last round of reasoning task into the matched target physical block; and performing hybrid grouping compression on key cache information and value cache information in the first target key value cache information to obtain second target key value cache information after quantization compression. According to the invention, triple balance of video memory-calculation performance-precision can be realized.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Industrial personal computer and multi-graphics card collaborative parallel operation acceleration system

The invention discloses an industrial personal computer and multi-graphics card collaborative parallel computation acceleration system, which relates to the technical field of industrial resource allocation and parallel computation, and comprises a resource monitoring and predicting module, a resource management module and a prediction type resource preparation module, the task splitting and collaborative execution module comprises a task splitting module, a cross-node collaborative module and a collaborative operation engine; the intelligent scheduling and dynamic resource allocation module comprises an intelligent scheduler, a dynamic resource allocation module and a conflict avoidance module. According to the method, the GPU video memory utilization rate, the core utilization rate, the temperature, the video memory fragment rate, the available video memory total amount, the CPU core total utilization rate, the load condition and the idle core number index are collected in real time through a resource monitoring module, and the video memory capacity, the GPU core occupancy rate and the CPU load requirement are predicted in advance before a task is submitted in combination with a gradient boosting decision tree and a neural network prediction model; and resources are reserved, so that the scheduling delay is remarkably reduced, and the scheduling hit rate is improved.
Owner:ZHUHAI SHININGDA TECH CO LTD

Large language model progressive field fine tuning and knowledge fusion method oriented to shield engineering

The invention discloses a large language model progressive field fine tuning and knowledge fusion method for shield engineering. The method comprises the following steps: constructing a layered shield training course containing a wide-area academic theory and a proprietary enterprise construction method; parallelly training a plurality of physically isolated parameter efficient adapters based on the frozen base; performing singular value decomposition on the adapter, extracting a geometric feature subspace representing knowledge distribution, and calculating a conflict correlation degree; based on this, a uniform adaptation mechanism of resource awareness is constructed. The mechanism not only can generate a static fusion model for conflict removal, but also can dynamically activate a specific rank slice of an adapter through a routing network based on real-time hardware resource budget (video memory / FLOPs) and geometry-resource signature. According to the method, multi-source knowledge is reserved, and adaptive dynamic scheduling of edge hardware resources by model reasoning is realized.
Owner:CHINA RAILWAY 14TH BUREAU GRP LARGE SHIELD ENG CO LTD +1

Scene rendering method and system based on three-dimensional Gaussian splashing

The invention discloses a scene rendering method and system based on three-dimensional Gaussian splashing. The method comprises the following steps: organizing and constructing an original 3D Gaussian set to obtain a spatial hierarchical structure; generating corresponding level details for the primitives in the spatial hierarchical structure to obtain a spatial hierarchical structure associated with the level details; traversing the spatial hierarchical structure associated with level details, and executing hierarchical view cone cutting and shielding elimination to obtain a visible node list; traversing the visible node list, calculating a level detail selection standard, and generating a level detail activity Gaussian set; performing optimization sorting on the level detail activity Gaussian set to obtain an activity primitive list; and performing tile-based rasterization on the active primitive list, and performing adaptive processing according to level details of the primitives to obtain a final color value of each pixel. According to the method, the rendering performance can be greatly improved, the occupation of a memory and a video memory is remarkably reduced, the rendering quality is improved, visual flaws are reduced, and the expandability of the 3DGS rendering method is enhanced.
Owner:CHINA ORDNANCE SCI INST

Hybrid expert model reasoning optimization method based on speculative preloading

The invention provides a hybrid expert model reasoning optimization method based on speculative preloading, and belongs to the field of artificial intelligence safety. The method aims at improving the reasoning performance of a hybrid expert model (MoE for short) in a resource limited scene, and solves the technical problems of high video memory occupation and bottleneck in calculation and data transmission of the hybrid expert model in a reasoning stage. According to the method, through hot expert identification, speculative preloading and pipeline execution, a fine-grained expert parameter preloading strategy and an asynchronous loading mechanism based on a priority queue are designed, and effective overlapping of a calculation process and a data transmission process is realized, so that the reasoning efficiency is improved. The method comprises the steps that expert parameters are divided into hot experts and non-hot experts, the hot experts are loaded to a computing device memory in advance, parameters of the next layer are asynchronously preloaded according to prediction for activation of experts of the future layer, and meanwhile computing of the current layer is executed. According to the method, resource optimization reasoning is realized on hybrid expert models with different scales and structures on the premise of not changing the structure and precision of the model.
Owner:BEIHANG UNIV

Asynchronous parallel reasoning method, system and equipment for hybrid expert model and medium

The invention discloses an asynchronous parallel reasoning method, system and equipment for a hybrid expert model and a medium, which are corresponding schemes: decoupling synchronization of calculation and communication between GPUs (Graphics Processing Unit) caused by all-to-all set communication in expert parallelism, allowing asynchronous parallelism of model calculation and lexical metadata communication, and solving the problem of asynchronous parallelism of the model calculation and lexical metadata communication. Data communication overhead caused by expert parallelization is fully masked, and synchronization waiting overhead is eliminated; aiming at the phenomenon of uneven cold and heat of experts in reasoning, the hot experts are preferentially placed in the GPU, the cold experts are laterally loaded in the CPU so as to release the video memory space of the GPU, and the calculation efficiency of the GPU can be improved by increasing the batch size during reasoning; efficient resource scheduling is realized by dynamically selecting a computing unit which is most suitable for execution and a cold expert which needs to be loaded; generally speaking, the communication overhead and the waiting overhead during parallel reasoning of experts can be remarkably reduced, meanwhile, the calculation efficiency of the GPU is improved, and the overall throughput performance in the reasoning process is optimized.
Owner:UNIV OF SCI & TECH OF CHINA

Method for improving long text processing efficiency and accuracy

The invention discloses a method for improving long text processing efficiency and accuracy, and relates to the technical field of natural language processing and large language models.According to the method, text word segmentation embedding, sliding block preprocessing, YaRN position code injection, dynamic sparse attention calculation, multi-level attention fusion, graded KV cache management and output generation are sequentially executed; position drift is inhibited through logarithmic scaling, and key contexts are adaptively screened according to the attention activeness, so that the attention calculation complexity is close to linearity; in million-level Token reasoning, the video memory occupation of the method is reduced, the remote dependency recall rate is improved, and the method is suitable for scenes such as document analysis, code auditing and multi-mode streaming understanding.
Owner:BEI JING JING YUE KE JI YOU XIAN GONG SI

High-fidelity lightweight world model construction method for end-to-end automatic driving test

The invention relates to a world model construction method, in particular to a high-fidelity lightweight world model construction method for an end-to-end automatic driving test, which is used for constructing a high-fidelity world model and performing knowledge distillation on the world model to solve the problems of huge parameters, low reasoning efficiency and the like of the world model. Model parameters are reduced on the basis of reserving the world model generation capability, and the reasoning efficiency is improved; a CUDA operator is developed in a user-defined mode for the bottleneck part of world model calculation, video memory allocation is optimized, and the reasoning efficiency of a high-fidelity world model is improved based on a single-device multi-thread scheduling and multi-device cooperative calculation method. According to the method, the high-fidelity lightweight world model for the end-to-end automatic driving test can be constructed, the problems that an existing world model is low in multi-modal information alignment precision, poor in cross-view and cross-frame consistency, low in reasoning efficiency and the like are effectively solved, the confidence coefficient of the end-to-end automatic driving system test process is improved, and the test efficiency of the end-to-end automatic driving system is improved. And the testing efficiency of the end-to-end automatic driving system is greatly improved.
Owner:JILIN UNIVERSITY

Distributed streaming multi-mode fusion adaptive gradient compression optimization method and system

The invention belongs to the field of multi-modal RGB-T data fusion, and discloses a distributed streaming multi-modal fusion-oriented adaptive gradient compression optimization method and system, and the method comprises the steps: carrying out the personalized compression of data features of different modals through employing a modal-sensitive gradient compression strategy, and enabling the modals to comprise an RGB modal and an infrared modal; introducing physical constraint time sequence alignment loss, and performing time sequence alignment on the multi-modal data subjected to personalized compression by utilizing a physical model based on a thermal diffusion equation; the dynamic batch processing strategy is sensed through the video memory, the use condition of the video memory is monitored in real time, and the batch processing size is dynamically adjusted according to the residual capacity of the video memory. According to the method, a brand new solution is provided for efficient training of distributed streaming multi-modal data, and the precision of a multi-modal data fusion task and the resource utilization rate are improved.
Owner:CHONGQING NORMAL UNIVERSITY

Large model batch reasoning and data flow optimization system oriented to MOE architecture

The invention relates to the technical field of project management, in particular to a large-model batch reasoning and data flow optimization system oriented to an MOE architecture. According to the method, a collaborative architecture of the request access module, the environment sensing module, the expert routing engine, the resource scheduling module and the dynamic optimization control module is set, the text length and the subject type are extracted by using the request access module, a basis is provided for accurate routing, and the GPU video memory, the I / O bandwidth and the request queue depth are acquired in real time through the environment sensing module, so that the real-time routing is realized. The system load is comprehensively monitored, meanwhile, an expert sub-network is activated through an expert routing engine according to request features, invalid calculation is avoided, weight loading and resource allocation are managed through a resource scheduling module, the I / O bottleneck is reduced, and finally an optimization strategy is intelligently triggered through a dynamic optimization control module based on routing conflict factors. The problems of large reasoning delay fluctuation and unbalanced resource utilization rate mentioned in the background technology are solved, and stable low-delay response and resource collaborative optimization in a high-concurrency scene is realized.
Owner:VIRTAI TECH BEIJING CO LTD

Lightweight display interface rendering optimization system

The invention discloses a lightweight display interface rendering optimization system, which relates to the technical field of image processing, and comprises an acquisition module, an interface analysis module, a behavior analysis module and an attention analysis and rendering module, dynamically calculating an effective view radius by utilizing a visual tunneling effect; the interface analysis module performs character string fuzzy matching in combination with the context input by the user, identifies the search intention of the user and improves the weight of a corresponding component; the attention analysis module performs multi-modal weighted fusion on the physiological fixation data and the sparse content saliency thermodynamic diagram; according to the method, by recognizing the semantic intention and the physiological fixation point of the user, on the premise that core visual experience continuity is guaranteed, GPU load and video memory bandwidth occupation are remarkably reduced, and balance of high-performance display and low-power-consumption operation is achieved.
Owner:SHENZHEN ZHILINTAI ELECTRONIC TECH CO LTD

GPU (Graphics Processing Unit) virtualized video memory management method and device, storage medium and program product

The invention relates to the technical field of GPUs, in particular to a GPU virtualization video memory management method and device, a storage medium and a program product. The method comprises the steps that a virtual machine transmits a page table root address corresponding to a target application program running in the virtual machine to a host machine; the virtual machine calls the non-memory operation type API in response to the target application program, intercepts calling of the non-memory operation type API and forwards the calling to the host machine; and in the host machine, accessing a video memory area corresponding to a target application program based on the page table root address through the drive of the non-memory operation type API, executing an instruction of the non-memory operation type API, and returning a processing result of the instruction of the non-memory operation type API to the virtual machine. According to the method and the device, the rear-end driver in the host machine can directly access the video memory resources allocated by the virtual machine, so that the resource utilization rate and the task execution efficiency are improved.
Owner:MOORE THREADS TECHNOLOGY (SHANGHAI) CO LTD

Multi-round dialogue method, device and equipment based on large model and storage medium

The invention discloses a multi-round dialogue method and device based on a large model, equipment and a storage medium, and relates to the field of natural language processing, and the method comprises the steps: screening out a target round dialogue corresponding to a current round dialogue from historical dialogues; extracting triple information of the current round of dialogue and the target round of dialogue, and constructing a target cue word based on the domain data and the triple information corresponding to the current round of dialogue; the triple information comprises intention information, entity information and slot position information; inputting the target cue word into a preset dialogue large model, and determining a target key value cache of the target cue word and a historical key value cache corresponding to the target round of dialogue; and generating a dialogue reply of the current round of dialogue based on the target key value cache and the historical key value cache. By calculating the semantic similarity of the historical dialogue and the current round, combining entity recognition and only retaining part of content, the memory used by the dialogue is effectively reduced, logic breakage is avoided, and the use efficiency of the video memory is effectively improved through key value cache multiplexing.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Large language model semantic query acceleration method based on sparse KV Cache index

The invention discloses a large language model semantic query acceleration method based on a sparse KV Cache index. The method comprises a KV Cache semantic pruning strategy based on an attention mechanism and a set of asynchronous pipeline reasoning architecture with overlapped calculation and I / O. According to the method, the attention sparsity characteristic of the large language model in the reasoning stage and the asynchronous transmission capacity between the Host memory and the GPU video memory are fully utilized, calculation redundancy and video memory occupation in repetitive semantic query are greatly reduced, and high-throughput, low-delay and high-performance batch semantic data processing service is provided. Through a mechanism for mapping a static text into a compressed semantic index and combining a prefix cache technology in a reasoning process to realize state multiplexing, high-performance reasoning acceleration and high-efficiency storage compression are provided for a data-intensive semantic analysis task in a resource-constrained environment.
Owner:EAST CHINA NORMAL UNIV

Heterogeneous intelligent computing power optimization management scheduling system for accelerating large model reasoning task

The invention discloses a heterogeneous intelligent computing power optimization management scheduling system for accelerating a large model reasoning task, and relates to the technical field of computing power optimization management scheduling. The video memory fragmentation problem in a long sequence scene is converted into a controllable block migration task, and the performance bottleneck of a traditional video memory exchange mechanism is broken through; based on an operator-level scheduling strategy of a hardware capability fingerprint database, position coding and other compute-intensive tasks are accurately matched with vector instruction set hardware, and resource mismatch loss caused by black-box scheduling is eliminated; an expert selection process is reconstructed by an integer routing and counting sorting algorithm, near-lossless reasoning is realized at a limited node of an instruction set, and the potential value of an old computing power pool is activated.
Owner:BEIJING HUAHONG DIGITAL TECH CO LTD

OpenGL-based large-scale three-dimensional model object rapid pickup method

The invention relates to an OpenGL (Open Graphics Library)-based large-scale three-dimensional model object quick pickup method. The method comprises the following steps: converting an interactive operation of a user into a three-dimensional space structure; a CPU constructs octree space acceleration structures based on a bounding box of an object in a large-scale three-dimensional model, and leaf nodes, intersecting with a pickup ray or a pickup view pyramid, of an axis alignment bounding box in each octree space acceleration structure are screened out to serve as candidate leaf nodes; collecting triangles contained in all candidate leaf nodes, and constructing a candidate triangle set; dynamically allocating thread groups and threads of the GPU based on the number of candidate triangles in the candidate triangle set; and performing highlight display on the candidate triangles intersected with the pickup ray or the pickup view cone, and feeding back IDs of the corresponding candidate triangles. According to the method, the data transmission quantity and the video memory access pressure are remarkably reduced, the pickup response time is greatly shortened in a large-scale three-dimensional model scene, and the requirements of real-time interaction and visual operation can be met.
Owner:SHENZHEN MAIXI SOFTWARE CO LTD

Key value cache grouping quantification method and device, storage medium and electronic equipment

The invention discloses a key value cache grouping quantification method and device, a storage medium and electronic equipment. When the key value cache grouping quantification method provided by the specification is adopted to quantify key value cache data, key vector and value vector data are divided based on channel dimensions to obtain a plurality of grouped data; determining quantization parameters of the grouped data, and performing asymmetric quantization on elements in the grouped data based on the quantization parameters; finally, quantization results and corresponding quantization parameter partitions may be stored in physical blocks. According to the method, on the premise of ensuring the model generation precision, the video memory occupation of key value cache is greatly compressed, and meanwhile, the reasoning throughput is improved. Through fusion of dynamic channel grouping quantization and implicit inverse quantization, a key value cache quantization solution capable of guaranteeing the generation precision is provided for edge device end-side deployment, and the technical defect that the occupation scale of a video memory is too large when the occupation amount of the key value cache video memory is linearly increased along with the sequence length in the autoregression decoding process of a large language model is relieved.
Owner:ZHEJIANG LAB

Heterogeneous cluster-oriented resource allocation method and device and storage medium

The invention relates to a heterogeneous cluster-oriented resource allocation method and device and a storage medium, and the method comprises the steps: executing a plurality of iterations of an initial model on a single node of each heterogeneous cluster, and obtaining the performance characteristics of the model; generating a plurality of parallel strategy combinations comprising the first parallelism degree and the second parallelism degree; generating a plurality of assembly line parallel combinations based on the first degree of parallelism, calculating the load imbalance rate of each assembly line parallel combination, and screening out N combinations with the minimum load imbalance rate from the assembly line parallel combinations; traversing the second parallelism degree and the micro-processing batch in each stage corresponding to each parallel combination to generate a plurality of stage strategy combinations; and calculating video memory occupation and execution time under each stage strategy combination, screening out K stage strategy combinations of which the video memories are smaller than or equal to a video memory threshold value and the execution time is shortest, and carrying out resource allocation on the heterogeneous cluster based on the stage strategy combinations. According to the invention, the problems of low training efficiency and low resource utilization rate are solved.
Owner:ZHEJIANG LAB

Distributed rendering task scheduling method based on dynamic load balancing

The invention provides a distributed rendering task scheduling method based on dynamic load balancing, which relates to the technical field of virtual reality simulation, and comprises the following steps: analyzing and preprocessing input physical simulation data through a multi-source model, and converting the format of the processed data into a manageable virtual set data format; loading the texture data to a video memory as required in a virtual mapping mode; dynamic memory loading and coloring are carried out based on the virtual set, and LOD generation is completed; dynamic task fragmentation and scheduling are carried out on the virtual set rendering task and the heterogeneous multi-source rendering task, real-time monitoring is carried out on GPU / CPU / memory loads of heterogeneous resources, and task scheduling is carried out according to a monitoring result; through a distributed cluster mode, a high-resolution rendering task is divided into a plurality of low-resolution rendering tasks, and the low-resolution rendering tasks are rendered on different nodes. According to the invention, high-fidelity, high-response and high-reliability three-dimensional scene simulation capability is provided for high-complexity scenes such as a hybrid system.
Owner:TAIHANG LABORATORY

Video memory management method and device, electronic equipment and storage medium

The invention relates to a video memory management method and device, electronic equipment and a storage medium, the method is applied to a GPU drive in a virtual machine, and the method comprises the steps that when it is detected that the remaining amount of a GPU video memory is lower than a set remaining amount threshold value, cold pages in a GPU page table are recognized; wherein the cold page is a page table item which is not accessed in a recent preset time interval; and under the condition that the cold page is a video memory, replacing a physical address corresponding to the virtual address in the cold page to a physical address in a system memory so as to replace video memory data corresponding to the cold page to the system memory. The embodiment of the invention can improve the overall utilization rate of the video memory.
Owner:MOORE THREADS TECHNOLOGY (SHANGHAI) CO LTD

Algorithm of tensor parallel computing large model based on cpu + gpu

The invention discloses a cpu + gpu-based tensor parallel computing algorithm for a large model, which comprises the following steps of: S1, combining row parallelism and column parallelism for a linear layer in the model by adopting a mixed-dimension tensor segmentation mode, and dynamically adjusting a segmentation proportion according to a model structure and hardware resources; s2, a CPU and GPU cooperative computing mechanism is constructed, part of tasks which have large video memory requirements and are relatively simple in computation are allocated to the CPU, the GPU is responsible for computing intensive tasks, communication between CPU-GPU is optimized, and overlapping of CPU-GPU communication and GPU computation is achieved; and S3, implementing a dynamic resource allocation and load balancing strategy, monitoring load conditions of the CPU and the GPU in real time, dynamically adjusting task allocation according to calculation requirements of different layers of the model and use conditions of hardware resources, and adopting a self-adaptive batch processing size adjustment strategy. The invention obviously reduces the occupation of the video memory, reduces the communication cost and improves the utilization rate of computing resources.
Owner:GUANGDONG UNIV OF TECH +1

Distributed training system and distributed training method

The embodiment of the invention provides a distributed training system and a distributed training method. The distributed training system and the distributed training method are used for improving the training efficiency of model distributed training. Training nodes in the distributed training system are used for storing check points in video memories of GPUs of the training nodes and sending the check points to memories of CPUs of the training nodes; the training node is also used for sending the check points in the memory of the CPU to the storage system through asynchronous operation; the training node is also used for sending abnormal information to the management node under the condition that the node state information of the training node is abnormal information; the management node is also used for determining whether the training node is a fault node based on the abnormal information, and isolating the fault node under the condition that the training node is the fault node; and the management node is also used for loading a target check point corresponding to the fault node from the storage system to other training nodes under the condition that the training node is the fault node, so that other training nodes complete a model training task of the fault node based on the target check point.
Owner:XFUSION DIGITAL TECH CO LTD

Three-dimensional scene data generation method, analysis method and rendering equipment

The invention relates to the field of three-dimensional scene rendering, in particular to a three-dimensional scene data generation method, an analysis method, rendering equipment and a computer storage medium. According to the method, LOD (Level of Detail) grading processing is performed on three-dimensional scene data, foreground data in the three-dimensional scene data is divided into a plurality of LOD layers, each LOD layer is divided into a plurality of grid nodes, and each grid node is independently stored, so that the data can be loaded as required. Meanwhile, different compression modes are adopted for different data, the GPU can directly analyze and read the compressed data, decompression in a memory is not needed, occupation of the memory and video memory bandwidth is greatly reduced, and rendering efficiency is improved.
Owner:SHENZHEN XGRIDS-INNOVATION CO LTD

Method and system for dynamically adjusting display interface

The invention relates to the technical field of display analysis, in particular to a method and system for dynamically adjusting a display interface, and the method comprises the following steps: analyzing the rendering data stream of each screen in a multi-display system, identifying a cross-screen repeated static region and a screen dynamic update region, and outputting a cross-screen repeated mark set and a dynamic region coordinate set; calculating a backlight compensation demand value of each screen according to ambient light sensor data, and applying a shared rendering weight value to the cross-screen repeated static region in combination with the cross-screen repeated mark set to realize energy consumption correction of the cross-screen repeated static region; creating a single instance frame buffer for the cross-screen repeated static region; and configuring an ambient light adaptive rendering pipeline for the dynamic updating area of the screen. According to the method, video memory resources and rendering energy consumption are remarkably saved, the occupancy rate of the static content video memory is reduced, and the overall scheduling efficiency and the energy efficiency level of the system are improved.
Owner:SHANGHAI TIANLONG DIGITAL TECH CO LTD