Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

112 results about "Massively parallel" patented technology

In computing, massively parallel refers to the use of a large number of processors (or separate computers) to perform a set of coordinated computations in parallel (simultaneously). In one approach, e.g., in grid computing the processing power of many computers in distributed, diverse administrative domains, is opportunistically used whenever a computer is available. An example is BOINC, a volunteer-based, opportunistic grid system, whereby the grid provides power only on a best effort basis.

Large model calculation acceleration chip architecture

The invention discloses a large model calculation acceleration chip architecture, and relates to the field of large models. Comprising a normalized network-on-chip, an interconnection transmission system, a storage control system and a calculation acceleration core, the storage control system is connected with the Memory video memory group and the external SSD memory based on different built-in storage controllers; the normalized network-on-chip reads on-chip model parameters of a target position based on the storage control system, sends the on-chip model parameters to the calculation acceleration core for calculation, restores on-chip model data, carries out off-chip interaction with an external management server based on the interconnection transmission system, receives an instruction task, reads and restores external model parameters, and sends the external model parameters to the calculation acceleration core for calculation. And performing inter-chip interaction with other computing acceleration chips based on an interconnection transmission system, and reading and restoring inter-chip model parameters. According to the scheme, the hybrid video memory computing chip architecture supports multi-parallel SSD direct connection access, the capacity is expanded at low cost under the condition of keeping the bandwidth, and high cost performance is kept for both large-scale parallel tasks and small-sized low-cost all-in-one machines.
Owner:STORAGEX TECH INC

Intelligent water operation management and control system based on Internet of Things

The invention discloses an intelligent water operation management and control system based on the Internet of Things, and relates to the technical field of industrial Internet of Things crossing, and the system comprises the steps: collecting multi-dimensional operation data in a water supply network, carrying out the preliminary time mark alignment and preprocessing, and generating a multi-source synchronous perception data set; importing the comprehensive risk list into a digital twinborn body, dynamically injecting corresponding fault and abnormal scene parameters in a digital twinborn environment, and performing large-scale parallel deduction through a Monte Carlo simulation algorithm to generate a rehearsal consequence data set; and carrying out multi-objective decision analysis on the rehearsal consequence data set, carrying out prediction comprehensive plan utility scoring based on a preset safety weight and an economic weight, and calculating an optimized management and control plan through iterative optimization. According to the invention, the multi-source heterogeneous data is deeply fused through the graph neural network of the risk identification module, the dynamic impedance characteristics of the pipe network are accurately extracted, the early abnormality is identified, and advanced diagnosis of the multi-modal coupling risk is realized.
Owner:JIANGSU ZHONGKE MONENG INTELLIGENT ENVIRONMENTAL TECH CO LTD

Optical Chip-to-Chip Interconnect and Method of Integration

In order to enable applications such as artificial intelligence (AI) and machine learning in a large scale, a large amount of information needs to be processed very fast. To improve the speed of the computation, not only faster graphical and central processors are required, but also the communication between processors and high bandwidth memories should be fast enough to reduce the latency. Optical interconnect can be a replacement for copper interconnect which is providing larger bandwidth and lower latency. As an efficient light source, micro-LEDs can be used as light source for chip-to-chip optical interconnect. Using a massively parallel memory-processor interconnect using low-power micro-LED, the computation speed will increase dramatically, and less amount of energy will be wasted through copper self-heating.
Owner:HYPERLUME INC

Inspection data processing method and system

The invention discloses an inspection data processing method and system, relates to the technical field of data processing, and ensures the accuracy and stability of authentication scheduling in a multi-interface cluster environment by responding to concurrent tasks and dynamically selecting authentication tokens according to interface features. The acquisition request is constructed based on the interface characteristics, so that the adaptive access to the heterogeneous interface path is realized, and the compatibility of large-scale parallel acquisition is guaranteed; by uniformly converting original data into a structured format and carrying out anomaly detection and linked list storage, the problems that multi-source heterogeneous data is difficult to process and an abnormal state is difficult to systematically track are solved; and finally, a standardized report is automatically generated according to a preset rule, and accurate marking of abnormal information is realized, so that full-process automation from data acquisition to report output is completed, and the reliability, the processing efficiency and the intelligent level of inspection operation are remarkably improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Rapid data segmentation method for large-scale parallel graph neural network calculation

The invention discloses a rapid data segmentation method for large-scale parallel graph neural network calculation, and relates to the field of data segmentation. The method comprises the steps that when streaming graph processing is carried out on vertexes, a super adjacency list of k-Hop is constructed, and therefore the complex dependency relationship in the GNN calculation process can be better captured. The influence of different types of neighbors on vertex distribution is particularly concerned, in the GNN training process, nodes are usually divided into training points, test points and verification points, and the influences of different types of neighbors on vertexes are different. Due to the fact that multiple rounds of iteration are usually carried out in the training stage and only one round is carried out in the testing stage, for a certain node, the node and a training neighbor are more expected to be located in the same partition, and communication overhead is reduced. Through partition information of the distributed nodes, expected values of neighbors in the partitions are increased. Namely, through the combined action of the distribution information of the inner neighbors of each node and the distribution information of the out neighbors in the k-Hop adjacency list, the effectiveness of division is enhanced.
Owner:OCEAN UNIV OF CHINA +2

Unstructured grid division method based on structured grid and electronic equipment

The invention provides a structured grid-based unstructured grid division method and electronic equipment, and the method comprises the steps: determining a computational domain range of a combustion detonation simulation task, carrying out the surface grid division of a geometric model, obtaining a geometric representation surface grid, generating a structured grid according to the geometric representation surface grid and the computational domain range, and carrying out the calculation of a combustion detonation simulation task. The method comprises the following steps: constructing a structured grid, performing simulation calculation on the structured grid, and determining a grid encryption parameter according to a simulation calculation result of the structured grid, so that an unstructured grid is generated according to the grid encryption parameter, a calculation domain range and a geometric representation surface grid, and unstructured grid division based on the structured grid is realized. The method can replace the trial-and-error iterative optimization or complex dynamic adaptive process of the unstructured grid, improves the generation efficiency of the unstructured grid, guarantees the generation quality of the unstructured grid, balances the calculation efficiency while meeting the high-precision requirement, and is suitable for a large-scale parallel environment.
Owner:国家超级计算天津中心 +1

Action-based graph framework and query method

Disclosed is a network and method for creating and searching a graph in a massively parallel manner. The network includes a plurality of graph storage instances that collectively store a graph formed from a plurality of entities and comprising a vertex for each entity, edges connecting pairs of vertices, and adjacency relations. Each graph storage instance stores a subgraph of the graph and comprises a partition with a non-overlapping vertex set of all the vertices in the respective subgraph, an edge set of all the edges between vertices in the respective vertex set, and the adjacency relations for the edge set and to any edges outside the partition, to which the edge set connects. Each graph storage instance also includes an executor for executing one or more of the actions over the respective subgraph, and an action monitor for supplying actions to the respective executor.
Owner:GRABTAXI HOLDINGS PTE LTD

Massively parallel in-network compute

Efficient scaling of in-network compute operations to large numbers of compute nodes is disclosed. Each compute node is connected to a same plurality of network compute nodes, such as compute-enabled network switches. Compute processes at the compute nodes generate local gradients or other vectors by, for instance, performing a forward pass on a neural network. Each vector comprises values for a same set of vector elements. Each network compute node is assigned to, based on the local vectors, reduce vector data for a different a subset of the vector elements. Each network compute node returns a result chunk for the elements it processed back to each of the compute nodes, whereby each compute node receives the full result vector. This configuration may, in some embodiments, reduce buffering, processing, and / or other resource requirements for the network compute node or network at large.
Owner:INNOVIUM INC

Structural grid based unstructured mesh partitioning method and electronic device

The application provides a structure grid-based unstructured grid division method and an electronic device. The method determines a calculation domain range of a combustion detonation simulation task, performs surface grid division on a geometric model to obtain a geometric surface representation grid, generates a structure grid according to the geometric surface representation grid and the calculation domain range, performs simulation calculation on the structure grid, determines a grid encryption parameter according to a simulation calculation result of the structure grid, generates an unstructured grid through the grid encryption parameter, the calculation domain range and the geometric surface representation grid, realizes the structure grid-based unstructured grid division, can replace a trial-and-error iterative optimization or a complex dynamic self-adaptive process of the unstructured grid, improves the generation efficiency of the unstructured grid and guarantees the generation quality of the unstructured grid, balances the calculation efficiency while meeting the high-precision requirement, and is suitable for a large-scale parallel environment.
Owner:国家超级计算天津中心 +1

Key switching method and system based on graphics processor

The invention discloses a key switching method and system based on a graphics processor, and belongs to the technical field of information security, and the key switching method comprises the following steps: setting a corresponding sequence for assembling a number-theory transformation kernel function, a precise basis transformation kernel function and an inner product kernel function which are realized on the graphics processor according to key switching requirements, therefore, the key switching in the high-throughput fully homomorphic encryption technology based on the graphics processor is realized. On one hand, large-scale parallelism of three kernel functions among coefficients and moduli of fully homomorphic encryption ciphertext polynomials is realized on a graphics processor, so that large-scale parallelism and high-throughput design of fully homomorphic encryption key switching kernel functions are realized, and the calculation efficiency is improved; and on the other hand, flexible parameter selection is supported, and the user is allowed to realize high-throughput key switching of any password parameter under the condition that no code modification is needed.
Owner:HUAZHONG UNIV OF SCI & TECH

Method and device for large-scale parallel processing of database

Embodiments of this specification provide a method and apparatus for large-scale parallel database processing, wherein the method comprises: receiving a search instruction, wherein the search instruction includes a vector to be searched corresponding to the item data to be searched; using an approximate nearest neighbor search method to determine at least one search feature vector corresponding to the vector to be searched in a pre-generated feature vector index table, and obtaining a search computation node corresponding to each search feature vector, wherein the feature vector index table records multiple feature vectors and the computation node corresponding to each feature vector; searching for an approximate vector corresponding to the vector to be searched in each search computation node; and determining a target vector corresponding to the vector to be searched from at least one approximate vector. This method significantly reduces computational overhead while ensuring search accuracy.
Owner:ALIBABA GROUP HOLDING LTD

A sparse attention calculation method, device and medium for a GPU

The present application relates to the technical field of GPU computing optimization, and in particular to a sparse attention computing method, device and medium for GPU, wherein the method realizes the high efficiency of long context reasoning through the geometric perception sparse attention framework of ball hashing, and combines a large-scale parallel hashing optimization algorithm and a load adaptive computing kernel. Compared with the existing sparse attention methods based on heuristics or gradient learning, the present application realizes higher retrieval recall rate, lower preprocessing overhead and efficient hardware adaptation to irregular sparse patterns.
Owner:CENT SOUTH UNIV

Heterogeneous symmetric matrix eigenvalue decomposition method, system and equipment based on GPU parallel acceleration

The invention provides a heterogeneous symmetric matrix eigenvalue decomposition method, system and device based on GPU parallel acceleration, and belongs to the technical field of high-performance computing. According to the method, the design that five sub-processes can only be strictly and serially executed in the traditional symmetric matrix eigenvalue decomposition process is broken through, the problem that CPU and GPU computing resources are idle in a serial algorithm is solved, and the computing efficiency is greatly improved. Meanwhile, aiming at the BC-Backk stage, the invention provides a novel BC-Backk method realized based on a GPU (Graphics Processing Unit), the BC-Backk process is transferred from a CPU (Central Processing Unit) with limited parallel capability to a GPU capable of realizing large-scale parallel, and the parallel efficiency of the BC-Backk stage is greatly improved. According to the method, heterogeneous architecture characteristics and GPU parallelism are fully utilized, parallel between symmetric matrix eigenvalue solving sub-processes and internal parallel of BC-Back are achieved, CPU and GPU resources are fully utilized, and the calculation performance of the interior of a computer for the symmetric matrix eigenvalue solving process is greatly improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

A parallel situation gravity calculation chip and method based on in-memory computing

PendingCN122509213AEliminate bottlenecksreduce power consumptionMassively parallelCrossbar switch
This invention discloses a parallel situational gravity calculation chip and method based on in-memory computing, belonging to the field of artificial intelligence chips and analog in-memory computing technology. The chip includes: an event polarity input unit, which converts externally input event information loads (six-dimensional polarity vectors) into analog voltage signals; a situational reference storage array, employing a memristor cross-switch array, with each row storing a six-dimensional binary encoded vector of a reference situational type (such as the 64 hexagrams); a parallel gravity calculation unit, which performs large-scale parallel multiply-accumulate operations simultaneously in the analog domain, following the gravity calculation rule of "like lines cooperate, dissimilar lines repel," generating a multi-dimensional gravity intensity distribution; and a gravity readout and post-processing unit, which converts the gravity distribution into situational awareness results and outputs them. This invention adopts an in-memory computing architecture, eliminating the "memory wall" bottleneck of the traditional von Neumann architecture, achieving nanosecond-level fully parallel gravity calculation and ultra-low power consumption operation. It can be deployed as an independent situational awareness processor at sensor terminals, IoT terminals, and wearable devices.
Owner:PUTIAN ZIXU LIFE TECHNOLOGY CO LTD

Massively parallel on-chip coalescence of microemulsions

Embodiments disclosed herein are directed to microfluidic devices that allow for scalable on-chip screening of combinatorial libraries and methods of use thereof. Droplets comprising individual molecular species to be screened are loaded onto the microfluidic device. The droplets are labeled by methods known in the art, including but not limited to barcoding, such that the molecular species in each droplet can be uniquely identified. The device randomly sorts the droplets into individual microwells of an array of microwells designed to hold a certain number of individual droplets in order to derive combinations of the various molecular species. The paired droplets are then merged in parallel to form merged droplets in each microwell, thereby avoiding issues associated with single stream merging. Each microwell is then scanned, e.g., using microscopy, such as high content imaging microscopy, to detect the optical labels, thereby identifying the combination of molecular species in each microwell.
Owner:MASSACHUSETTS INST OF TECH +1

A parallel finite element mesh stitching method for flow channels of reactor unit components

The present invention discloses a parallelized finite element mesh stitching method for flow channels of reactor unit assemblies, belonging to the field of finite element mesh technology; the present invention classifies cavities according to their characteristics, and then stitches cavities of different categories using different methods. In addition to the original point stitching algorithm, the algorithm also adaptively applies a variety of algorithms such as particle swarm optimization, shortest path problem, and advancing wavefront method, and achieves high-efficiency stitching of large-scale meshes while ensuring the overall mesh quality through parallel processing of the algorithm and tetrahedron quality detection optimization algorithm. The present invention helps to optimize mesh quality, promote the application scenarios of large-scale parallel mesh division algorithms, fill the gap in three-dimensional mesh stitching technology, and further promote the application and development of mesh technology in the field of numerical calculation.
Owner:UNIV OF SCI & TECH BEIJING

Paillier homomorphic encryption acceleration board card supporting multi-chip cooperative computing and scheduling method of Paillier homomorphic encryption acceleration board card

The invention discloses a Paillier homomorphic encryption acceleration board card system supporting multi-chip cooperative computing. The Paillier homomorphic encryption acceleration board card system comprises a central scheduling control module, a plurality of isomorphic encryption computing chips, an inter-chip communication module, a host interface module and an auxiliary function module. And the central scheduling control module is used for analyzing a ciphertext structure and dynamically scheduling chip resources according to a task scale, and supports multiple cooperative computing modes such as ciphertext shunting, task fragmentation and pipeline processing. And the inter-chip communication module adopts AXI interconnection or an NoC structure to realize high-speed data exchange. Parallel calculation of modular exponent operation is completed among the slave chips in a master-slave cooperation mode, and dynamic modular basis configuration, chip activation control and a thermal management mechanism are supported. The system is suitable for being deployed on a cloud server, an edge device and a multi-card cluster platform, large-scale parallel acceleration processing of the Paillier encryption algorithm is achieved, and the system has high expansibility and high reliability.
Owner:HANGZHOU TOPOLOGY MACRO SEMICONDUCTOR CO LTD

Efficient coupling parallel method for dense matrix and sparse matrix

The invention discloses a dense matrix and sparse matrix efficient coupling parallel method. The method comprises the following steps: S1, carrying out feature analysis on an input dense matrix and sparse matrix; s2, carrying out adaptive partitioning on the dense matrix and the sparse matrix; s3, creating a uniform matrix descriptor for each matrix block; s4, constructing a calculation dependency graph; s5, executing parallel computing; according to the dense matrix and sparse matrix efficient coupling parallel method, through data rearrangement of self-adaptive hybrid storage and cache perception, the cache hit rate is greatly increased, the bottleneck of memory access mode conflicts in hybrid calculation is overcome, and meanwhile, based on an accurate prediction model and dynamic scheduling, the performance data in the running process are collected. The method has the advantages that combined optimization of computing load and communication is achieved, high load balance and resource utilization rate under large-scale parallel are guaranteed, redundant format conversion overhead is almost eliminated by the aid of unified descriptors and inert conversion strategies, a computing pipeline is smoother, and overall performance is improved remarkably.
Owner:NAT SUPERCOMPUTING WUXI CENT

Sinusoidal signal generation method and device, and cosine signal generation method and device

According to the sinusoidal signal generation method and device and the cosine signal generation method and device provided by the invention, based on triangle and angle formula expansion, a sinusoidal signal is generated in a manner of generating an iterative value through preconfiguration of a basic value and pipeline multiplication and addition iteration, and complex logics such as data cache management of a DDS and multi-round iterative calculation of Cordic are avoided. Moreover, the multiply-add operation module of the m-level assembly line splits the logic combination of multiplication and addition operation in multiply-add operation into m clock periods, so that the operation frequency of the FPGA is optimized. The method only depends on common logic operations such as accumulative counting, data selection and basic multiplication and addition, a large number of BRAMs for storing sine wave complete data in a DDS scheme are not needed, large-scale parallel iterative logic needed by a Cordic scheme for guaranteeing precision and throughput rate is also avoided, only a small number of LUT tables and DSP resources are needed in actual implementation, and hardware resource occupation of an FPGA is greatly reduced.
Owner:SICHUAN CHUANGZHI LIANHENG TECH CO LTD

Multimodal low-precision inner product computation circuit for massively parallel neural inference engines

A neural inference chip for computing neural activations is provided. In various embodiments, the neural inference chip is adapted to: receive an input activation tensor comprising a plurality of input activations; receive a weight tensor comprising a plurality of weights; Booth-recode each of the plurality of weights into a plurality of Booth-encoded weights, each Booth-encoded value having an order; multiply the input activation tensor by the Booth-encoded weights to produce a plurality of results for each input activation, each result in the plurality of results corresponding to an order of the Booth-encoded weight; for each order of the Booth-encoded weights, sum the corresponding results to produce a plurality of partial sums, one partial sum for each order; and compute the neural activation from the sum of the plurality of partial sums.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Action-based graph framework and query method

Disclosed is a network and method for creating and searching a graph in a massively parallel manner. The network includes a plurality of graph storage instances that collectively store a graph formed from a plurality of entities and comprising a vertex for each entity, edges connecting pairs of vertices, and adjacency relations. Each graph storage instance stores a subgraph of the graph and comprises a partition with a non-overlapping vertex set of all the vertices in the respective subgraph, an edge set of all the edges between vertices in the respective vertex set, and the adjacency relations for the edge set and to any edges outside the partition, to which the edge set connects. Each graph storage instance also includes an executor for executing one or more of the actions over the respective subgraph, and an action monitor for supplying actions to the respective executor.
Owner:GRABTAXI HOLDINGS PTE LTD

Large-scale parallel sampling environment construction method and device for reinforcement learning of intelligent computing cloud platform through computing power

The invention provides a large-scale parallel sampling environment construction method and device for reinforcement learning of an intelligent computing cloud platform through computing power, and the method comprises the steps: obtaining a plurality of physical disks of a plurality of computing power nodes of the intelligent computing cloud platform, and constructing the plurality of physical disks into a bottom storage pool through an RAID technology, storing the basic mirror image of the operating system in a qcow2 format, and further generating a plurality of virtual machine instances and running the virtual machine instances; packaging the QEMU process and the running environment of the QEMU process into a container, and performing distributed scheduling on the container according to the disk topological relation among the plurality of bottom storage pools on the plurality of computing power nodes and the real-time I / O load condition. Therefore, the construction of a large-scale parallel sampling environment of reinforcement learning can be realized by utilizing the computing power of an intelligent computing cloud platform and integrating storage optimization, hardware virtualization and container arrangement technologies.
Owner:DATACANVAS LTD

Massively parallel cell analysis and sorting apparatus and method

A massively parallel microfluidic chip is provided having multiple sections stacked or layered along a stacking direction to form multiple microchannels at least partially oriented for flow along the stacking direction. The multiple sections can include a transfer section for introducing a sample fluid containing particles, a particle focusing section configured to focus particles in the sample fluid, and an actuation section including multiple interrogation regions and multiple actuators. Each interrogation region and actuator is associated with at least one microchannel within the multiple microchannels. The arrangement of the microfluidic channels along the stacking direction allows for a very high packing density of channels and interrogation regions on a single chip, providing massively parallel processing of particles.
Owner:SITE NOME ESTÉE LELSEY

Debugging system and method

The invention discloses a debugging system and method, and relates to the technical field of equipment testing, and a link controller is connected with a crossbar switch matrix controller and is used for receiving a debugging instruction issued by a debugging host. And determining link configuration parameters according to the debugging instruction, the priority of the debugging instruction and the current link state. The link controller converts the link configuration parameters into control signals and transmits the control signals to the crossbar switch matrix controller. And the crossbar switch matrix controller adjusts the connection mode of the crossbar switch matrix and each test access port according to link configuration parameters contained in the control signal to construct a matched debugging link, so that an optimal debugging link conforming to the current debugging instruction is formed, and effective and rapid transmission of debugging data is ensured. The efficiency boundary of the serial topology is broken through while the protocol compatibility is maintained, the quick response capability and multi-task parallel support are provided for high-performance equipment, and the overall debugging efficiency of the system is improved. And efficient cooperative processing of large-scale parallel debugging tasks is supported.
Owner:JINAN MAIWEI INTELLIGENT TECHNOLOGY CO LTD

Method and device for predicting performance of application program in heterogeneous system

The invention discloses an application program performance prediction method and device under a heterogeneous system, and belongs to the technical field of system performance prediction.The method comprises the steps that S100, small-scale programs are extracted from large-scale parallel programs to be predicted; s200, running the small-scale program on a target heterogeneous system, and collecting performance data related to a hardware architecture; s300, constructing a performance model based on the performance data; s400, training the performance model by adopting a staged migration training strategy; and S500, inputting the program parameters of the large-scale parallel program into the performance model, and outputting a performance trend prediction result on the current heterogeneous system. According to the method, large-scale performance prediction with low cost, high precision and strong extrapolation is realized through a representative small-scale program, a dual-channel neural network and staged migration training.
Owner:COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI

Large-scale parallel simulation method based on in-situ decentration field data aggregation, processing architecture and storage medium

The invention discloses a large-scale parallel simulation method based on in-situ decentralized field data aggregation, a processing architecture and a chip, and relates to the field of high-performance computing and artificial intelligence. The architecture comprises a memory coupling array, a data distribution network and a parallel sampling interface. Wherein the storage and calculation array is integrated with an analog resistance network field generation layer, and global field calculation of O (1) complexity is realized by using the Kirchhoff's law; a hierarchical dynamic bypass bus (VST) is integrated, so that zero-delay jump transmission of sparse data is realized; a congestion sensing type asynchronous clock island is integrated, and the relativistic time expansion effect is simulated in a hardware layer to balance loads. According to the method, complex vector mechanics and interactive calculation are converted into a linear signal processing process through the in-situ polymerization characteristic of a physical storage array, the computing power and bandwidth bottlenecks in large-scale multi-body simulation are solved, and the method can be widely applied to large-model (LLM) reasoning, automatic driving point cloud processing and neuromorphic calculation. And the system has extremely high energy efficiency ratio and universality.
Owner:王俊鹏

Massively parallel in-network compute

Efficient scaling of in-network compute operations to large numbers of compute nodes is disclosed. Each compute node is connected to a same plurality of network compute nodes, such as compute-enabled network switches. Compute processes at the compute nodes generate local gradients or other vectors by, for instance, performing a forward pass on a neural network. Each vector comprises values for a same set of vector elements. Each network compute node is assigned to, based on the local vectors, reduce vector data for a different a subset of the vector elements. Each network compute node returns a result chunk for the elements it processed back to each of the compute nodes, whereby each compute node receives the full result vector. This configuration may, in some embodiments, reduce buffering, processing, and / or other resource requirements for the network compute node or network at large.
Owner:INNOVIUM INC

A resource scheduling method for cloud rendering clusters based on AI distributed deployment

This invention provides a resource scheduling method for cloud rendering clusters based on AI distributed deployment. First, the original instruction stream is divided into instruction particles carrying energy consumption labels, latency labels, and complementarity weights. A cross-task instruction resonance graph is constructed to locate complementary candidate resonance particle groups. Placeholder particles are generated based on time-series prediction to fill upcoming instruction stream gaps. Particles are exchanged and spliced ​​across GPUs in a dynamic splicing network, solidifying into a continuous target instruction stream across GPUs. Simultaneously, real-time power consumption and performance indicators are collected, and a joint performance-energy consumption evaluation is performed on multiple candidate splicing schemes to select and apply the optimal scheme. Without compromising business correctness, this invention significantly improves GPU execution unit utilization, reduces end-to-end latency, and lowers energy consumption through continuous pipeline, zero-copy data reuse, and rollback anchor point guarantees. It is suitable for large-scale parallel scenarios such as film and television rendering, cloud gaming, and digital twins.
Owner:SHANGHAI ITHELP NETWORK TECH CO LTD

Compression of bitstream indexes for wide scale parallel entropy coding in neural-based video codecs

Systems and techniques are described herein for processing video data. For example, an encoding device can obtain a sequence of video data and determine a minimum value in the sequence of video data. The encoding device can, based on the minimum value, identify positions in the sequence of video data associated with entry points for individually entropy codable parcels of a parallel entropy codable sequence of video data. The encoding device can generate the parallel entropy codable sequence of video data. The encoding device can further generate an index for the parallel entropy codable sequence of video data, the index identifying the individually entropy codable parcels within the parallel entropy codable sequence of video data.
Owner:QUALCOMM INC