Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

67 results about "Thread (computing)" patented technology

In computer science, a thread of execution is the smallest sequence of programmed instructions that can be managed independently by a scheduler, which is typically a part of the operating system. The implementation of threads and processes differs between operating systems, but in most cases a thread is a component of a process. Multiple threads can exist within one process, executing concurrently and sharing resources such as memory, while different processes do not share these resources. In particular, the threads of a process share its executable code and the values of its dynamically allocated variables and non-thread-local global variables at any given time.

Inter-thread data sharing method, electronic equipment, storage medium and program product

The invention relates to the technical field of high-performance computing, and provides an inter-thread data sharing method, electronic equipment, a storage medium and a program product.The method comprises the steps that a plurality of threads in a thread block are divided into at least one cooperative thread group, and the threads in each cooperative thread group are from at least two different thread bundle groups; when a write operation of a first thread in any cooperative thread group is received, writing target data into a shared register resource associated with any cooperative thread group; and when a read operation of a second thread in any cooperative thread group for the target data is received, executing a synchronous operation, and reading the target data from the shared register resource after the execution is completed. According to the method, the cooperative thread group spanning different thread beam groups is constructed, and the shared register resources are utilized for data exchange, so that efficient data sharing is directly carried out among the threads spanning the thread beam groups through the register, and the method has the advantages of low delay, high bandwidth and capability of remarkably improving parallel computing performance.
Owner:SHANGHAI BIREN TECH CO LTD

Load diagnosis and processing method and device, electronic equipment, medium and program product

The invention provides a load diagnosis and processing method and device, electronic equipment, a medium and a program product, which can be applied to the technical field of cloud computing. The method comprises the following steps: acquiring a performance index sequence of at least one candidate application process, and in response to meeting a high-load triggering condition, determining a target process to be diagnosed and generating an abnormal event identifier; obtaining thread resource occupation information of the target process in response to the abnormal event identifier, determining an abnormal thread based on the thread resource occupation information, and obtaining execution stack information corresponding to the abnormal thread; thread features are extracted based on the thread resource occupation information and the execution stack information, and a root cause classification label is determined based on the thread features; and generating processing suggestion information for the target application process based on the root cause classification label.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Thread-based sandboxing for untrusted software execution

This disclosure describes approaches for sandboxing a thread and memory resources within non-secure / secure processing environments such as in a TrustZone-M processor architecture. An example method of controlling memory access includes: providing a memory locking service in a computing device having a secure processing environment and a non-secure processing environment, and executing the memory locking service in the secure processing environment; receiving a request with the memory locking service to establish a sandbox for a particular thread that executes in the non-secure processing environment and is associated with at least one specified memory region; and associating other threads of the non-secure processing environment with the secure processing environment, such that the particular thread is unable to access memory resources of the other threads while the particular thread is sandboxed.
Owner:ANALOG DEVICES INC

Universal computing unit and computing device

The invention provides a general purpose computing unit and computing equipment. The general-purpose computing unit comprises a plurality of execution units, wherein each execution unit is used for executing an operation instruction by taking a thread bundle as a unit; and a thread local register for each execution unit, including a plurality of thread beam registers allocated for a plurality of thread beams run by the execution unit, the plurality of thread beam registers being non-overlapping, where the thread local register further includes a shared register shared by the plurality of thread beams.
Owner:SHANGHAI BIREN TECH CO LTD

Artificial intelligence chip and collaborative thread beam calculation method

The invention provides an artificial intelligence chip and a collaborative thread bundle calculation method. The cooperative thread bundle calculation method comprises the following steps of: initializing a plurality of thread bundle units by a thread bundle resource allocation unit, so that the plurality of thread bundle units have the same initial vector register base address to form a group of cooperative thread bundles; the thread beam resource allocation unit notifies a thread beam scheduling and instruction transmitting unit to send a plurality of execution instructions to an execution unit; running a plurality of thread bundles of the plurality of thread bundle units by the execution unit according to the plurality of execution instructions; reading different scalar parameters from the respective scalar registers of the plurality of thread bundle units by the plurality of thread bundles; and accessing, by the plurality of thread bundles, the same register space of the vector register according to the same initial vector register base address. According to the artificial intelligence chip and the cooperative thread beam calculation method, the efficient cooperative thread beam calculation function of the cross-thread beam sharing vector register block can be achieved.
Owner:SHANGHAI BIREN TECH CO LTD

Universal computing unit and instruction scheduling method

The invention provides a general purpose computing unit and an instruction scheduling method for the general purpose computing unit. The general-purpose computing unit comprises a plurality of execution units, wherein each execution unit is used for executing an operation instruction by taking a thread bundle as a unit; and an instruction scheduler for determining a priority of the plurality of operation instructions based on whether the multiplexing flag exists, and scheduling each operation instruction to one of the plurality of execution units based on the priority; wherein each execution unit comprises one or more reuse registers, each reuse register is used for registering an operand of a previous operation instruction and a label, and the label is used for indicating an address and a thread bundle of the operand; and the instruction decoding unit is used for decoding the received operation instruction to determine the reading position of the operand of the operation instruction. The instruction is preferentially scheduled by adding a reuse mark to the whole operation instruction, so that the hit rate of the operand is improved, and excessive instruction bit fields do not need to be occupied.
Owner:SHANGHAI BIREN TECH CO LTD

Directive generative thread-based user assistance system

Embodiments of the disclosed technologies include generating a first thread classification prompt based on a first thread portion of an online dialog involving a user of a computing device, sending the first thread classification prompt to a first large language model, receiving a first thread classification generated and output by the first large language model based on the first thread classification prompt, formulating a plan execution prompt based on the first thread classification, sending the plan execution prompt to a second large language model, receiving a second thread portion generated and output by the second large language model based on the plan execution prompt and the online dialog, and generating a label for a third thread portion of the online dialog.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

User assistance system based on instructive generative thread

Embodiments of the disclosed techniques include generating a first thread classification cue based on a first thread portion of an online conversation involving a user of a computing device; sending the first thread classification prompt word to a first large language model; receiving a first thread classification generated and output by the first large language model based on the first thread classification prompt word; formulating a plan execution cue word based on the first thread classification; sending the plan execution cue word to a second large language model; receiving a second thread part generated and output by the second large language model based on the plan execution cue word and the online dialogue; and generating a tag for a third thread portion of the online conversation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Concurrent throttling while managing upstream resources

Systems, apparatuses, and methods are disclosed for arbitrating threads in a computing system. The computing system includes a processor having a plurality of cores that are each capable of concurrently processing instructions of a plurality of threads. When a thread throttling unit receives an indication that a shared cache has a resource contention, the throttling unit sets a cache miss threshold for the cache. If the number of cache misses exceeds the threshold, the throttling unit notifies a particular upstream computing unit to throttle the processing of instructions of the thread. After a time period elapses, if the cache continues to exceed the threshold, the throttling unit notifies the upstream computing unit to more restrictively throttle the thread by performing one or more of decreasing a selection rate and increasing the time period. Otherwise, the unit notifies the upstream computing unit to less restrictively throttle the thread.
Owner:ONESTA IP LLC

Data screening, large language model output acceleration method, device, medium and product

The application discloses a data screening, large language model output acceleration method, device, medium and product. The data screening method comprises the following steps: selecting a thread bundle and determining a local comparison value of the thread bundle; triggering each thread bundle to execute an operation of comparing the numerical value between each current probability value and the baseline value of itself, and updating the key value of itself using each current probability value greater than the baseline value of itself; comparing the screening threshold of a target screening algorithm with the key value of each thread bundle; if the screening threshold and the key value of each thread bundle do not match, updating the target set and the current probability value set according to the identified first thread bundle and second thread bundle; and if the screening threshold and the key value of the third thread bundle match, updating the target set according to the third thread bundle as a screening result. The technical scheme of the embodiment of the application effectively reduces the calculation complexity of the data screening algorithm in a soft and hard cooperative manner, and maximally utilizes the parallel computing hardware in the artificial intelligence acceleration computing chip.
Owner:SHANGHAI SUIYUAN TECH CO LTD

Federal learning acceleration method based on parallel sampling and training of in-batch real-time data of assembly line

PendingCN121684100AMachine learningKnowledge based modelsEvent synchronizationAlgorithm
The invention discloses a federated learning acceleration method and system based on pipeline in-batch real-time data sampling and training. According to the method, in-batch data sampling is provided on the algorithm level aiming at the problems that computing resources of edge equipment are limited and importance is outdated and gradient deviation is caused by existing static sampling: the importance of samples in a fixed mini-batch is evaluated in real time by utilizing a latest model in each round of iteration, and a dynamic micro-batch is constructed; and through a gradient correction coefficient based on a sampling probability reciprocal, distribution deviation is eliminated, and unbiased training is realized. On the system level, an assembly line parallel mechanism based on a CPU-GPU heterogeneous architecture is designed, overlapping execution of sampling and training is achieved through a double-thread-double-flow concurrent model, data competition is solved through an annular buffer area and an event synchronization mechanism, and overhead is reduced in cooperation with mixing precision reasoning. According to the method, the hardware utilization rate can be remarkably improved, the model convergence precision is improved while the training time is greatly shortened, and the method is suitable for various edge computing scenes.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Hardware acceleration of relational operations

This disclosure describes an implementation of a computing system that utilizes accelerator hardware configured to perform parallel computation of matching pairs of a join operation of a database program. This computation is performed at least in part by, for each thread that determines a matching pair, computing an offset in an output tuple by adding a global rank of a respective block, an intra block rank of a respective warp, and an intra warp rank of a respective thread. This computation is further performed by storing index values for a primary key and a foreign key in a matching pair at the computed offset location in an output tuple, and outputting the output tuple.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Priority based power allocation

Apparatuses, systems, and techniques to allocate power to one or more processors. In at least one embodiment, processors or computing systems perform an API to allocate power to one or more processors based, at least in part, on indications of priority of one or more threads to be performed by the one or more processors.
Owner:NVIDIA CORP

Lock circuit for competing kernels in a hardware accelerator

A computing system (102), comprises a system memory (116) including a lock array (204), a processor (112) coupled to the system memory (116), a peripheral bus (115) coupled to the system memory (116), and a hardware accelerator (122). The hardware accelerator (102) comprises a bus interface (141) coupled to the peripheral bus (115), a lock circuit (140) coupled to the bus interface (141), and a plurality of kernel circuits (138) coupled to the lock circuit (140) and the bus interface (141). The plurality of kernel circuits (138) provide lock requests to the lock circuit (140), the lock requests for data (202) stored in system memory (116) of the computing system (102). The lock circuit (140) is configured to process the lock requests from the plurality of kernel circuits (138) and to issue atomic transactions over the peripheral bus (115) through the bus interface (141) based on the lock requests. The lock circuit (140) is further configured to issue atomic transactions in competition with competing threads (139) executed by the processor (112) to check the lock array (204) in the system memory (116) for exclusive access to the data (202). The lock array (204) is indexed by identifiers for certain portions of the data (202) stored in the system memory (116). The lock circuit (140) provides a single source of atomic transactions over the bus interface (141) and maintains a kernel lock array (206) indexed by identifiers for the data (202) stored in the system memory (116).
Owner:XILINX INC

A method and system for dividing GPU space resources based on persistent threads

The application discloses a GPU space resource partitioning method and system based on a persistent thread, and belongs to the field of high-performance computing and computer system structure.The method extracts original PTX code from a binary module by dynamically intercepting a CUDA interface at a driver layer; a static compiling conversion technology is used to recombine an instruction stream, construct a persistent thread structure, and innovatively introduce a "kernel fusion" mechanism in the persistent thread; through instruction body multiplication and virtual index remapping, the static demand of a single physical thread block for registers and shared memory is artificially doubled, so that the single physical thread block reaches the upper limit of hardware resources of a streaming multiprocessor; finally, a kernel starting interface is intercepted, and the original kernel is transparently replaced by the converted kernel.The application effectively solves the defect that traditional persistent threads cannot realize strict physical exclusive on a streaming multiprocessor due to unsaturated resource occupation, and realizes fine-grained and deterministic partitioning and isolation of GPU space resources without modifying user source code.
Owner:XI AN JIAOTONG UNIV

Artificial intelligence chip and cooperative thread bundle computing method

The application provides an artificial intelligence chip and a cooperative thread bundle calculation method. The artificial intelligence chip comprises M thread bundle units and a thread bundle resource allocation unit corresponding to M thread bundles respectively, wherein M is a positive integer. The thread bundle resource allocation unit is coupled with the M thread bundle units. The thread bundle resource allocation unit initializes the M thread bundle units to allocate M thread bundle numbers and M available register space sizes to the M thread bundle units. The M available register space sizes correspond to the M thread bundle numbers respectively. The thread bundle resource allocation unit sets M vector register base addresses of the M thread bundle units. An Nth vector register base address of the M vector register base addresses plus an Nth available register space size of the M available register space sizes is equal to an N+1th vector register base address of the M vector register base addresses. The artificial intelligence chip provided by the application can guarantee the register resource use efficiency.
Owner:SHANGHAI BIREN TECH CO LTD

Operator execution method, electronic device, storage medium, and program product

The present application relates to the technical field of artificial intelligence, and provides an operator execution method, an electronic device, a storage medium and a program product, the method comprising: determining the number of processing batches corresponding to each thread block based on the batch size of a query tensor; in any thread block, aggregating the query tensors of all attention heads corresponding to the number of processing batches, performing attention calculation with a key tensor and a value tensor to obtain a thread block calculation result; and integrating the thread block calculation results of each thread block to obtain an output tensor. The method provided by the present application determines the outer parallel loop by the batch size of the query tensor, changes the core operation of calculation from the originally low-efficiency vector x matrix operation to the matrix x matrix operation with higher calculation density, fully utilizes the hardware characteristics of the tensor core optimized for matrix multiplication, greatly improves the throughput and execution efficiency of the calculation unit, and further gives full play to the potential computing power of the GPU, thereby further improving the execution speed of the overall operator.
Owner:SHANGHAI BIREN TECH CO LTD

General purpose computing unit and instruction scheduling method

This disclosure provides a general-purpose computing unit and an instruction scheduling method for the general-purpose computing unit. The general-purpose computing unit includes: multiple execution units, each execution unit executing arithmetic instructions in units of thread bundles; and an instruction scheduler for determining the priority of multiple arithmetic instructions based on whether they have a reuse tag, and scheduling each arithmetic instruction to one of the multiple execution units based on the priority; wherein each execution unit includes: one or more reuse registers, each reuse register storing an operand and a tag of a previous arithmetic instruction, the tag indicating the address of the operand and the thread bundle; and an instruction decoding unit for decoding the received arithmetic instruction to determine the read location of the operand of the arithmetic instruction. By attaching a reuse tag to the entire arithmetic instruction to prioritize the scheduling of the instruction, the operand hit rate is improved, and excessive instruction bit fields are not required.
Owner:SHANGHAI BIREN TECH CO LTD

Artificial intelligence chip and operation method thereof

The invention provides an artificial intelligence chip and an operation method thereof. The artificial intelligence chip comprises a shared memory, a plurality of computing cores, a plurality of thread beam scheduling and instruction transmitting units and a thread block scheduling unit. The thread block scheduling unit is used for segmenting each thread block into a memory access type thread bundle and a calculation type thread bundle; and issuing the memory access type thread bundle of the first thread block so as to load the first corresponding data of the first thread block to a data preparation area in the shared memory. And issuing the calculation type thread bundle of the first thread block to enable the calculation type thread bundle of the first thread block to take the first corresponding data in the data preparation area. And after the first corresponding data in the data preparation area is taken and used and before the first thread block exits, issuing the memory access type thread bundle of the second thread block in advance so as to pre-load second corresponding data of the memory access type thread bundle of the second thread block to the data preparation area in the shared storage.
Owner:SHANGHAI BIREN TECH CO LTD

Thread block based warp scheduling method and GPU scheduling device

This invention discloses a thread bundle scheduling method and a GPU scheduling device based on thread blocks. By employing a granular scheduling strategy of prioritizing the oldest thread bundle among thread bundles and round-robin scheduling within thread blocks, it ensures consistent execution progress of thread bundles within the same thread block, reducing synchronization waiting overhead and meeting high-concurrency synchronization requirements. Simultaneously, it achieves progress staggering between different thread blocks, avoiding collective resource contention and collective memory access requests, resolving resource conflicts and memory access latency-induced failures, and perfectly adapting to the load characteristics of large models. The GPU scheduling device can be directly integrated into existing GPU architectures, requiring only the addition of a few functional units to the existing thread bundle scheduling module. It does not require modification of the GPU's computing cores, registers, memory access modules, or other hardware structures, resulting in low hardware modification costs and ease of implementation and deployment.
Owner:BEIJING FENGHUA CHUANGZHI TECHNOLOGY CO LTD

Safe, secure, virtualized, domain specific hardware accelerator

This disclosure relates to various implementations an embedded computing system. The embedded computing system comprises a hardware accelerator (HWA) thread user and a second HWA thread user that creates and sends out message requests. The HWA thread user and the second HWA thread user is communication with a microcontroller (MCU) subsystem. The embedded computing system also comprises a first inter-processor communication (IPC) interface between the HWA thread user and the MCU subsystem and a second IPC interface between the second HWA thread user and the MCU subsystem, where the first IPC interface is isolated from the second IPC interface. The MCU subsystem is also in communication with a first domain specific HWA and a second domain specific HWA.
Owner:TEXAS INSTRUMENTS INC

Processing system of thread block, method and relative device

The embodiments of the present disclosure provide a processing system of a thread block, a method and a relative device. The processing system includes: a first computing unit for running the first sub-thread block and a second computing unit for running the second sub-thread block; the first computing unit is used for obtaining the data to be processed of the thread block, and the second computing unit is used for executing the processing task of the thread block according to the data to be processed obtained by the first computing unit.
Owner:HYGON INFORMATION TECH CO LTD

Artificial intelligence chip and cooperative thread bundle computing method

The application provides an artificial intelligence chip and a cooperative thread bundle calculation method. The cooperative thread bundle calculation method comprises the following steps: initializing a plurality of thread bundle units by a thread bundle resource allocation unit, so that the plurality of thread bundle units have the same initial vector register base address to form a group of cooperative thread bundles; informing a thread bundle scheduling and instruction emission unit to send a plurality of execution instructions to an execution unit by the thread bundle resource allocation unit; running a plurality of thread bundles of the plurality of thread bundle units according to the plurality of execution instructions by the execution unit; reading different scalar parameters from respective scalar registers of the plurality of thread bundle units by the plurality of thread bundles respectively; and accessing the same register space of the vector register according to the same initial vector register base address by the plurality of thread bundles. The artificial intelligence chip and the cooperative thread bundle calculation method can realize the cooperative thread bundle calculation function of efficiently sharing the vector register group across the thread bundles.
Owner:SHANGHAI BIREN TECH CO LTD

Computing device and its operating method

This invention provides a computing device and its operating method. The computing device includes multiple execution units, a global cache, and a data memory access unit. Each execution unit includes a first type of computing core, a second type of computing core, and a shared cache. The execution unit executes a thread group. The thread group includes multiple thread groups. The first thread group in the thread group controls the data memory access unit to transfer data between the global cache and the shared cache. Other thread groups in the thread group use the first type of computing core and the second type of computing core to perform calculations on the data located in the shared cache. The first operating time period of the first thread group partially overlaps with the second operating time periods of other thread groups to run in parallel. This invention's computing device and its operating method enable different types of computing cores in the same execution unit to execute in parallel, thereby making full use of computing resources.
Owner:SHANGHAI BIREN TECH CO LTD

Loading elements in a computing environment

A system determines a trigger for loading a set of elements on a background data area. In response to determining the trigger for loading the set of elements on the background data area, the system executes a first thread to load the set of elements on the background data area. The first thread accesses a set of element identifiers that identify the set of elements arranged in a sequence corresponding to a traversal of the set of elements to transitive closure. Based on the set of element identifiers, the first thread accesses a dataset that includes data for loading the set of elements and loads the set of elements on the background data area in the sequence corresponding to the traversal of the set of elements to transitive closure. A second thread maps the set of elements from the background data area to a runtime data area.
Owner:ORACLE INT CORP

GPU collision detection acceleration method based on logic group filtering and load balancing

The invention discloses a GPU (Graphics Processing Unit) collision detection acceleration method based on logic group filtering and load balancing, and aims to solve the efficiency bottleneck caused by uneven thread load and semantic irrelevant calculation redundancy in collision detection in large-scale multi-robot simulation. In a logic layer, a robot is endowed with a group identifier, a logic group filtering rule is designed, object pairs which do not need to be detected are removed in a preprocessing stage, and the calculated amount is reduced from the source; and secondly, when a computing layer runs on a GPU (Graphic Processing Unit), grouping candidate geometry pairs screened out in a wide stage according to computing complexity implied in a geometry type, and dynamically distributing an exclusive kernel function for each task group to carry out accurate collision detection, so that the thread divergence problem is radically solved, and the extreme load balancing is realized. Through the two-stage optimization, heterogeneous computing tasks are converted into isomorphic task groups, the GPU utilization rate is remarkably improved, and an efficient collision detection solution is provided for large-scale robot simulation training.
Owner:ZHEJIANG UNIV

Compiler-based ordered view method, electronic device, and storage medium

The present application relates to the field of rasterization technology in chip design, in particular to an ordered view method based on a compiler, an electronic device and a storage medium, which obtains unique identifiers of each thread group within a current rendering target range, and makes the thread groups access the buffer in turn according to the order through the processing of the global thread group order mark and the unique identifier of the thread group; each thread group in the target tile queue distributes the current maximum order mark in the tile and takes it as an execution order mark, and executes and updates the graphic resource in turn according to the order of the execution order mark, so that the order of the thread group accessing the buffer and executing is constrained, the thread groups in the tiles are executed in parallel, and the thread groups in the tiles are executed in series, so that the parallel advantage of the GPU is fully utilized under the premise of guaranteeing the ordered view, the computing resources are fully utilized, and the rendering efficiency is improved.
Owner:METAX INTEGRATED CIRCUITS (SHANGHAI) CO LTD

Ordered view method based on compiler, electronic equipment and storage medium

The invention relates to the technical field of rasterization in chip design, in particular to an ordered view method based on a compiler, electronic equipment and a storage medium. Enabling the thread groups to access the buffer area in sequence through the processing of the global thread group sorting marks and the unique identifiers of the thread groups; the method comprises the following steps of: distributing a current maximum sorting mark in a graph block in each thread group queued by a target graph block, taking the current maximum sorting mark as an execution sequence mark, and sequentially executing and updating graph resources according to the sequence of the execution sequence mark, so that parallel execution of the thread groups among the graph blocks is realized by constraining the access of the thread groups to a buffer area and the execution sequence; and the thread groups in the blocks are executed in series, so that the aim of fully exerting the parallel advantages of the GPU on the premise of ensuring the ordered view is fulfilled, the computing resources are fully utilized, and the rendering efficiency is improved.
Owner:METAX INTEGRATED CIRCUITS (SHANGHAI) CO LTD