Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1939results about "Concurrent instruction execution" patented technology

Two-level context caching and eviction for scatter-gather DMA

One aspect of the instant disclosure may provide a system and method for processing scatter-gather direct memory access (S-G DMA) instructions. During operation, the system may receive an S-G DMA instruction associated with a message and gather instruction context for the S-G DMA instruction. An S-G DMA processor may process the S-G DMA instruction based on the gathered instruction context and determine whether there exists a pending S-G DMA instruction associated with the message. In response to the presence of the pending S-G DMA instruction, the system stores the instruction context in a hot context cache at an address corresponding to the pending S-G DMA instruction. In response to the absence of the pending S-G DMA instruction, the system stores the instruction context in a cold context cache.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

Neural network large model efficient reasoning method based on multiple GPGPUs

The invention belongs to the technical field of artificial intelligence and high-performance computing, and particularly relates to a neural network large model efficient reasoning method based on multiple GPGPUs. The method aims to solve the problems of high communication overhead, non-uniform load, low resource utilization rate, high data transmission delay and the like among multiple processors. Dividing a calculation task into a plurality of sub-graphs through static analysis and mixed granularity partitioning of a model calculation graph; distributing the sub-graphs to the optimal GPGPU based on a weighted cost function in combination with heterogeneous resource perception and a dynamic mapping strategy; a global pipeline scheduling plan is constructed by using communication topology perception, and calculation and communication overlap are maximized; data are loaded in advance through a host side hierarchical caching and asynchronous prefetching mechanism, and transmission delay is hidden; multi-stream concurrent execution and event-based lightweight synchronization are adopted on each GPGPU, so that waiting overhead is reduced. According to the method, the reasoning delay can be remarkably reduced, the throughput and the hardware utilization rate are improved, and the method has good adaptivity and expandability.
Owner:BEIJING TOPMOO TECH

Memory architecture-oriented dual-precision general matrix multiplication optimization method and system

The invention belongs to the related technical field of high-performance computing, and provides a memory architecture-oriented dual-precision general matrix multiplication optimization method and system in order to solve the problems of limited computing power and access efficiency and the like in the prior art. Decomposing the matrix into a plurality of sub-matrix blocks according to the slave core array topology; the slave core receives the sub-matrix blocks issued by the master core, divides the sub-matrix blocks into small sub-matrix blocks based on a uniform blocking rule, loads the small sub-matrix blocks to an independent buffer area of a local data memory based on a DMA double-buffer protocol, divides the small sub-matrix blocks in the buffer area into SIMD vectors according to the SIMD unit characteristics of the slave core, and sends the SIMD vectors to the slave core; vectorization calculation and caching operation are alternately switched according to an iteration period through different independent buffer areas; and after all the slave cores finish calculation, the master core collects results written back to the master memory by the slave cores to obtain a final operation result, and double breakthrough of calculation power and memory access efficiency is realized.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Heterogeneous computing power cooperative scheduling system and method for mixed precision training

The invention discloses a heterogeneous computing power cooperative scheduling system and method for mixed precision training, and belongs to the technical field of artificial intelligence computing. The system comprises a computational graph analysis and operator portrait module which is used for analyzing and dividing a model computational graph and extracting operator features; the heterogeneous hardware capability sensing and matching module is used for managing performance files and real-time states of heterogeneous hardware in the cluster and matching optimal execution hardware for each calculation partition; and the data flow coordination and pipeline parallel controller is used for generating a global execution plan, managing cross-device data dependence and communication and calculating overlapping optimization execution efficiency through communication. According to the method, the problem of low scheduling efficiency of mixed precision training in a heterogeneous environment is solved, automatic and accurate mapping from a calculation task to heterogeneous hardware is realized, the training speed is remarkably improved, the training cost is reduced, and the overall resource utilization rate of a cluster is improved.
Owner:HANHOU (BEIJING) TECH CO LTD

Multi-dimensional data logic processing method based on artificial intelligence algorithm

The invention relates to the technical field of electric digital data processing, and discloses a multi-dimensional data logic processing method based on an artificial intelligence algorithm, which comprises the following steps: receiving a target discrete data packet, extracting metadata, and calculating to generate a dimension entropy feature vector representing data logic complexity; inputting the vector into a preset topological mapping model, and outputting an initial adjacent matrix; calling a feedback suppression mask matrix generated based on a historical operator utility state, and executing bitwise logic AND operation with the initial adjacent matrix to generate a corrected effective topological matrix; the matrix is analyzed, a logic operator function pointer is dynamically indexed in an instruction cache, and a directed acyclic execution linked list is constructed; according to the method, redundant logic nodes in AI prediction are definitely eliminated through a bit operation mask mechanism based on historical feedback, and deterministic convergence of processing delay and optimal matching of computing power resources are achieved.
Owner:SHENJIANG UNIVERSAL DATA INFORMATION CO LTD

Data processor, method, electronic device and storage medium

The invention provides a data processor and method, electronic equipment and a storage medium, the data processor comprises an instruction scheduler and a copy engine, the instruction scheduler comprises a first tensor access register, and the first tensor access register is used for storing sub-tensor description information transmitted externally; the replication engine comprises a first queue structure, a second tensor access register and an instruction decoder, the second tensor access register is used for storing sub-tensor description information received from the first queue structure, and the instruction decoder is used for analyzing an instruction received from the first queue structure and generating a control signal to drive data handling; wherein the instruction scheduler is configured to dynamically detect a change parameter item of the sub-tensor description information, and write the change parameter item into the first queue structure, so that the copy engine incrementally updates the sub-tensor description information in the second tensor access register based on the change parameter item, therefore, the data transmission redundancy can be reduced, and the data transmission efficiency is improved. And the instruction processing efficiency is improved.
Owner:SUZHOU YIZHU INTELLIGENT TECH CO LTD

Instruction speculation execution method and device of vector processor and storage medium

The invention discloses an instruction speculation execution method and device of a vector processor and a storage medium. The method comprises the following steps: in response to a vector configuration instruction identified at the front end of a vector processor pipeline, querying a vector prediction configuration table based on instruction information of the vector configuration instruction, and determining a prediction vector configuration parameter; updating a speculative vector state of a back end of the vector processor pipeline based on the predictive vector configuration parameter; performing a speculative process on a vector instruction following the vector configuration instruction based on the speculative vector state; when an actual execution result of the vector configuration instruction is obtained, determining an actual vector configuration parameter; comparing the actual vector configuration parameters with the prediction vector configuration parameters, and when the actual vector configuration parameters are consistent with the prediction vector configuration parameters, determining that speculation processing is effective and updating confidence information in the vector prediction configuration table; and if not, performing a flushing operation on the vector processor assembly line, and updating the vector prediction configuration table by using the actual vector configuration parameters. The processing efficiency of the vector processor assembly line can be improved.
Owner:SHANGHAI LINGRUI INTELLIGENT CORE COMPUTING TECHNOLOGY CO LTD

High-real-time interrupt management system and method based on RISC-V architecture

The invention relates to the technical field of integrated circuit design, computer system structures and embedded systems, in particular to a high-real-time interrupt management system and method based on an RISC-V. The system comprises an improved platform-level interrupt controller and an optimized processor internal interrupt processing unit. The system and the method are in tight coupling cooperative work through a special interruption field interface and a system bus, and the system further comprises a core local interrupter, an SRAM, an APB bus matrix and other components. The improved platform-level interrupt controller is responsible for sampling, gating, arbitration, flow control and direct memory access transmission of external interrupt, and a plurality of functional modules are arranged in the improved platform-level interrupt controller; an interrupt processing unit in the processor is responsible for hardware handshake, vector jump, nested control and tail biting mechanism implementation. Through hardware field management, multi-level priority arbitration, hardware nesting and tail biting mechanisms, delay and overhead caused by software intervention in a traditional architecture are eliminated, nanosecond response is achieved, the system throughput rate is increased, the software development threshold is lowered, and the method is suitable for strong real-time scenes.
Owner:FUDAN UNIVERSITY

Heterogeneous processor-oriented reciprocal calculation instruction sequence generation method

The invention discloses a reciprocal calculation instruction sequence generation method oriented to a heterogeneous processor, and belongs to the field of compilation optimization and code generation. Aiming at the problems of instruction redundancy, weak precision control, poor hardware adaptation and high manual dependence of an existing method in a heterogeneous environment, characteristics of a reciprocal instruction and an operand are accurately identified by linearly scanning heterogeneous object codes (including vectorization, scalar and complex instruction sequences); in combination with hardware characteristics of RISC / SIMD / VLIW / DSP and the like, a multi-round iteration precision improvement and temporary register optimization allocation strategy is adopted, differential generation logic is formulated, and a high-precision low-redundancy instruction sequence is generated. The method comprises linear code scanning classification, reciprocal instruction and operand identification, cross-architecture generation logic rule formulation, instruction sequence generation and legality verification. Full-process automation is achieved, manual intervention is reduced, the execution efficiency and precision of reciprocal calculation of the heterogeneous processor are improved, and the method is suitable for embedded systems, high-performance calculation and other scenes.
Owner:HUNAN UNIV OF SCI & TECH

Thread group scheduling method and device for GPU (Graphics Processing Unit), graphics processing unit and equipment

The invention discloses a thread group scheduling method and device for a GPU (Graphics Processing Unit), the GPU and equipment. The method comprises the following steps of: 1) receiving a scheduling request of a thread group; 2) task type priority scheduling; according to a thread group weight value set by a user, obtaining execution priorities of the vertex thread group and the fragment thread group in the current scheduling period; 3) instruction type priority scheduling; pre-analyzing to-be-executed instructions of the thread group, and determining a priority sequence of schedulable instruction types; and 4) priority scheduling of the thread groups: selecting the thread group with the highest priority for scheduling according to the specified task type priority and instruction type priority in combination with the thread group generation time. The invention provides a thread group three-level scheduling strategy so as to improve the instruction throughput rate and the key task response speed of the GPU under the complex load.
Owner:WUHAN LINGJIU MICROELECTRONICS CO LTD

Expert parallelism processing method and system of large language model based on MoE

The invention belongs to the field of machine learning, discloses an expert parallelism processing method and system for a large language model based on MoE, and realizes efficient parallel processing of the MoE model by dynamically distributing expert quantization bit width and sparse mode, predicting and prefetching to-be-activated expert parameters, grouping tokens to generate task queues and dynamically configuring hardware accelerators. Firstly, an importance score is calculated based on expert historical activation frequency, weight distribution and a model structure, so that quantization precision and a sparse proportion are adaptively allocated, and resource waste and precision loss of a unified strategy are avoided; secondly, predicting an expert to be activated by using a current layer hidden state and a historical activation sequence, loading parameters to a special cache in advance, reducing high-bandwidth memory access, and relieving bandwidth peak scrambling; moreover, tokens are grouped through a token-expert mapping table, a task queue is created, and dynamic configuration of a systolic array is combined, so that the problem of expert heterogeneity after compression is solved, and efficient and parallel hybrid precision matrix operation is ensured.
Owner:NINGBO ORIENTAL UNIVERSITY OF TECHNOLOGY +2

Automatic compiling vectorization method, terminal and medium

The invention provides a compiling automatic vectorization method, a terminal and a medium, and the method comprises the steps: extracting each seed instruction group in a program block, and constructing an instruction group queue based on each extracted seed instruction group; sequentially executing SIMD (Single Instruction Multiple Data) judgment on each instruction group in the instruction group queue, and executing a corresponding vectorization design according to an SIMD judgment result to obtain an initial vectorization scheme; performing redundant structure rewriting on the initial vectorization scheme to obtain a simplified vectorization scheme; respectively calculating a scalar total execution cost and a vector total execution cost corresponding to the simplified vectorization scheme by using a preset cost model; when it is detected that the scalar total execution cost is larger than the vector total execution cost, setting the simplified vectorization scheme as a final vectorization scheme, and conducting vectorization compiling on the program block based on the final vectorization scheme; according to the method, the compiling efficiency of program vectorization compiling can be effectively improved.
Owner:SHANGHAI ADVANCED RES INST CHINESE ACADEMY OF SCI

Instruction processing method and device, chip, equipment and storage medium

The embodiment of the invention discloses an instruction processing method and device, a chip, equipment and a storage medium. The method comprises the steps that a to-be-executed branch instruction and a target thread bundle used for executing the branch instruction are acquired; the target thread bundle comprises K threads, P effective threads exist in the K threads, and the branch instruction needs to be executed and processed in each effective thread; wherein K > = P > = 1, and K and P are integers; performing thread compression processing on the target thread bundle to obtain a sub-thread bundle; the sub-thread bundle comprises P effective threads of the target thread bundle; obtaining the number of available execution units in the target processor, and selecting corresponding effective threads for the execution units from the sub-thread bundles according to the number of the available execution units; and calling the execution unit to execute the branch instruction in the corresponding effective thread. The instruction processing efficiency of the target processor can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Branch prediction unit power consumption management method and device

The invention discloses a branch prediction unit power consumption management method and device, and the method comprises the steps: setting a multi-stage branch predictor and a low-power-consumption control logic coupled with the multi-stage branch predictor in a processor core micro-architecture, and enabling the multi-stage branch predictor to comprise three stages of predictors which are sequentially accessed in an assembly line, the hardware resource consumption and the prediction accuracy of each predictor are gradually increased from front to back; monitoring the prediction condition of each predictor in real time through low-power-consumption control logic, and closing all predictors behind the first predictor meeting the prediction requirement within a first preset time according to a front-to-back sequence; branch jump addresses are divided into high-order address fields and low-order address fields by BTBs in the multiple levels of predictors except the first-level predictor to be stored in different storage blocks, and label registers used for taking the high-order addresses as label indexes are arranged. The power consumption unit of each predictor can be finely controlled, the energy consumption is greatly reduced, system-level cooperation is not needed, and the universality is high.
Owner:NANJING YINGQI INTELLIGENT TECH CO LTD

Time axis phase unwrapping method and system based on FPGA

The invention belongs to the technical field of distributed optical fiber sensing (DAS) system signals, and discloses a time axis phase unwrapping method and system based on an FPGA (Field Programmable Gate Array), and the method comprises the steps of phase data acquisition, stable group formation, phase progressive unwrapping, assembly line and cache cooperative processing and quality measurement closed-loop correction. According to the system, real-time phase extraction of interference signals is achieved through the FPGA demodulation unit, phase change is dynamically unwrapped through the reliable period screening unit and the phase progressive unwrapping unit, a self-adaptive unwrapping mechanism is formed in combination with anomaly judgment and threshold updating logic, and through collaborative design of an annular cache and a three-level assembly line comparator, the interference signal is obtained. According to the method, parallel transmission and time sequence synchronization of data are achieved, the precision and stability of phase unwrapping can be effectively improved in a complex noise environment, hardware delay and resource occupation are remarkably reduced, and the method is suitable for real-time signal unwrapping processing of a high-precision distributed optical fiber sensing system.
Owner:BEIJING ZHONGTUO XINYUAN TECH CO LTD

Artificial intelligence chip, parallel method for vector and scalar execution pipeline, computing device, medium and program product

The invention relates to an artificial intelligence chip, a method for parallel vector and scalar execution assembly lines, a computing device, a medium and a program product. The artificial intelligence chip comprises an execution unit, the execution unit is configured with a vector execution assembly line and a scalar execution assembly line, and the scalar execution assembly line at least comprises a scalar instruction decoding unit which is configured to at least obtain an operand type, address information and scalar operation control information of a scalar instruction; a scalar instruction operand acquisition unit configured to acquire an operand source of a scalar instruction; and a scalar instruction operation unit configured to execute scalar calculation at least based on an operand type, an operand source and scalar operation control information of the scalar instruction, and write a calculation result to the scalar register group included in the execution unit. According to the method, the utilization rate and the actual computing power of hardware resources of the execution unit of the artificial intelligence chip can be remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD

RISC-V simulation resource dynamic generation method and system

The invention belongs to the technical field of integrated circuit simulation verification, particularly relates to an RISC-V simulation resource dynamic generation method and system, and solves the problems of instruction consistency and variable-length instruction truncation through a metadata double-table structure and a cross-boundary instruction splicing mechanism. Virtualized two-stage translation support and abnormal injection are realized through a recursive multiple hit detection and dynamic attribute bit modification mechanism. According to the method, complete storage is replaced with lightweight metadata, memory occupation is remarkably reduced, the problem that memory occupation is linearly increased along with time is solved, simulation efficiency and consistency are improved, and the method is suitable for full-system verification of a high-performance RISC-V processor.
Owner:SHANDONG UNIV

Automatic illegal character cleaning system

ActiveCN121807413AResource allocationRegister arrangementsTopology mappingSpeculative execution
The invention relates to the technical field of computer data processing and network security, in particular to an illegal character automatic cleaning system, which comprises a rule compiling module for monitoring rule change, performing semantic fusion and topological mapping on a rule set, constructing a deterministic finite automaton and mapping the deterministic finite automaton into a state transition table; the state switching module is used for constructing a double-buffer context and operating lock-free switching through an atomic pointer to realize hot updating; the speculation execution module is used for carrying out vectorization pre-scanning based on a state transition table by utilizing single-instruction multi-data stream parallel loading, identifying a walk path, falling into a safe state for releasing and falling into a trap state for triggering external verification; the self-adaptive feedback module is used for counting trap state triggering frequency, generating a rule allergy report and dynamically adjusting the size of a read fragment; according to the method, the contradiction between rule flexibility and execution efficiency is solved, and high-performance cleaning based on speculative execution is realized.
Owner:北京啄木鸟云健康科技有限公司

Task processing device, task processing method, electronic equipment and storage medium

The embodiment of the invention provides a task processing device, a task processing method, electronic equipment and a storage medium. The task processing device comprises a calculation scheduling core particle and a function core particle which are integrated in a single package, the calculation scheduling core particle comprises a hardware subsystem and a calculation core, and the hardware subsystem is configured to receive and analyze task information to obtain at least one task. And assigning at least one task to at least one of the functional core and the computing core according to the task type; the computing core is configured to execute computing operations according to tasks allocated by the hardware subsystem, and the functional core particles comprise at least one of a quantum computing management core particle and a ray tracing core particle. According to the task processing device, through a unified task scheduling distribution mechanism, the adaptive capacity of the task processing device in an end-side mixed task scene is remarkably improved, so that a single computing chip can efficiently support various heterogeneous workloads.
Owner:SHANGHAI BIREN TECH CO LTD

Hybrid branch predictor capable of dynamically selecting different prediction methods and application

The invention discloses a hybrid branch predictor capable of dynamically selecting different prediction methods and application, and relates to the technical field of high-performance processor design, the hybrid branch predictor mainly comprises a branch direction predictor and a branch target predictor; the branch direction predictor is used for predicting whether a branch instruction jumps or not; and the branch target predictor is used for predicting a specific jump target address when the branch direction is predicted to be jump. By implementing the hybrid branch predictor capable of dynamically selecting different prediction methods and the application provided by the invention, the branch prediction accuracy can be improved, and the hardware resource overhead can be reduced.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719 +1

Non-blocking vector instruction dispatch with micro-operations

A processor core is coupled to a memory hierarchy. The processor core is configured to execute vector instructions, scalar instructions, and micro-operations. A dispatch unit within the processor core receives a vector memory operation. The dispatch unit sends the vector memory operation to a first vector input queue of multiple vector input queues. The sending is based on the memory addressing mode. A micro-operation sequencer splits the vector memory operation into one or more memory micro-operations, which includes forwarding each micro-operation within the one or more micro-operations to a first memory queue within multiple memory queues. A memory operation is then issued to a load-store unit within the processor core. The issuing includes selecting, from the multiple memory queues, the memory operation. The vector memory operation comprises either a vector load operation or a vector store operation.
Owner:AKEANA INC

Tokenized data validation in artificial intelligence operational pipelines

An example operation may include one or more of executing an artificial intelligence (AI) pipeline including an AI model via a software application, storing a token data model of the AI model via a storage of the software application, receiving input data via the AI pipeline of the software application, converting the input data into tokens via execution of a tokenizer within the AI pipeline on the input data, determining whether the tokenizer is valid based on a comparison of the tokens and the token data model of the AI model, and continuing execution of the AI pipeline of the software application based on whether the tokenizer is valid.
Owner:THE TORONTO DOMINION BANK

Data processing apparatus, processor, and data processing method

The embodiment of the invention provides a data processing device, a processor and a data processing method. The data processing apparatus includes an instruction distribution unit, an instruction information recording unit, an instruction processing unit, and a write-back control unit, the instruction distribution unit being configured to split a vector instruction into a plurality of microinstructions, and to distribute the plurality of microinstructions to the instruction processing unit; the instruction information recording unit is configured to record instruction state information corresponding to a plurality of microinstructions respectively; the instruction processing unit is configured to receive a plurality of micro-instructions, process each of the plurality of micro-instructions to obtain a corresponding execution result, and feed back a processing progress to the instruction information recording unit; and the write-back control unit is configured to select at least one to-be-written-back microinstruction from the instruction information recording unit to perform write-back operation based on the instruction state information, and the write-back operation is used for writing an execution result of the to-be-written-back microinstruction back to the register. The device can realize out-of-order write-back of the instruction and improve the execution efficiency of the instruction.
Owner:XIAN YISIWEI COMPUTING TECH CO LTD +1

Pipeline systems and methods for use in data analytics platforms

A data analytics system including an append-only first data store accessible to multiple clients and a second data store is disclosed. The data analytics system can be configurable to, in response to receiving first instructions from a first target system of a first client, the first target system separate from the data analytics system, create a first pipeline between the append-only first data store and the second data store. The first pipeline can be configured according to the first instructions to generate a client-specific data object and store the client-specific data object in the second data store. The data analytics system can be configurable to tear down the first pipeline upon completion of storing the client-specific data object in the second data store.
Owner:FIDELITY INFORMATION SERVICES LLC

Driving task arrangement generation method and device, vehicle, medium and program product

The invention relates to the technical field of intelligent network connection, in particular to a driving task arrangement generation method and device, a vehicle, a medium and a program product, and the method comprises the steps: obtaining task information and resource information of at least two tasks in driving task arrangement; determining at least one optimization target of task arrangement based on the task information and the resource information; and inputting the task information and the resource information into a pre-trained large language model to output initial task scheduling data meeting at least one optimization target, and verifying the initial task scheduling data until task scheduling data meeting a preset verification condition is obtained. Therefore, the problems that in the related technology, manual analysis is prone to errors, the management efficiency is low, and the actual use scene cannot be adapted are solved; model processing depends on a fixed rule, is lack of dynamic scheduling capability, increases response delay of a system, causes task timeout, and cannot guarantee reliability and safety of a driving task.
Owner:BEIJING AUTOMOBILE RES GENERAL INST

Enhanced ultra low-latency, high-throughput matching engine for electronic trading systems

A high-speed matching-engine architecture is disclosed that sustains deterministic sub-microsecond latency while processing more than 10 million order messages per second per core on commodity multi-core processors. Orders reside in cache-aligned Data Holder Nodes whose occupancy and price-level boundaries are tracked with constant-time bitmask operations, eliminating pointer-chasing penalties. Per-core huge-page pools, SIMD copy kernels, and lock-free, cache-line-aligned queues further minimize TLB misses and coherence overheads. Overflow is handled by Push Back / Push Forward cascades that relocate the least- or most-prioritized orders between adjoining nodes without violating price-time priority. Node capacities vary monotonically with book depth and are re-tuned online by a lightweight machine-learning controller that maximizes cache-hit probability under changing market micro-structure. The design tightens spreads, raises match-rate revenue, and complies with stringent regulatory latency caps using standard x86-64, Arm, or other architectures.
Owner:YOON JIN SEOK

Return address backup method and device, computer equipment and storage medium

The invention provides a return address backup method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining a return address stack (RAS) corresponding to a first pipeline processing unit, and configuring RASs corresponding to at least one second pipeline processing unit; obtaining a to-be-processed instruction sequence, and performing pipeline processing on the to-be-processed instruction sequence; in the process of performing assembly line processing on the instruction sequence to be processed, responding to assembly line flushing operation initiated by a target assembly line processing unit, copying an RAS corresponding to the target assembly line processing unit into an RAS corresponding to the first assembly line processing unit, the target assembly line processing unit is one of the at least one second assembly line processing unit. Based on the scheme of the invention, the RAS backup can be realized by setting the RAS in the second assembly line processing unit, so that the influence of assembly line scouring on the RAS is eliminated, and the accuracy of the RAS for predicting the return address is improved.
Owner:芯来智融半导体科技(上海)股份有限公司

Circuitry and methods for a conditional fence instruction

Circuitry and methods for implementing conditional fence instructions are described. In certain examples, a hardware processor (e.g., core) includes a branch predictor to predict one of a taken path and a not taken path for a conditional branch instruction; decoder circuitry to decode an instruction into a decoded instruction, the instruction comprising a field that indicates a condition to be set by execution of another instruction, and an opcode that indicates execution circuitry is to, in response to the condition being satisfied, implement an execution fence to delay execution of the instruction until prior instructions in program order execute and delay execution of instructions after the instruction in program order until the instruction executes; and the execution circuitry to execute the decoded instruction according to the opcode.
Owner:INTEL CORP

Cooperative parallel memory allocation

Apparatuses, systems, and techniques to perform multi-threaded memory allocation in parallel by one or more software programs being performed on a parallel processing unit (PPU), such as a graphics processing unit (GPU), or any other processing unit capable of supporting multi-threaded software execution. In at least one embodiment, one or more software programs expressed in part by code using an application programming interface for parallel computing, such as CUDA, perform allocation, search, and deallocation of memory efficiently and in parallel on a GPU.
Owner:NVIDIA CORP