Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

366results about "Instruction analysis" patented technology

Instruction speculation execution method and device of vector processor and storage medium

The invention discloses an instruction speculation execution method and device of a vector processor and a storage medium. The method comprises the following steps: in response to a vector configuration instruction identified at the front end of a vector processor pipeline, querying a vector prediction configuration table based on instruction information of the vector configuration instruction, and determining a prediction vector configuration parameter; updating a speculative vector state of a back end of the vector processor pipeline based on the predictive vector configuration parameter; performing a speculative process on a vector instruction following the vector configuration instruction based on the speculative vector state; when an actual execution result of the vector configuration instruction is obtained, determining an actual vector configuration parameter; comparing the actual vector configuration parameters with the prediction vector configuration parameters, and when the actual vector configuration parameters are consistent with the prediction vector configuration parameters, determining that speculation processing is effective and updating confidence information in the vector prediction configuration table; and if not, performing a flushing operation on the vector processor assembly line, and updating the vector prediction configuration table by using the actual vector configuration parameters. The processing efficiency of the vector processor assembly line can be improved.
Owner:SHANGHAI LINGRUI INTELLIGENT CORE COMPUTING TECHNOLOGY CO LTD

Thread group scheduling method and device for GPU (Graphics Processing Unit), graphics processing unit and equipment

The invention discloses a thread group scheduling method and device for a GPU (Graphics Processing Unit), the GPU and equipment. The method comprises the following steps of: 1) receiving a scheduling request of a thread group; 2) task type priority scheduling; according to a thread group weight value set by a user, obtaining execution priorities of the vertex thread group and the fragment thread group in the current scheduling period; 3) instruction type priority scheduling; pre-analyzing to-be-executed instructions of the thread group, and determining a priority sequence of schedulable instruction types; and 4) priority scheduling of the thread groups: selecting the thread group with the highest priority for scheduling according to the specified task type priority and instruction type priority in combination with the thread group generation time. The invention provides a thread group three-level scheduling strategy so as to improve the instruction throughput rate and the key task response speed of the GPU under the complex load.
Owner:WUHAN LINGJIU MICROELECTRONICS CO LTD

System and method for determining communication pathway

PendingUS20260081014A1Medical communicationDrug and medicationsIntelligent NetworkPrior authorization
Systems and methods for providing an intelligent network to onboard patients to specialty medications are described herein. A server system for managing communication is configured to receive from a user at a healthcare practice, an electronic request to obtain a medical prior authorization (medPA) for a specialty drug to administer for a patient of the healthcare practice; select, using a rules service, a communication pathway to submit the medPA to a healthcare payer from a plurality of communication pathways; and initiate a connection with the healthcare payer using the selected communication pathway.
Owner:CAREMETX LLC

RISC-V-oriented eBPF loop vectorization compiling method and device

The invention provides an RISC-V-oriented eBPF cycle vectorization compiling method and device, and the method comprises the steps: generating an RVV vector machine code, executing the parallel operation when a condition is satisfied, achieving the parallel processing of a single instruction and multiple data, carrying out the serial conversion of the time dimension of eBPF cycle into the parallel spatial dimension of an RVV vector, remarkably improving the processing throughput, and improving the compiling efficiency. And the task response time delay and the CPU occupancy rate are reduced. According to the RVV vector execution mode, a large number of repeated loop control instructions and memory read-write instructions can be reduced, the register overflow problem caused by eBPF scalar register resource shortage is relieved, the cache pressure and the branch prediction failure rate are reduced, and the energy efficiency ratio is increased. Meanwhile, an RVV vector machine code and a scalar fallback code are generated, a parallel priority and scalar bottom double execution path is formed, and it is ensured that under the conditions that a hardware environment is not supported, the data size does not reach a threshold value, vectorization verification fails or operation is abnormal, the scalar fallback code can be seamlessly switched to be executed.
Owner:BEIJING VCORE TECH CO LTD

Multi-account asynchronous bookkeeping data consistency collaboration system based on instruction driving

PendingCN121614483ADatabase updatingResource allocationInformation systems engineeringLogisim
The invention relates to the field of computer software and information system engineering, in particular to a multi-account asynchronous bookkeeping data consistency collaboration system based on instruction driving, which comprises the following steps of: packaging a request into a standard instruction and routing according to an account to ensure ordered processing of the request and improve the processing efficiency and expandability of the system; repeated requests are avoided through idempotent verification, business is mapped into atomic operation through rule analysis, and the accuracy of business logic and data consistency are guaranteed; atomic operation is executed in a single transaction, asynchronous retry is performed during conflict, the atomicity of cross-account operation is ensured, and the concurrent processing capacity is improved; by buffering high-frequency account change and batch aggregation updating, write conflicts are reduced, the system performance is stabilized, and the throughput is improved; through account checking based on daily switching snapshots, pipeline playback and reverse account flushing are supported, and a rapid repair capability is provided; through dynamic configuration and versioned management of the bookkeeping rule, visual monitoring is provided, the operation and maintenance complexity is reduced, and the system flexibility is enhanced.
Owner:HANGZHOU YIZI INFORMATION TECHNOLOGY CO LTD

Instruction processing method and device, computer equipment and storage medium

Embodiments of the invention provide an instruction processing method and apparatus, a computer device and a storage medium. The method comprises the steps of obtaining a setting instruction and a vector calculation instruction from a preset instruction local memory or an instruction buffer; updating a preset shadow register according to the setting instruction to obtain a first update value; performing data recovery on the first update value by using a preset real register to obtain a second update value; and executing the vector calculation instruction according to the second update value. According to the method, the preset shadow register is used for speculative execution according to the instruction of the assembly line, an update value is obtained in advance, and when the instruction needing to be scoured is executed, the wrong update value is backed up and recovered by using the preset real register, so that the register values of the shadow register and the real register are kept consistent; therefore, the performance and the efficiency when the vector calculation instruction is executed are improved.
Owner:芯来智融半导体科技(上海)股份有限公司

Command processing circuit, command processing method, and PIM device

The invention provides a command processing circuit, a command processing method and a PIM device. The command processing circuit comprises a PIM command identification module which is electrically connected with a command address bus of the PIM device and is configured to receive a command signal through the command address bus and judge the type of the currently received command signal; if it is judged that the currently received command signal is a DRAM conventional command, the command signal is sent to a conventional command processing module; if it is judged that the currently received command signal is a PIM related command, the command signal is sent to a PIM command processing module; the PIM command processing module is electrically connected with the PIM command identification module and is configured to receive a PIM related command, decode the PIM related command and perform corresponding operation according to a decoding result; and the conventional command processing module is electrically connected with the PIM command identification module and is configured to receive the DRAM conventional command and decode the DRAM conventional command. The method is at least beneficial to avoiding mode switching and improving system performance.
Owner:RUILI INTEGRATED CIRCUIT CO LTD

Supporting 8-bit floating point format operands in a computing architecture

An apparatus to facilitate supporting 8-bit floating point format operands in a computing architecture is disclosed. The apparatus includes a processor comprising: a decoder to decode an instruction fetched for execution into a decoded instruction, wherein the decoded instruction is a matrix instruction that operates on 8-bit floating point operands to cause the processor to perform a parallel dot product operation; a controller to schedule the decoded instruction and provide input data for the 8-bit floating point operands in accordance with an 8-bit floating data format indicated by the decoded instruction; and systolic dot product circuitry to execute the decoded instruction using systolic layers, each systolic layer comprises one or more sets of interconnected multipliers, shifters, and adder, each set of multipliers, shifters, and adders to generate a dot product of the 8-bit floating point operands.
Owner:INTEL CORP

Circuitry and methods for capability informed prefetches

Systems, methods, and apparatuses for implementing capability informed prefetches are described. In certain examples, a hardware processor comprises an execution circuit to execute an instruction that generates a memory access request for an element in memory via a first capability; a capability management circuit to check the first capability for the memory access request, the first capability comprising an address field of the element in the memory, a validity field, and a bounds field that is to indicate a lower bound and an upper bound of a first object to which the first capability authorizes access; a cache; and a prefetch circuit to: prefetch an additional element from the memory to the cache, determine if the additional element is a second capability comprising an address field of a second element in the memory, a validity field, and a bounds field that is to indicate a lower bound and an upper bound of a second object to which the second capability authorizes access, and prefetch the second element from the memory to the cache based on the additional element being the second capability.
Owner:INTEL CORP

Instruction control method, data caching method, and related products

This disclosure discloses an instruction control method, instruction control apparatus, processor, chip, and board card. The processor can be included as a computing apparatus in a combination processing apparatus, which may also include interface devices and other processing apparatus. The computing apparatus interacts with other processing apparatuses to jointly complete computing operations specified by the user. The combined processing apparatus may also include a storage apparatus, which is connected to the computing apparatus and other processing apparatuses respectively, for storing data of the computing apparatus and other processing apparatuses. The disclosed solution provides an instruction control method that can enhance instruction level parallelism and improve processing efficiency.
Owner:SHANGHAI CAMBRICON INFORMATION TECH CO LTD

Loosely coupled slice target file data

The system may determine that two instructions can be combined based on the processing power of the processor and the sizes of the instructions, fuse the two instructions into a pair, map the two instructions using a single register tag, write the register tag to a mapper with a bit indicating that the register tag is for a first instruction of the two instructions, write the register tag to the mapper with a bit indicating that the register tag is for a second instruction of the two instructions, write the fused instruction pair to an issue queue, issue the fused instruction pair to a vector scalar translation unit (VSU), and execute the two instructions.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Providing physical register (PR) swap memory renaming in processor-based devices

Providing physical register (PR) swap memory renaming in processor-based devices is disclosed herein. In some exemplary aspects, a processor provides an instruction processing circuit comprising a scheduling stage circuit and an execution stage circuit. The scheduling stage circuit comprises a reservation station circuit, while the execution stage circuit comprises a PR swap table storing a plurality of PR swap table entries. The scheduling stage circuit issues a first instruction that is associated with a store dependency ID. The execution stage circuit, in response to the issuing of the first instruction, identifies a PR swap table entry among the plurality of PR swap table entries corresponding to the store dependency ID, retrieves a load dependency ID of the PR swap table entry, and broadcasts the load dependency ID to the reservation station circuit to wake a second instruction that is associated with the load dependency ID.
Owner:QUALCOMM INC

A method for dynamic voltage and frequency regulation of a RISC-V processor core

This invention relates to a dynamic voltage and frequency adjustment method for a RISC-V processor core. The method includes: monitoring and recording the critical path time of memory access requests for each missing state processing register to update the global critical path counter; monitoring prefetch instructions and extracting their delay parameters; collecting cache access events based on a time window to calculate bandwidth utilization and dynamically determining a bandwidth threshold according to preset performance parameters; and correcting the global critical path counter value by combining the prefetch delay and the bandwidth threshold, thereby triggering a dynamic voltage and frequency adjustment instruction. This method achieves fine-grained adaptive DVFS for the RISC-V processor core under complex loads, significantly improving energy efficiency and performance stability.
Owner:CHAORUI TECH (CHANGSHA) CO LTD

Systems and methods to provide instructions to coprocessors

A method (400) may include a processor core fetching (402) a packet of machine code instructions and then determining (404) whether a first machine code instruction of the packet corresponds to a coprocessor operation. In response to determining that the first machine code instruction corresponds to a coprocessor operation, the processor core may treat the other machine code instructions of the packet as no operations (NOOPs) and transmit (406) the machine code instructions of the packet to a coprocessor. The coprocessor may then decode and execute the machine code instructions. The method may further include the processor core keeping responsibility for load and store operations and, in the case of coprocessor operations, using (408, 410) registers of the coprocessor as source and destination for load and store operations.
Owner:TEXAS INSTRUMENTS INC

Speculation barrier

An apparatus comprises processing circuitry configured to perform data processing and instruction decoding circuitry configured to decode instructions to control the processing circuitry to perform the data processing. The instruction decoding circuitry is responsive to a speculation barrier variant of a conditional instruction having an instruction outcome depending on a condition, to control the processing circuitry to: while the condition is not yet resolved, impose a speculation barrier requirement to restrict speculative handling of a subsequent operation appearing in program order after the speculation barrier variant of the conditional instruction, and in response to resolution of the condition, relax the speculation barrier requirement imposed by the speculation barrier variant of the conditional instruction.
Owner:ARM LTD

Data structure processing

An apparatus includes an instruction decoder and a processing circuitry system. In response to a data structure processing instruction specifying at least one input data structure identifier and an output data structure identifier, the instruction decoder controls the processing circuitry system to perform processing operations on at least one input data structure to produce an output data structure. Each input / output data structure includes a data arrangement corresponding to multiple memory addresses. The apparatus includes two or more sets of one or more data structure metadata buffers, each set associated with a corresponding data structure identifier and designated to store metadata indicating the memory address of the data structure identified by the corresponding data structure identifier.
Owner:ARM LTD

Data execution method and device, storage medium and product

The invention relates to the technical field of data processing, in particular to a data execution method and device, a storage medium and a product. The method comprises the following steps: receiving a data execution request sent by an upper computer; after instruction skipping is carried out according to data execution data contained in the data execution request, a control core corresponding to the execution request is determined; the data execution request is fed back to the control core, so that the control core executes a data execution command contained in the data execution request; according to the method and the device, each control core can execute different data processing tasks in a segmental manner, the data response capability of multiple data cores is improved, full resource utilization of each control core in the surgical robot is ensured, and the data processing efficiency of the surgical robot is improved.
Owner:CORE MOTION MEDICAL ROBOT (SHENZHEN) CO LTD

Systems, methods, and apparatuses for tile transpose

Embodiments detailed herein relate to matrix operations. In particular, support for a matrix transpose instruction is detailed. In some embodiments, decode circuitry to decode an instruction having fields for an opcode, a source matrix operand identifier, and a destination matrix operand identifier; and execution circuitry to execute the decoded instruction to transpose each row of elements of the identified source matrix operand into a corresponding column of the identified destination matrix operand are detailed.
Owner:INTEL CORP

A coprocessor-based method and system for taint propagation

The application discloses a kind of based on coprocessor's taint propagation method and system.The steps of the method include:1) pre-analysis instruction semantics, construct the taint propagation mask table of instruction;2) by coprocessor, intervene CPU instruction pipeline decoding, write-back process;3) in decoding stage, the opcode and operand of current instruction are acquired in real time quickly, and the taint propagation mask corresponding to instruction is found in taint propagation mask table by coprocessor;4) in write-back stage, the mask table corresponding to current instruction is searched in real time quickly, and taint propagation calculation is implemented.The application can analyze the semantics of instruction, calculate taint propagation process by using the characteristics of hardware fast and accurate when CPU instruction is executed by configuring the input function return value of monitoring.
Owner:INST OF SOFTWARE - CHINESE ACAD OF SCI

Accelerator for sparse-dense matrix multiplication

Disclosed embodiments relate to an accelerator for sparse-dense matrix instructions. In one example, a processor to execute a sparse-dense matrix multiplication instruction, includes fetch circuitry to fetch the sparse-dense matrix multiplication instruction having fields to specify an opcode, a dense output matrix, a dense source matrix, and a sparse source matrix having a sparsity of non-zero elements, the sparsity being less than one, decode circuitry to decode the fetched sparse-dense matrix multiplication instruction, execution circuitry to execute the decoded sparse-dense matrix multiplication instruction to, for each non-zero element at row M and column K of the specified sparse source matrix generate a product of the non-zero element and each corresponding dense element at row K and column N of the specified dense source matrix, and generate an accumulated sum of each generated product and a previous value of a corresponding output element at row M and column N of the specified dense output matrix.
Owner:INTEL CORP

Apparatuses, methods, and systems for instructions for structured-sparse tile matrix fma

Systems, methods, and apparatuses relating sparsity based FMA. In some examples, an instance of a single FMA instruction has one or more fields for an opcode, one or more fields to identify a source / destination matrix operand, one or more fields to identify a first plurality of source matrix operands, one or more fields to identify a second plurality of matrix operands, wherein the opcode is to indicate that execution circuitry is to select a proper subset of FP8 data elements from the first plurality of source matrix operands based on sparsity controls from a first matrix operand of the second plurality of matrix operands and perform a FMA.
Owner:INTEL CORP

On-device neural processing unit with heterogeneous cores for speculative decoding

According to the present disclosure, a device is provided. The device includes a first memory of a first capacity configured to store a first generative neural network model comprising a first parameters, and a first neural processing unit configured to generate a response corresponding to an input query utilizing the first generative neural network model stored in the first memory, and wherein the first neural processing unit may be configured to store a first execution code of the first generative neural network model compiled to process speculative decoding.
Owner:DEEPX CO LTD

Increasing data-dependent memory prefetcher accuracy using memory metadata

PendingEP4718243A1Instruction analysis
Data-dependent (memory) prefetcher support is described. In some examples, the data-dependent (memory) prefetcher utilizes metadata to predict if a data word is a valid pointer before doing a prefetch. Data words that are not deemed to be valid pointers are not used to prefetch. The metadata may be linear or physical depending on the example.
Owner:INTEL CORP

Apparatus and method for number-theoretical transformation instructions

The invention relates to an apparatus and a method for number-theoretical transformation instructions. Instructions are executed for performing a number-theory transform (NTT). For example, one embodiment includes: a decode circuit to decode the instruction; one or more packed data registers to store a first, second, and third plurality of source packed data elements and corresponding result packed data elements; and execution circuitry to execute the NTT using a butterfly operation including Montgomery multiplication of a source packed data element of the second plurality of source packed data elements and a corresponding source packed data element of the third plurality of source packed data elements using a modulus value or an inverse of the modulus value to generate a product. The product is added to a corresponding source packed data element of the first plurality of source packed data elements to obtain a first result, and the product is subtracted from the corresponding source packed data element of the first plurality of source packed data elements to obtain a second result.
Owner:INTEL CORP

Tensor data transformation method, tensor processing unit, processor, system on chip, and computing device

Provided in the embodiments of the present disclosure are a tensor data transformation method, a tensor processing unit, a processor, a system on a chip, and a computing device. The tensor processing unit comprises: a register array unit, which comprises at least a source address register, a destination address register, and a source tensor parameter register; an instruction decoding unit, which parses an acquired transformation operation instruction, so as to obtain a transformation operation type indicated by an operation type field and obtain respective register serial numbers of the source address register, the destination address register, and the source tensor parameter register; and an instruction execution unit, which accesses the source address register, the destination address register, and the source tensor parameter register on the basis of their respective register serial numbers, reads a source memory address field, a destination memory address field, and a source tensor parameter, reads a source tensor on the basis of the source memory address field, performs, on the source tensor, a transformation operation indicated by the transformation operation type, and writes, on the basis of the source tensor parameter, a destination tensor obtained after the transformation operation into the destination memory address field.
Owner:ALIBABA DAMO (HANGZHOU) TECH CO LTD

Using memory metadata to improve accuracy of data dependent memory prefetcher

The invention relates to using memory metadata to improve accuracy of a data dependent memory prefetcher. Data dependent (memory) prefetcher support is described. In some examples, a data dependent (memory) prefetcher utilizes metadata to predict whether a data word is a valid pointer prior to prefetch. Data words that are not regarded as valid pointers are not used for prefetch. The metadata may be linear or physical, depending on the example.
Owner:INTEL CORP

Command processing circuit, command processing method, and PIM apparatus

Provided in the present disclosure are a command processing circuit, a command processing method, and a PIM apparatus. The command processing circuit comprises: a PIM command identification module, which is electrically connected to a command address bus of a PIM apparatus and is configured to receive a command signal by means of the command address bus, determine the type of the currently received command signal, send, when it is determined that the currently received command signal is a conventional DRAM command, the command signal to a conventional command processing module, and send, when it is determined that the currently received command signal is a PIM-related command, the command signal to a PIM command processing module; the PIM command processing module, which is electrically connected to the PIM command identification module and is configured to receive the PIM-related command, decode the PIM-related command, and perform a corresponding operation on the basis of a decoding result; and the conventional command processing module, which is electrically connected to the PIM command identification module and is configured to receive the conventional DRAM command and decode the conventional DRAM command. The present disclosure is at least conducive to avoiding mode switching and improving system performance.
Owner:RUILI INTEGRATED CIRCUIT CO LTD