Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

645results about "Instruction analysis" patented technology

Matrix multiplication and accumulation operation unit and operation method, hardware accelerator and electronic equipment

The embodiment of the invention provides a matrix multiplication and accumulation operation unit and method, a hardware accelerator and electronic equipment, and the matrix multiplication and accumulation operation unit comprises a data loading storage engine, a tensor register file and a matrix multiplication engine. The data loading and storage engine is used for loading data of a plurality of matrixes to be subjected to matrix multiplication and accumulation calculation; the tensor register file is used for storing data of a plurality of matrixes acquired from the data loading storage engine; the tensor register file comprises at least three tensor register groups, each tensor register group comprises a plurality of tensor registers, and different tensor register groups are used for storing data of different matrixes in the plurality of matrixes; and the matrix multiplication engine is used for carrying out matrix multiplication accumulation calculation based on the data of the plurality of matrixes stored in the tensor register file. According to the embodiment of the invention, more efficient MMA calculation is realized under the conditions of low cost, low power consumption and less occupied space.
Owner:ALIBABA (CHINA) CO LTD

Technique for controlling stashing of data

An apparatus has decoder circuitry within a first processing element to decode instructions, in order to respond to a sequence of instructions by generating control signals. Processing circuitry within the first processing element is responsive to the control signals to perform operations defined by the sequence of instructions. The decoder circuitry is responsive to a stashee hint instruction in the sequence of instructions to issue control signals to cause the processing circuitry to issue a stash interest request via an interface used to couple the first processing element to interconnect circuitry. The stash interest request is arranged to provide a given memory address indication determined from the stashee hint instruction and to trigger stashing control circuitry accessible via the interconnect circuitry to cause updated data associated with the given memory address indication to be made available for stashing in an associated storage structure of the first processing element.
Owner:ARM LTD

Instruction speculation execution method and device of vector processor and storage medium

The invention discloses an instruction speculation execution method and device of a vector processor and a storage medium. The method comprises the following steps: in response to a vector configuration instruction identified at the front end of a vector processor pipeline, querying a vector prediction configuration table based on instruction information of the vector configuration instruction, and determining a prediction vector configuration parameter; updating a speculative vector state of a back end of the vector processor pipeline based on the predictive vector configuration parameter; performing a speculative process on a vector instruction following the vector configuration instruction based on the speculative vector state; when an actual execution result of the vector configuration instruction is obtained, determining an actual vector configuration parameter; comparing the actual vector configuration parameters with the prediction vector configuration parameters, and when the actual vector configuration parameters are consistent with the prediction vector configuration parameters, determining that speculation processing is effective and updating confidence information in the vector prediction configuration table; and if not, performing a flushing operation on the vector processor assembly line, and updating the vector prediction configuration table by using the actual vector configuration parameters. The processing efficiency of the vector processor assembly line can be improved.
Owner:SHANGHAI LINGRUI INTELLIGENT CORE COMPUTING TECHNOLOGY CO LTD

Matrix multiply accumulation operation unit and operation method, and hardware accelerator and electronic device

Provided in the embodiments of the present disclosure are a matrix multiply accumulation (MMA) operation unit and operation method, and a hardware accelerator and an electronic device. The MMA operation unit comprises a data load / store engine, a tensor register file and a matrix multiplication engine, wherein the data load / store engine is used for loading data of a plurality of matrixes to be subjected to MMA computation; the tensor register file is used for storing the data of the plurality of matrixes that is acquired from the data load / store engine; the tensor register file comprises at least three tensor register groups, each tensor register group comprising a plurality of tensor registers, and different tensor register groups being used for storing data of different matrixes among the plurality of matrixes; and the matrix multiplication engine is used for performing MMA computation on the basis of the data of the plurality of matrixes that is stored in the tensor register file. By means of the embodiments of the present disclosure, more efficient MMA computation is realized with low cost, low power consumption and a smaller footprint.
Owner:ALIBABA (CHINA) CO LTD

Instruction checking method and device based on semantic matching, terminal equipment and storage medium

The invention discloses an instruction checking method and device based on semantic matching, terminal equipment and a storage medium, and belongs to the technical field of instruction checking. The method comprises the steps that a current scheduling instruction and an original scheduling instruction are input into an instruction checking model, according to the method, the instruction checking model is used for respectively capturing semantic information of the current scheduling instruction and the original scheduling instruction, so that a plurality of different semantic embedding vectors are obtained, the semantic deviation degree of the current instruction and the original instruction is obtained by comparing the plurality of groups of semantic embedding vectors of the current instruction and the original instruction, and different checking prompt information is generated according to different deviation conditions; and a dispatcher is prompted to correct the instruction with semantic deviation in time. Therefore, the semantic level of the instruction can be deeply understood, the semantic difference between the instructions can be accurately recognized, the accuracy of instruction checking is improved, and the problem that the accuracy of instruction checking is low due to the fact that the semantics of the instruction cannot be deeply understood in the prior art can be solved.
Owner:POWER DISPATCHING CONTROL CENT OF GUANGDONG POWER GRID CO LTD

Thread group scheduling method and device for GPU (Graphics Processing Unit), graphics processing unit and equipment

The invention discloses a thread group scheduling method and device for a GPU (Graphics Processing Unit), the GPU and equipment. The method comprises the following steps of: 1) receiving a scheduling request of a thread group; 2) task type priority scheduling; according to a thread group weight value set by a user, obtaining execution priorities of the vertex thread group and the fragment thread group in the current scheduling period; 3) instruction type priority scheduling; pre-analyzing to-be-executed instructions of the thread group, and determining a priority sequence of schedulable instruction types; and 4) priority scheduling of the thread groups: selecting the thread group with the highest priority for scheduling according to the specified task type priority and instruction type priority in combination with the thread group generation time. The invention provides a thread group three-level scheduling strategy so as to improve the instruction throughput rate and the key task response speed of the GPU under the complex load.
Owner:WUHAN LINGJIU MICROELECTRONICS CO LTD

Efficient compression instruction handling in a processing pipeline

Systems and methods related to efficient compression instruction handling in a processing pipeline are disclosed herein. A processing pipeline may accept a mask vector from a first register, the mask vector including a set of set bits, and may accept a payload from a second register. The processing pipeline may determine a set of prefix sums for a set of portions of the mask vector by applying the portions to a set of adders in parallel to add the set bits independently for each of the portions. The processing pipeline may determine a set of indexes for the set bits, with respect to the mask vector, in parallel, using the set of prefix sums and the set of portions of the mask vector. The processing pipeline may store a set of identified values from the payload, as identified by the set of indexes, in at least one destination register.
Owner:TENSTORRENT USA INC

Batch processing job automatic debugging method, device and equipment and storage medium

The invention discloses a batch processing job automatic debugging method and device, equipment and a storage medium, and relates to the technical field of data processing and debugging verification, and the method comprises the following steps: obtaining a debugging copy; replacing a target table in the debugging copy, and determining debugging copy information; deploying a corresponding batch processing script based on the debugging copy information by using the allocated temporary account necessary permission, and determining a target batch processing operation script; and performing batch processing job debugging based on the target batch processing job script to generate a batch processing job debugging result. According to the method, the target table in the debugging copy is replaced, the batch processing script is deployed according to the distributed temporary account minimum permission, batch processing operation debugging is carried out on the target batch processing operation script, the batch processing operation debugging result is generated, and the temporary account only having the minimum necessary permission is dynamically generated for each time of debugging. And a special temporary table is automatically created as a safe output environment, so that authority distribution according to needs and operation sandbox isolation are realized.
Owner:CHINA MERCHANTS BANK

System and method for scalable stream encryption and decryption

Information security systems and methods are presented including synchronized state machines operating with private data and tamper-evident identifiers that expand the problem space of attacks by unauthorized consumers of data in flight or rest to near infinity. A quantum cryptography-resistant distribution network for identity, trust relationship, encryption, and transcoding of valuable information is described where a priori knowledge of the hardware or software is irrelevant to the safe time for the data protected and modified as needed within a multikey encryption and data control and availability assurance domains. Due to its low compute design, the methods described within are well suited for a variety of tasks including real-time data streams such as video and audio distribution.
Owner:NEURSCIENCES LLC

System and method for determining communication pathway

PendingUS20260081014A1Medical communicationDrug and medicationsIntelligent NetworkPrior authorization
Systems and methods for providing an intelligent network to onboard patients to specialty medications are described herein. A server system for managing communication is configured to receive from a user at a healthcare practice, an electronic request to obtain a medical prior authorization (medPA) for a specialty drug to administer for a patient of the healthcare practice; select, using a rules service, a communication pathway to submit the medPA to a healthcare payer from a plurality of communication pathways; and initiate a connection with the healthcare payer using the selected communication pathway.
Owner:CAREMETX LLC

Microprocessor that builds multi-fetch block macro-op cache entries in two-stage process

A microprocessor includes a prediction unit (PRU) that predicts a sequence of fetch blocks (FBlks) in a program instruction stream, a macro-op (MOP) cache (MOC) that comprises MOC entries (MEs), a decode unit, and a fusion engine. Each ME indicates whether it is a single-FBlk ME (SF-ME) that holds MOPs associated with a single FBlk whose architectural instructions have been decoded into the MOPs of the SF-ME or a multi-FBlk ME (ME-ME) that holds MOPs associated with multiple FBlks whose architectural instructions have been decoded into the MOPs of the MF-ME. For each FBlk of one or more FBlks in the program instruction stream: the decode unit decodes the architectural instructions of the FBlk into MOPs, and the fusion engine builds a SF-ME in the MOC using the decoded MOPs, and the fusion engine builds a MF-ME in the MOC using the MOPs of a series of SF-MEs.
Owner:VENTANA MICRO SYSTEMS INC

System for processing heterogeneous generative artificial intelligence models

In accordance with the present disclosure, an apparatus is provided. The apparatus includes: a first memory having a first capacity configured to store a first generative neural network model including a first parameter; and a first neural processing unit configured to generate a response corresponding to the input query using the first generative neural network model stored in the first memory; and wherein the first neural processing unit may be configured to store first execution code of the first generative neural network model, the first execution code being compiled to process the speculative decoding.
Owner:DEEPX CO LTD

Systolic array scheduling processing method, device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a systolic array scheduling processing method, device, equipment and medium, and the method comprises the steps: obtaining a data processing model, and compiling the data processing model into an initial processing task based on hardware architecture features of a systolic array, acquiring to-be-processed data and analyzing data characteristics of the to-be-processed data, monitoring a real-time operation state of the systolic array processing device, inputting the real-time operation state, the data characteristics and an initial processing task into a scheduling decision module to generate a scheduling strategy, executing the scheduling strategy to complete data processing, and collecting performance data; and updating optimization parameters of the compiling module and strategy generation parameters of the scheduling decision module based on the performance data. According to the method, the initial processing task is generated through compiling, the scheduling strategy is dynamically generated in combination with the data features and the running state, parameter updating is achieved through performance data feedback, self-adaptive closed-loop optimization of compiling and scheduling is formed, and the calculation efficiency and the energy efficiency ratio are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Processor, method, device and storage medium for data processing

A processor and a method, device and storage medium for data processing are provided. The processor includes an instruction decoder configured to decode a target instruction for a vector operation. The target instruction involves a target opcode, a source operand, and a target operand. The target opcode indicates a vector operation specified by the target instruction. The source operand specifies a source storage location in the memory for reading to-be-processed data. The target operand specifies a target storage location in the memory for writing a processed result. The processor further includes an arithmetic logic unit configured to: read the to-be-processed data from the source storage location of the memory; perform, on the to-be-processed data, an arithmetic logic operation associated with the vector operation specified by the target instruction; and write the processed result to the target storage location of the memory.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Processing heterogeneous generative artificial intelligence models

According to the present disclosure, a device is provided. The device includes a first memory of a first capacity configured to store a first generative neural network model comprising a first parameters, and a first neural processing unit configured to generate a response corresponding to an input query utilizing the first generative neural network model stored in the first memory, and wherein the first neural processing unit may be configured to store a first execution code of the first generative neural network model compiled to process speculative decoding.
Owner:DEEPX CO LTD

Instruction Deltas For Processing-In-Memory Divergence

Instruction deltas for processing-in-memory divergence are described. In one or more implementations, a system includes a memory and a processing-in-memory component configured to identify an instruction delta based on one or more undefined portions of an instruction of a PIM command and decode the instruction delta into one or more defined portions of the instruction to be used in place of the undefined portions to execute the instruction. In one or more implementations, a processing-in-memory component includes at least one computational unit of an in-memory processor that identifies an instruction delta based on one or more undefined portions of an instruction of a PIM command, decodes the instruction delta into one or more defined portions of the instruction to be used during execution in place of the undefined portions, and executes the instruction based on the defined portions.
Owner:ADVANCED MICRO DEVICES INC

RISC-V-oriented eBPF loop vectorization compiling method and device

The invention provides an RISC-V-oriented eBPF cycle vectorization compiling method and device, and the method comprises the steps: generating an RVV vector machine code, executing the parallel operation when a condition is satisfied, achieving the parallel processing of a single instruction and multiple data, carrying out the serial conversion of the time dimension of eBPF cycle into the parallel spatial dimension of an RVV vector, remarkably improving the processing throughput, and improving the compiling efficiency. And the task response time delay and the CPU occupancy rate are reduced. According to the RVV vector execution mode, a large number of repeated loop control instructions and memory read-write instructions can be reduced, the register overflow problem caused by eBPF scalar register resource shortage is relieved, the cache pressure and the branch prediction failure rate are reduced, and the energy efficiency ratio is increased. Meanwhile, an RVV vector machine code and a scalar fallback code are generated, a parallel priority and scalar bottom double execution path is formed, and it is ensured that under the conditions that a hardware environment is not supported, the data size does not reach a threshold value, vectorization verification fails or operation is abnormal, the scalar fallback code can be seamlessly switched to be executed.
Owner:BEIJING VCORE TECH CO LTD

Multi-account asynchronous bookkeeping data consistency collaboration system based on instruction driving

PendingCN121614483ADatabase updatingResource allocationInformation systems engineeringLogisim
The invention relates to the field of computer software and information system engineering, in particular to a multi-account asynchronous bookkeeping data consistency collaboration system based on instruction driving, which comprises the following steps of: packaging a request into a standard instruction and routing according to an account to ensure ordered processing of the request and improve the processing efficiency and expandability of the system; repeated requests are avoided through idempotent verification, business is mapped into atomic operation through rule analysis, and the accuracy of business logic and data consistency are guaranteed; atomic operation is executed in a single transaction, asynchronous retry is performed during conflict, the atomicity of cross-account operation is ensured, and the concurrent processing capacity is improved; by buffering high-frequency account change and batch aggregation updating, write conflicts are reduced, the system performance is stabilized, and the throughput is improved; through account checking based on daily switching snapshots, pipeline playback and reverse account flushing are supported, and a rapid repair capability is provided; through dynamic configuration and versioned management of the bookkeeping rule, visual monitoring is provided, the operation and maintenance complexity is reduced, and the system flexibility is enhanced.
Owner:HANGZHOU YIZI INFORMATION TECHNOLOGY CO LTD

Vector extract and merge instruction

There is provide an apparatus, method and medium. The apparatus comprises decoder circuitry to generate control signals in response to a vector extract and merge instruction specifying a control parameter, a first vector register, a second vector register, and a destination vector register. The apparatus comprises processing circuitry responsive to the control signals, to perform plural beats of processing, each beat comprising processing corresponding to a portion of at least the first vector register and the destination vector register. The processing, for a Kth beat comprises: extracting bits, specified by the control parameter, from a Kth portion of the first vector register, concatenating the bits with further bits, and storing the result in the Kth portion of the destination register. The further bits are, for a first portion, extracted from a first portion of the second vector register and, otherwise, from a (K−1)th portion of the first vector register.
Owner:ARM LTD

Matrix multiplication in dynamically spatially and dynamically temporally divisible architectures

A data processing apparatus includes a first vector register and a second vector register, both of which are dynamically spatially and dynamically temporally divisible. A decode circuit receives one or more matrix multiplication instructions indicating a set of first elements in the first vector register and a set of second elements in the second vector register, and generates a matrix multiplication operation in response to receiving the matrix multiplication instructions. The matrix multiplication operation causes one or more execution units to perform a matrix multiplication of the set of first elements and the set of second elements, and an average bit width of the first elements is different from an average bit width of the second elements.
Owner:ARM LTD

Microprocessor that performs partial fallback abort processing of multi-fetch block macro-op cache entries

A microprocessor includes a macro-op (MOP) cache (MOC) that holds MOC entries (MEs), including single-fetch block MEs (SF-MEs) and multi-fetch block MEs (MF-MEs), comprising MOPs decoded from architectural instructions. A prediction circuit generates fetch block (FBlk) start addresses (FBSAs) to make predictions of a sequence of FBlks fetched from an instruction cache and MEs fetched from the MOC. A back end detects that execution of a MOP of an MF-ME needs an abort and generates an abort request. A control circuit flushes all the MF-ME's MOPs and signals the prediction circuit to restart prediction at the MF-ME's FBSA. For each current FBSA of N current FBSAs used to make N predictions starting with the FBSA of the MF-ME, the prediction circuit ignores a hit on any MF-ME and instead, if the current FBSA hits on an SF-ME, predicts the SF-ME, and otherwise predicts a FBlk at the current FBSA.
Owner:VENTANA MICRO SYSTEMS INC

Instruction processing method and device, computer equipment and storage medium

Embodiments of the invention provide an instruction processing method and apparatus, a computer device and a storage medium. The method comprises the steps of obtaining a setting instruction and a vector calculation instruction from a preset instruction local memory or an instruction buffer; updating a preset shadow register according to the setting instruction to obtain a first update value; performing data recovery on the first update value by using a preset real register to obtain a second update value; and executing the vector calculation instruction according to the second update value. According to the method, the preset shadow register is used for speculative execution according to the instruction of the assembly line, an update value is obtained in advance, and when the instruction needing to be scoured is executed, the wrong update value is backed up and recovered by using the preset real register, so that the register values of the shadow register and the real register are kept consistent; therefore, the performance and the efficiency when the vector calculation instruction is executed are improved.
Owner:芯来智融半导体科技(上海)股份有限公司

Converting a stream of data using a lookaside buffer

A stream of data is accessed from a memory system by an autonomous memory access engine, converted on the fly by the memory access engine, and then presented to a processor for data processing. A portion of a lookup table (LUT) containing converted data elements is preloaded into a lookaside buffer associated with the memory access engine. As the stream of data elements is fetched from the memory system each data element in the stream of data elements is replaced with a respective converted data element obtained from the LUT in the lookaside buffer according to a content of each data element to thereby form a stream of converted data elements. The stream of converted data elements is then propagated from the memory access engine to a data processor.
Owner:TEXAS INSTRUMENTS INC

Data processing device, method and equipment

The invention discloses a data processing device, method and equipment, and belongs to the field of computers. The device comprises; the instruction decoder is used for decoding a four-word loading instruction, a first memory address and a target register are specified in the four-word loading instruction, and the processing circuit is used for responding to the four-word loading instruction and determining a next register of the target register; and the data loading module is used for reading continuous four-word data from the first memory address and orderly loading the continuous four-word data to the target register and the next register. Therefore, on the premise that the instruction format of the processor architecture is followed, the four-word instruction is loaded, and conflicts with existing instructions of the architecture are avoided. Based on similar principles, the device can store four-word instructions.
Owner:BEIJING ESWIN COMPUTING TECH CO LTD

Instructions to convert from FP16 to BF8

Techniques for converting FP16 data elements to BF8 data elements using a single instruction are described. An exemplary apparatus includes decoder circuitry to decode a single instruction, the single instruction to include a one or more fields to identify a source operand, one or more fields to identify a destination operand, and one or more fields for an opcode, the opcode to indicate that execution circuitry is to convert packed half-precision floating-point data from the identified source to packed bfloat8 data and store the packed bfloat8 data into corresponding data element positions of the identified destination operand; and execution circuitry to execute the decoded instruction according to the opcode to convert packed half-precision floating-point data from the identified source to packed bfloat8 data and store the packed bfloat8 data into corresponding data element positions.
Owner:INTEL CORP

Integer divider and method for realizing Radix-16 operation, chip and electronic equipment

The invention discloses an integer divider and method for realizing Radix-16 operation, a chip and electronic equipment, and belongs to the technical field of electronics. The integer divider comprises a preprocessing module, an iterative calculation module and a post-processing module, wherein the iterative calculation module comprises four levels of Radix-2 sub-operation modules which are connected in sequence; the step of generating the 1-bit quotient and the remainder update by the current-level Radix-2 in the primary iteration Radix-16 comprises the following steps of: obtaining an initialized divisor d and a partial remainder wj; determining a 1-bit quotient qj of the current level Radix-2 according to the partial remainder wj; obtaining an updated partial remainder wj + 1 according to a preset iteration recursion rule, and outputting the updated partial remainder wj + 1 to a next-level Radix-2 sub-operation module; and the post-processing module combines the 1-bit quotients generated by each level of Radix-2 to generate 4-bit quotients, and combines the 4-bit quotients as a final quotient value of Radix-16 division operation.
Owner:SANECHIPS TECH CO LTD

Command processing circuit, command processing method, and PIM device

The invention provides a command processing circuit, a command processing method and a PIM device. The command processing circuit comprises a PIM command identification module which is electrically connected with a command address bus of the PIM device and is configured to receive a command signal through the command address bus and judge the type of the currently received command signal; if it is judged that the currently received command signal is a DRAM conventional command, the command signal is sent to a conventional command processing module; if it is judged that the currently received command signal is a PIM related command, the command signal is sent to a PIM command processing module; the PIM command processing module is electrically connected with the PIM command identification module and is configured to receive a PIM related command, decode the PIM related command and perform corresponding operation according to a decoding result; and the conventional command processing module is electrically connected with the PIM command identification module and is configured to receive the DRAM conventional command and decode the DRAM conventional command. The method is at least beneficial to avoiding mode switching and improving system performance.
Owner:RUILI INTEGRATED CIRCUIT CO LTD

Apparatus and method for hiding vector load latency in a time-based vector coprocessor

A processor includes a time counter and a vector coprocessor for executing vector instructions for statically dispatching vector instructions with preset execution times based on a write time of a register in a coprocessor register scoreboard and a time counter provided to a vector execution pipeline. The processor also provides a method for hiding the latency of the vector load instructions.
Owner:SIMPLEX MICRO INC