Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

940results about "Instruction analysis" patented technology

Cache-line retention hint information for conditional write instruction

In response to instruction decoding circuitry decoding a conditional write instruction, processing circuitry determines whether a predetermined condition is satisfied for a target cache line corresponding to a target address specified by the conditional write instruction. If the predetermined condition is satisfied for the target cache line, a write request is issued to update the target cache line. If the predetermined condition is not satisfied for the target cache line, a failure indication is returned. The processing circuitry selects, depending on whether the sequence of instructions specifies cache-line-retention hint information applicable to the conditional write instruction, whether to prevent a unique coherency state of the target cache line being relinquished by a local cache associated with the processing circuitry for a retention period following processing of the conditional write instruction. The unique coherency state comprises a coherency state in which the processing circuitry has exclusive right to update the target cache line.
Owner:ARM LTD

Matrix multiplication and accumulation operation unit and operation method, hardware accelerator and electronic equipment

The embodiment of the invention provides a matrix multiplication and accumulation operation unit and method, a hardware accelerator and electronic equipment, and the matrix multiplication and accumulation operation unit comprises a data loading storage engine, a tensor register file and a matrix multiplication engine. The data loading and storage engine is used for loading data of a plurality of matrixes to be subjected to matrix multiplication and accumulation calculation; the tensor register file is used for storing data of a plurality of matrixes acquired from the data loading storage engine; the tensor register file comprises at least three tensor register groups, each tensor register group comprises a plurality of tensor registers, and different tensor register groups are used for storing data of different matrixes in the plurality of matrixes; and the matrix multiplication engine is used for carrying out matrix multiplication accumulation calculation based on the data of the plurality of matrixes stored in the tensor register file. According to the embodiment of the invention, more efficient MMA calculation is realized under the conditions of low cost, low power consumption and less occupied space.
Owner:ALIBABA (CHINA) CO LTD

Technique for controlling stashing of data

An apparatus has decoder circuitry within a first processing element to decode instructions, in order to respond to a sequence of instructions by generating control signals. Processing circuitry within the first processing element is responsive to the control signals to perform operations defined by the sequence of instructions. The decoder circuitry is responsive to a stashee hint instruction in the sequence of instructions to issue control signals to cause the processing circuitry to issue a stash interest request via an interface used to couple the first processing element to interconnect circuitry. The stash interest request is arranged to provide a given memory address indication determined from the stashee hint instruction and to trigger stashing control circuitry accessible via the interconnect circuitry to cause updated data associated with the given memory address indication to be made available for stashing in an associated storage structure of the first processing element.
Owner:ARM LTD

Instruction speculation execution method and device of vector processor and storage medium

The invention discloses an instruction speculation execution method and device of a vector processor and a storage medium. The method comprises the following steps: in response to a vector configuration instruction identified at the front end of a vector processor pipeline, querying a vector prediction configuration table based on instruction information of the vector configuration instruction, and determining a prediction vector configuration parameter; updating a speculative vector state of a back end of the vector processor pipeline based on the predictive vector configuration parameter; performing a speculative process on a vector instruction following the vector configuration instruction based on the speculative vector state; when an actual execution result of the vector configuration instruction is obtained, determining an actual vector configuration parameter; comparing the actual vector configuration parameters with the prediction vector configuration parameters, and when the actual vector configuration parameters are consistent with the prediction vector configuration parameters, determining that speculation processing is effective and updating confidence information in the vector prediction configuration table; and if not, performing a flushing operation on the vector processor assembly line, and updating the vector prediction configuration table by using the actual vector configuration parameters. The processing efficiency of the vector processor assembly line can be improved.
Owner:SHANGHAI LINGRUI INTELLIGENT CORE COMPUTING TECHNOLOGY CO LTD

Matrix multiply accumulation operation unit and operation method, and hardware accelerator and electronic device

Provided in the embodiments of the present disclosure are a matrix multiply accumulation (MMA) operation unit and operation method, and a hardware accelerator and an electronic device. The MMA operation unit comprises a data load / store engine, a tensor register file and a matrix multiplication engine, wherein the data load / store engine is used for loading data of a plurality of matrixes to be subjected to MMA computation; the tensor register file is used for storing the data of the plurality of matrixes that is acquired from the data load / store engine; the tensor register file comprises at least three tensor register groups, each tensor register group comprising a plurality of tensor registers, and different tensor register groups being used for storing data of different matrixes among the plurality of matrixes; and the matrix multiplication engine is used for performing MMA computation on the basis of the data of the plurality of matrixes that is stored in the tensor register file. By means of the embodiments of the present disclosure, more efficient MMA computation is realized with low cost, low power consumption and a smaller footprint.
Owner:ALIBABA (CHINA) CO LTD

Instruction checking method and device based on semantic matching, terminal equipment and storage medium

The invention discloses an instruction checking method and device based on semantic matching, terminal equipment and a storage medium, and belongs to the technical field of instruction checking. The method comprises the steps that a current scheduling instruction and an original scheduling instruction are input into an instruction checking model, according to the method, the instruction checking model is used for respectively capturing semantic information of the current scheduling instruction and the original scheduling instruction, so that a plurality of different semantic embedding vectors are obtained, the semantic deviation degree of the current instruction and the original instruction is obtained by comparing the plurality of groups of semantic embedding vectors of the current instruction and the original instruction, and different checking prompt information is generated according to different deviation conditions; and a dispatcher is prompted to correct the instruction with semantic deviation in time. Therefore, the semantic level of the instruction can be deeply understood, the semantic difference between the instructions can be accurately recognized, the accuracy of instruction checking is improved, and the problem that the accuracy of instruction checking is low due to the fact that the semantics of the instruction cannot be deeply understood in the prior art can be solved.
Owner:POWER DISPATCHING CONTROL CENT OF GUANGDONG POWER GRID CO LTD

Software defined super cores

Techniques for software defined super core usage are described. In some examples, a first and second processor core are to operate as a single virtual core enabled by the operating system to fetch the first set of instruction segments of the single threaded program and the second set of instruction segments of the single threaded program concurrently using flow control instructions that have been inserted into the single threaded program.
Owner:INTEL CORP

Pipelined Decoding Microarchitecture Design Method for RISC-V Vector Instructions

A pipelined decoding microarchitecture design method for RISC-V vector instructions includes the following steps: S1, extracting an instruction packet according to a PC, performing preprocessing, and writing a preprocessing result, each instruction field and preprocessed instruction information into a cache; S2, reading the cache, obtaining pre-decoding result data and recognizing each configuration instruction; S3, transmitting the pre-decoding result data to an instruction buffer, and acquiring field information, a split quantity and predicted vtype information; and splitting each instruction into multiple microoperations, and constructing a queue to store each microoperation; and S4, transmitting first, by means of the queue, each configuration instruction, performing decoding and execution, acquiring vtype information generated after execution, and when the vtype information generated after execution is identical with the predicted vtype information, controlling the number of microoperations input to multiple decoders, and transmitting a decoding result to an instruction slot.
Owner:JIANGSU HUACHUANG MICROSYSTEM CO LTD

Thread group scheduling method and device for GPU (Graphics Processing Unit), graphics processing unit and equipment

The invention discloses a thread group scheduling method and device for a GPU (Graphics Processing Unit), the GPU and equipment. The method comprises the following steps of: 1) receiving a scheduling request of a thread group; 2) task type priority scheduling; according to a thread group weight value set by a user, obtaining execution priorities of the vertex thread group and the fragment thread group in the current scheduling period; 3) instruction type priority scheduling; pre-analyzing to-be-executed instructions of the thread group, and determining a priority sequence of schedulable instruction types; and 4) priority scheduling of the thread groups: selecting the thread group with the highest priority for scheduling according to the specified task type priority and instruction type priority in combination with the thread group generation time. The invention provides a thread group three-level scheduling strategy so as to improve the instruction throughput rate and the key task response speed of the GPU under the complex load.
Owner:WUHAN LINGJIU MICROELECTRONICS CO LTD

Efficient compression instruction handling in a processing pipeline

Systems and methods related to efficient compression instruction handling in a processing pipeline are disclosed herein. A processing pipeline may accept a mask vector from a first register, the mask vector including a set of set bits, and may accept a payload from a second register. The processing pipeline may determine a set of prefix sums for a set of portions of the mask vector by applying the portions to a set of adders in parallel to add the set bits independently for each of the portions. The processing pipeline may determine a set of indexes for the set bits, with respect to the mask vector, in parallel, using the set of prefix sums and the set of portions of the mask vector. The processing pipeline may store a set of identified values from the payload, as identified by the set of indexes, in at least one destination register.
Owner:TENSTORRENT USA INC

Batch processing job automatic debugging method, device and equipment and storage medium

The invention discloses a batch processing job automatic debugging method and device, equipment and a storage medium, and relates to the technical field of data processing and debugging verification, and the method comprises the following steps: obtaining a debugging copy; replacing a target table in the debugging copy, and determining debugging copy information; deploying a corresponding batch processing script based on the debugging copy information by using the allocated temporary account necessary permission, and determining a target batch processing operation script; and performing batch processing job debugging based on the target batch processing job script to generate a batch processing job debugging result. According to the method, the target table in the debugging copy is replaced, the batch processing script is deployed according to the distributed temporary account minimum permission, batch processing operation debugging is carried out on the target batch processing operation script, the batch processing operation debugging result is generated, and the temporary account only having the minimum necessary permission is dynamically generated for each time of debugging. And a special temporary table is automatically created as a safe output environment, so that authority distribution according to needs and operation sandbox isolation are realized.
Owner:CHINA MERCHANTS BANK

System and method for scalable stream encryption and decryption

Information security systems and methods are presented including synchronized state machines operating with private data and tamper-evident identifiers that expand the problem space of attacks by unauthorized consumers of data in flight or rest to near infinity. A quantum cryptography-resistant distribution network for identity, trust relationship, encryption, and transcoding of valuable information is described where a priori knowledge of the hardware or software is irrelevant to the safe time for the data protected and modified as needed within a multikey encryption and data control and availability assurance domains. Due to its low compute design, the methods described within are well suited for a variety of tasks including real-time data streams such as video and audio distribution.
Owner:NEURSCIENCES LLC

A blockchain hybrid granularity parallel transaction processing method for federated learning

The present invention relates to the field of blockchain technology and specifically discloses a hybrid-granularity parallel transaction processing method for blockchains oriented to federated learning. The method comprises the following steps: Step 1: Obtaining BCFL transactions; Step 2: Traversing BCFL transactions and grouping transactions with the same BCFL task type and no data dependencies or read / write conflicts into the same transaction group, thereby obtaining multiple transaction groups; Step 3: Sequentially extracting the transaction groups and processing the transactions within the transaction groups based on a multi-EVM parallel processing mechanism. This method significantly improves the throughput and system response speed when executing BCFL tasks.
Owner:先进计算与关键软件(信创)海河实验室 +1

Cache replacement method and device, equipment and storage medium

ActiveCN120561032AAssociative processorsInstruction analysisCache accessParallel computing
The invention provides a cache replacement method and device, equipment and a storage medium, a cache is formed by connecting n paths of groups, and n is 4 or 8; if n is 4, dividing every two paths into one group to obtain a first group and a second group; 4-bit access information is recorded; if n is 8, every two paths are divided into one group, and a third group, a fourth group, a fifth group and a sixth group are obtained; every two groups are divided into a large group, and a first large group and a second large group are obtained; 10 bit access information is recorded; the method comprises the following steps: acquiring cached access information; determining a cache line which is not accessed for the longest time according to the access information; and replacing the cache line which is not accessed for the longest time. According to the method provided by the invention, the cache line can be quickly replaced through the access information.
Owner:SHANGHAI SMARTLOGIC TECHNOLOGY LTD

System and method for determining communication pathway

PendingUS20260081014A1Medical communicationDrug and medicationsIntelligent NetworkPrior authorization
Systems and methods for providing an intelligent network to onboard patients to specialty medications are described herein. A server system for managing communication is configured to receive from a user at a healthcare practice, an electronic request to obtain a medical prior authorization (medPA) for a specialty drug to administer for a patient of the healthcare practice; select, using a rules service, a communication pathway to submit the medPA to a healthcare payer from a plurality of communication pathways; and initiate a connection with the healthcare payer using the selected communication pathway.
Owner:CAREMETX LLC

Microprocessor that builds multi-fetch block macro-op cache entries in two-stage process

A microprocessor includes a prediction unit (PRU) that predicts a sequence of fetch blocks (FBlks) in a program instruction stream, a macro-op (MOP) cache (MOC) that comprises MOC entries (MEs), a decode unit, and a fusion engine. Each ME indicates whether it is a single-FBlk ME (SF-ME) that holds MOPs associated with a single FBlk whose architectural instructions have been decoded into the MOPs of the SF-ME or a multi-FBlk ME (ME-ME) that holds MOPs associated with multiple FBlks whose architectural instructions have been decoded into the MOPs of the MF-ME. For each FBlk of one or more FBlks in the program instruction stream: the decode unit decodes the architectural instructions of the FBlk into MOPs, and the fusion engine builds a SF-ME in the MOC using the decoded MOPs, and the fusion engine builds a MF-ME in the MOC using the MOPs of a series of SF-MEs.
Owner:VENTANA MICRO SYSTEMS INC

System for processing heterogeneous generative artificial intelligence models

In accordance with the present disclosure, an apparatus is provided. The apparatus includes: a first memory having a first capacity configured to store a first generative neural network model including a first parameter; and a first neural processing unit configured to generate a response corresponding to the input query using the first generative neural network model stored in the first memory; and wherein the first neural processing unit may be configured to store first execution code of the first generative neural network model, the first execution code being compiled to process the speculative decoding.
Owner:DEEPX CO LTD

Systolic array scheduling processing method, device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a systolic array scheduling processing method, device, equipment and medium, and the method comprises the steps: obtaining a data processing model, and compiling the data processing model into an initial processing task based on hardware architecture features of a systolic array, acquiring to-be-processed data and analyzing data characteristics of the to-be-processed data, monitoring a real-time operation state of the systolic array processing device, inputting the real-time operation state, the data characteristics and an initial processing task into a scheduling decision module to generate a scheduling strategy, executing the scheduling strategy to complete data processing, and collecting performance data; and updating optimization parameters of the compiling module and strategy generation parameters of the scheduling decision module based on the performance data. According to the method, the initial processing task is generated through compiling, the scheduling strategy is dynamically generated in combination with the data features and the running state, parameter updating is achieved through performance data feedback, self-adaptive closed-loop optimization of compiling and scheduling is formed, and the calculation efficiency and the energy efficiency ratio are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Processor, method, device and storage medium for data processing

According to an embodiment of the present disclosure, a processor and a method, device and storage medium for data processing are provided. The processor includes an instruction decoder configured to decode a target instruction for a vector operation. The target instruction involves a target opcode, a source operand and a target operand. The target opcode indicates the vector operation specified by the target instruction. The source operand specifies a source storage location in the memory for reading data to be processed. The target operand specifies a target storage location in the memory for writing a processing result. The processor also includes an arithmetic logic unit coupled to the instruction decoder and the memory. The arithmetic logic unit is configured to: read data to be processed from a source storage location in the memory; perform an arithmetic logic operation associated with the vector operation specified by the target instruction on the data to be processed; and write the processing result to a target storage location in the memory. In this way, the efficiency of vector calculations can be improved.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Dependence instruction screening method and device and medium

The invention provides a dependency instruction screening method and device and a medium, and relates to the technical field of computers, the method comprises the following steps: step 1, averagely splitting a test instruction set of a target chip into two parts of instructions; step 2, storing one part of instructions in the two parts of instructions into an instruction group, and judging whether the instructions of the instruction group influence a target instruction or not; step 3, if not, emptying the instruction group, storing the other part of instructions in the two parts of instructions into the instruction group, and judging whether the instructions of the instruction group influence the target instruction or not; 4, if yes, judging whether the number of instructions of the instruction group is 1 or not; step 5, if the number of the instructions is not 1, splitting the instructions of the instruction group into two parts of instructions, emptying the instruction group, and returning to the step 2; and 6, if the number of the instructions is 1, determining the instructions of the instruction group as dependent instructions of the target instruction. According to the splitting and screening method based on the dichotomy, the screening efficiency of the dependency instruction can be improved.
Owner:SICHUAN TIANYI COMHEART TELECOM

Processor, method, device and storage medium for data processing

A processor and a method, device and storage medium for data processing are provided. The processor includes an instruction decoder configured to decode a target instruction for a vector operation. The target instruction involves a target opcode, a source operand, and a target operand. The target opcode indicates a vector operation specified by the target instruction. The source operand specifies a source storage location in the memory for reading to-be-processed data. The target operand specifies a target storage location in the memory for writing a processed result. The processor further includes an arithmetic logic unit configured to: read the to-be-processed data from the source storage location of the memory; perform, on the to-be-processed data, an arithmetic logic operation associated with the vector operation specified by the target instruction; and write the processed result to the target storage location of the memory.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Processing heterogeneous generative artificial intelligence models

According to the present disclosure, a device is provided. The device includes a first memory of a first capacity configured to store a first generative neural network model comprising a first parameters, and a first neural processing unit configured to generate a response corresponding to an input query utilizing the first generative neural network model stored in the first memory, and wherein the first neural processing unit may be configured to store a first execution code of the first generative neural network model compiled to process speculative decoding.
Owner:DEEPX CO LTD

Instruction Deltas For Processing-In-Memory Divergence

Instruction deltas for processing-in-memory divergence are described. In one or more implementations, a system includes a memory and a processing-in-memory component configured to identify an instruction delta based on one or more undefined portions of an instruction of a PIM command and decode the instruction delta into one or more defined portions of the instruction to be used in place of the undefined portions to execute the instruction. In one or more implementations, a processing-in-memory component includes at least one computational unit of an in-memory processor that identifies an instruction delta based on one or more undefined portions of an instruction of a PIM command, decodes the instruction delta into one or more defined portions of the instruction to be used during execution in place of the undefined portions, and executes the instruction based on the defined portions.
Owner:ADVANCED MICRO DEVICES INC

RISC-V-oriented eBPF loop vectorization compiling method and device

The invention provides an RISC-V-oriented eBPF cycle vectorization compiling method and device, and the method comprises the steps: generating an RVV vector machine code, executing the parallel operation when a condition is satisfied, achieving the parallel processing of a single instruction and multiple data, carrying out the serial conversion of the time dimension of eBPF cycle into the parallel spatial dimension of an RVV vector, remarkably improving the processing throughput, and improving the compiling efficiency. And the task response time delay and the CPU occupancy rate are reduced. According to the RVV vector execution mode, a large number of repeated loop control instructions and memory read-write instructions can be reduced, the register overflow problem caused by eBPF scalar register resource shortage is relieved, the cache pressure and the branch prediction failure rate are reduced, and the energy efficiency ratio is increased. Meanwhile, an RVV vector machine code and a scalar fallback code are generated, a parallel priority and scalar bottom double execution path is formed, and it is ensured that under the conditions that a hardware environment is not supported, the data size does not reach a threshold value, vectorization verification fails or operation is abnormal, the scalar fallback code can be seamlessly switched to be executed.
Owner:BEIJING VCORE TECH CO LTD

Multi-account asynchronous bookkeeping data consistency collaboration system based on instruction driving

PendingCN121614483ADatabase updatingResource allocationInformation systems engineeringLogisim
The invention relates to the field of computer software and information system engineering, in particular to a multi-account asynchronous bookkeeping data consistency collaboration system based on instruction driving, which comprises the following steps of: packaging a request into a standard instruction and routing according to an account to ensure ordered processing of the request and improve the processing efficiency and expandability of the system; repeated requests are avoided through idempotent verification, business is mapped into atomic operation through rule analysis, and the accuracy of business logic and data consistency are guaranteed; atomic operation is executed in a single transaction, asynchronous retry is performed during conflict, the atomicity of cross-account operation is ensured, and the concurrent processing capacity is improved; by buffering high-frequency account change and batch aggregation updating, write conflicts are reduced, the system performance is stabilized, and the throughput is improved; through account checking based on daily switching snapshots, pipeline playback and reverse account flushing are supported, and a rapid repair capability is provided; through dynamic configuration and versioned management of the bookkeeping rule, visual monitoring is provided, the operation and maintenance complexity is reduced, and the system flexibility is enhanced.
Owner:HANGZHOU YIZI INFORMATION TECHNOLOGY CO LTD

Vector extract and merge instruction

There is provide an apparatus, method and medium. The apparatus comprises decoder circuitry to generate control signals in response to a vector extract and merge instruction specifying a control parameter, a first vector register, a second vector register, and a destination vector register. The apparatus comprises processing circuitry responsive to the control signals, to perform plural beats of processing, each beat comprising processing corresponding to a portion of at least the first vector register and the destination vector register. The processing, for a Kth beat comprises: extracting bits, specified by the control parameter, from a Kth portion of the first vector register, concatenating the bits with further bits, and storing the result in the Kth portion of the destination register. The further bits are, for a first portion, extracted from a first portion of the second vector register and, otherwise, from a (K−1)th portion of the first vector register.
Owner:ARM LTD

Instruction set control method for discrete semiconductor testing

The invention relates to the technical field of discrete semiconductor testing, and discloses an instruction set control method for discrete semiconductor testing, which comprises the following steps: packaging test steps and parameters of a tested device into an instruction set through an upper computer, and writing the instruction set into a memory of a lower computer. When the test starts, the upper computer sends a starting instruction, the lower computer analyzes and executes an instruction set in the memory, the corresponding hardware board card is operated to complete a test item, and finally result data is stored to a specified memory address and a completion flag bit is returned to the upper computer. And after receiving the completion mark, the upper computer reads and displays the test result in real time. According to the method, high-speed data transmission and parallel testing of multiple hardware board cards are achieved through PCIE driving, the testing efficiency, stability and anti-jamming capability are remarkably improved, the testing cost is reduced, and the method is suitable for the automatic testing process of various discrete semiconductor devices.
Owner:DONGGUAN NOLI SEMICONDUCTOR TECHNOLOGY CO LTD

Apparatus and method for vector packed signed / unsigned shift, round, and saturate

Apparatus and method for signed and unsigned shift, round and saturate using different data element values. For example, one embodiment of an apparatus comprises a decoder to decode an instruction having fields for a first packed data source operand to provide a first source data element and a second source data element, a second packed data source operand or immediate to provide a first shift value and a second shift value corresponding to the first source data element and second source data element, respectively, and a packed data destination operand to indicate a first result value and a second result value corresponding to the first source data element and second source data element, and execution circuitry to execute the decoded instruction to: shift the first source data element by an amount based on the first shift value to generate a first shifted data element; shift the second source data element by an amount based on the second shift value to generate a second shifted data element; update a saturation indicator responsive to detecting a saturation condition resulting from the shift of the first and / or second source data elements; round and / or saturate the first and second shifted data elements in accordance with a specified rounding mode and the saturation indicator, respectively, to generate the first and second result data elements; and store the first result value and the second result value in a first data element location and a second data element location in a destination register.
Owner:INTEL CORP

Matrix multiplication in dynamically spatially and dynamically temporally divisible architectures

A data processing apparatus includes a first vector register and a second vector register, both of which are dynamically spatially and dynamically temporally divisible. A decode circuit receives one or more matrix multiplication instructions indicating a set of first elements in the first vector register and a set of second elements in the second vector register, and generates a matrix multiplication operation in response to receiving the matrix multiplication instructions. The matrix multiplication operation causes one or more execution units to perform a matrix multiplication of the set of first elements and the set of second elements, and an average bit width of the first elements is different from an average bit width of the second elements.
Owner:ARM LTD