Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

71 results about "Speculative execution" patented technology

Speculative execution is an optimization technique where a computer system performs some task that may not be needed. Work is done before it is known whether it is actually needed, so as to prevent a delay that would have to be incurred by doing the work after it is known that it is needed. If it turns out the work was not needed after all, most changes made by the work are reverted and the results are ignored.

Vector configuration instruction implementation method and system and storage medium

The invention is suitable for the technical field of processors, and particularly relates to a vector configuration instruction implementation method and system and a storage medium. The implementation system comprises an instruction fetching unit, a decoding unit, a vector configuration instruction execution unit, an execution management unit and a vector execution unit. The vector configuration instruction execution unit is used for obtaining the number of the idle register, renaming the received vector configuration instruction according to the number of the idle register to obtain a renamed vector configuration instruction and renaming information, and executing the renamed vector configuration instruction to obtain the renamed vector configuration instruction. And generating vector configuration information according to the execution result and sending the vector configuration information to a vector execution unit. Compared with the prior art, the vector configuration instruction is realized by introducing a register renaming technology, and simple and quick recovery can be carried out when speculative execution errors occur while the speculative execution capability is reserved, so that the design complexity can be reduced and the execution efficiency of the instruction can be improved while the performance is considered.
Owner:RIVAI TECH (SHENZHEN) CO LTD

Speculative execution of kernel programs in a chiplet based architecture

One embodiment provides a multi-chiplet graphics processor comprising a plurality of chiplets, where a chiplet of the plurality of chiplets comprise a memory interface, processing resources configured to execute threads of a kernel, and thread dispatch circuitry to facilitate dispatch of threads of the kernel to the processing resources. The processing resources are configured to execute threads of a first kernel, receive dispatch of threads of a second kernel for execution before completion of the first kernel as threads of the first kernel retire, execute a first phase of the second kernel during completion of execution of the first kernel, via a thread of the first kernel, signal an event via an uncached write to a global memory, and execute a second phase of the second kernel based on detection of the event via an uncached read from the global memory.
Owner:INTEL CORP

Instruction speculation execution method and device of vector processor and storage medium

The invention discloses an instruction speculation execution method and device of a vector processor and a storage medium. The method comprises the following steps: in response to a vector configuration instruction identified at the front end of a vector processor pipeline, querying a vector prediction configuration table based on instruction information of the vector configuration instruction, and determining a prediction vector configuration parameter; updating a speculative vector state of a back end of the vector processor pipeline based on the predictive vector configuration parameter; performing a speculative process on a vector instruction following the vector configuration instruction based on the speculative vector state; when an actual execution result of the vector configuration instruction is obtained, determining an actual vector configuration parameter; comparing the actual vector configuration parameters with the prediction vector configuration parameters, and when the actual vector configuration parameters are consistent with the prediction vector configuration parameters, determining that speculation processing is effective and updating confidence information in the vector prediction configuration table; and if not, performing a flushing operation on the vector processor assembly line, and updating the vector prediction configuration table by using the actual vector configuration parameters. The processing efficiency of the vector processor assembly line can be improved.
Owner:SHANGHAI LINGRUI INTELLIGENT CORE COMPUTING TECHNOLOGY CO LTD

Instruction pipeline processing method of processor and processor

The invention provides an instruction pipeline processing method of a processor and the processor. A processor includes: a control module; the control module is set to execute the branch instruction speculatively according to the branch prediction direction when detecting that the first instruction subjected to initial decoding is the branch instruction, and execute the branch instruction if detecting that the second instruction subjected to initial decoding is the function call instruction or the function return instruction before the speculation execution result of the first instruction is generated. If yes, pausing all operations after the instruction processing assembly line performs initial decoding on the second instruction, and blocking the instruction fetching operation of the instruction processing assembly line on the next instruction until a speculation execution result of the first instruction is obtained; wherein the operation after the initial decoding comprises the step of carrying out a push-in or push-out operation of a return address stack (RAS) according to the second instruction. According to the technical scheme, the pollution risk caused by speculative execution of the branch instruction to the return address stack can be shielded, so that the hardware implementation logic of the return address stack is simplified, and the circuit area and power consumption are saved.
Owner:SUNMMIO SCIENCE & TECHNOLOGY (BEIJING) CO LTD

Speculative execution of kernel programs in chiplet-based architectures

The invention relates to speculative execution of kernel programs in a chiplet-based architecture. One embodiment provides a multi-chiplet graphics processor comprising a plurality of chiplets, where a chiplet of the plurality of chiplets comprises a memory interface; a processing resource configured to execute a thread of the kernel; and a thread dispatch circuitry module to facilitate dispatch of threads of the kernel to the processing resources. The processing resource is configured to: execute a thread of the first core; when the threads of the first core are revolved, receiving the dispatch of the threads of the second core for execution before the first core is completed; executing the first stage of the second core during execution completion of the first core; signaling, via a thread of the first core, an event via an uncached write to the global memory; and performing a second stage of the second kernel based on the detection of the event via the uncached read from the global memory.
Owner:INTEL CORP

Automatic illegal character cleaning system

The invention relates to the technical field of computer data processing and network security, in particular to an illegal character automatic cleaning system, which comprises a rule compiling module for monitoring rule change, performing semantic fusion and topological mapping on a rule set, constructing a deterministic finite automaton and mapping the deterministic finite automaton into a state transition table; the state switching module is used for constructing a double-buffer context and operating lock-free switching through an atomic pointer to realize hot updating; the speculation execution module is used for carrying out vectorization pre-scanning based on a state transition table by utilizing single-instruction multi-data stream parallel loading, identifying a walk path, falling into a safe state for releasing and falling into a trap state for triggering external verification; the self-adaptive feedback module is used for counting trap state triggering frequency, generating a rule allergy report and dynamically adjusting the size of a read fragment; according to the method, the contradiction between rule flexibility and execution efficiency is solved, and high-performance cleaning based on speculative execution is realized.
Owner:北京啄木鸟云健康科技有限公司

Consistent speculation of pointer authentication

In an embodiment, a processor includes hardware circuitry which may be used to authenticate instruction operands. The processor may execute instructions that perform operand authentication both speculatively and non-speculatively. During speculative execution of such instructions, the processor may execute authentication such that no differences in observable state of the processor, relative to authentication result, are detectable via a side channel. During speculative execution, a result of authentication may be deferred until speculative execution of the instruction, and additional instructions, may be completed. Upon resolution of a condition that indicates acceptance of the speculative execution, a speculative execution result may cause a processor exception and stalling of execution at the instruction to be performed.
Owner:APPLE INC

Instruction processing method and device, computer equipment and storage medium

Embodiments of the invention provide an instruction processing method and apparatus, a computer device and a storage medium. The method comprises the steps of obtaining a setting instruction and a vector calculation instruction from a preset instruction local memory or an instruction buffer; updating a preset shadow register according to the setting instruction to obtain a first update value; performing data recovery on the first update value by using a preset real register to obtain a second update value; and executing the vector calculation instruction according to the second update value. According to the method, the preset shadow register is used for speculative execution according to the instruction of the assembly line, an update value is obtained in advance, and when the instruction needing to be scoured is executed, the wrong update value is backed up and recovered by using the preset real register, so that the register values of the shadow register and the real register are kept consistent; therefore, the performance and the efficiency when the vector calculation instruction is executed are improved.
Owner:芯来智融半导体科技(上海)股份有限公司

Cache row rollback-based processor architecture optimization method free from side channel attack

The invention belongs to the technical field of computer system security. The invention provides a processor architecture optimization method free of side channel attack based on cache line rollback. According to the embodiment of the invention, whether speculative execution of the memory access instruction is consistent with an actual result is checked; and sending a Rollback or Commit signal according to a branch prediction result and an actual branch jump result so as to trigger rollback or submission operation of a cache line. And when the memory access instruction is in a speculative execution state and Cache miss occurs, recording detailed information of the replaced cache line in the Line Buffer. When a Rollback signal is received, the information stored in the Line Buffer is utilized to restore the cache line to a state before branch prediction; when a Commit signal is received, it is confirmed that the previous speculation execution is correct, and the corresponding cache line can continue to be reserved or updated, so that side channel attacks are effectively defended.
Owner:XIDIAN UNIV

Cooperative evaluation method for processor micro-architecture performance upper limit exploration and related device

The embodiment of the application discloses a kind of collaborative evaluation method for exploring the upper limit of processor micro-architecture performance and related device, the method includes that various idealized micro-architecture components are constructed according to the collaborative process of ideal model by real processor model, processor micro-architecture performance is collaboratively evaluated, and collaborative process includes: real processor model sends instruction query and information request to ideal model, and the request carries the memory address of dynamic instruction;Ideal model executes dynamic instruction according to memory address and obtains target value in normal working state, and target value is returned as the response of the request to real processor model;Real processor model executes dynamic instruction and obtains actual value in execution phase;Target value and actual value are verified in submission phase, and verification result is obtained.Using the embodiment of the application, the theoretical performance upper limit of value prediction and other speculative execution techniques can be accurately quantified, and the real performance bottleneck of the entire processor system integrated with the technique can be identified.
Owner:UNIV OF SCI & TECH OF CHINA

Low-overhead processor cache micro-architecture defense method and device and computer equipment

The invention relates to a low-overhead processor cache micro-architecture defense method and device and computer device.The low-overhead processor cache micro-architecture defense method comprises the steps that when a missing state keeping register module takes out a loading instruction waiting for data from a replay queue, a target branch mask of the loading instruction is obtained; the replay queue is used for storing an instruction which is stagnated due to miss of the cache; judging whether the loading instruction is in a speculative execution state or not according to the target branch mask; and when the loading instruction is in the speculative execution state, stopping a cache write-in operation corresponding to the loading instruction. Through the method and the device, the problem of sensitive information leakage caused by incapability of defending against the cache side channel attack is solved, defending against the cache side channel attack is realized, and sensitive information leakage is prevented.
Owner:HANGZHOU HIGH-TECH ZONE (BINJIANG) INSTITUTE OF BLOCKCHAIN & DATA SECURITY +1

Technique for generating tests for a processing device

A computer-implemented system and method is described for generating a supplemented test to be performed for a processing device under test. The method comprises receiving result information from performing an initial test on a representation of the processing device under test, the initial test providing a program to be executed and associated data, and determining from the result information a sequence of instructions that would be executed by the processing device under test when adopting in order execution of the program. A plurality of stress injection elements are provided, each having an associated hazard condition. One or more iterations of a test modification process are performed, comprising: identifying, from the sequence of instructions, a given control flow influencing event that causes a selection between a first path that will be chosen when adopting the in order execution of the program, and a second path; identifying one or more instructions in an instruction window associated with the given control flow influencing event, and one or more associated resources; selecting, in dependence on the one or more instructions and the one or more associated resources, a stress injection element from the plurality of stress injection elements; and employing the selected stress injection element to modify the program by introducing into the second path one or more additional instructions that, when executed, will provide a stimulus aimed at inducing the associated hazard condition during operation of the processing device under test when speculative execution of instructions to support out of order execution causes the instructions in the second path to be speculatively executed. A supplemented test is then generated using a modified program generated through performance of the one or more iterations of the test modification process.
Owner:ARM LTD

Computer systems and programs

The system flexibly presents additional questions based on the user's responses. [Solution] The computer system outputs output data to present a first question to the user, sequentially receives first stream data from the user including the answer to the first question, performs a speculative execution process for generating question candidates one or more times while receiving the first stream data, and when the input of the answer to the first question is finished, generates a second question based on the question candidates generated by the speculative execution process for generating the first question candidates, and outputs output data to present the second question to the user. In the speculative execution process for generating question candidates, a prompt is generated to cause a natural language processing program to generate question candidates that take into account the answers included in the received stream data, the prompt is input to the natural language processing program, and the question candidates generated by the natural language processing program are stored in a storage medium.
Owner:HITACHI SOFTWARE ENG

Computer System and Non-Transitory Computer-Readable Storage Medium

In order to flexibly present an additional question in response to an answer from a user, a computer system outputs output data for presenting a first question to a user, sequentially receives first stream data including an answer to the first question from the user, executes speculative execution processing for question candidate generation at least one time during reception of the first stream data, generates a second question based on a question candidate generated by the speculative execution processing for first question candidate generation when input of the answer to the first question is ended, and outputs output data for presenting the second question to the user. In the speculative execution processing for the question candidate generation, a prompt that causes a natural language processing program to execute generation of the question candidate in consideration of the answer included in the stream data received is generated, the prompt is input to the natural language processing program, and the question candidate generated by the natural language processing program is stored in a storage medium.
Owner:HITACHI SOFTWARE ENG

Device, method and system for detecting a misprediction of an instruction execution

Techniques and architectures for determining a target of a branch instruction. In an embodiment, a processor core detects a fall through event wherein multiple fetched instructions comprise one or more branch instructions. Based on the fall through event, a repository is provided with respective branch information for each of the one or more branch instructions. The repository functions as a cache that is available to an evaluation circuit at an instruction fetch stage of the processor core. Branch information at the repository is accessible to facilitate a relatively early identification of an instruction as being of a branch instruction type. In another embodiment, the early identification enables re-steering of a speculative execution sequence.
Owner:INTEL CORP

A Speculative Execution Scheduling Method for Heterogeneous MapReduce Clusters Based on Reinforcement Learning

The present invention relates to a speculative execution scheduling method for heterogeneous MapReduce clusters based on reinforcement learning, belonging to the field of big data processing. The present invention adopts a method for dynamically updating node weights based on Q-learning reinforcement learning, and realizes the adaptive adjustment of node weights based on historical information, effectively improving the estimation accuracy of the remaining running time of tasks; for the discrimination of whether a straggler is migrated, two conditions, namely the backup task ratio constraint and the running time constraint after migration, need to be satisfied simultaneously before the straggler can start the backup task; at the same time, the fast nodes of map tasks and the fast nodes of reduce tasks are combined, which improves the resource utilization rate of heterogeneous MapReduce clusters. The simulation test results based on typical data sets show that, compared with the existing algorithms, the algorithm proposed in this paper significantly improves the processing efficiency for large-scale data.
Owner:BEIJING INST OF COMP TECH & APPL

High-performance cache fault tolerance method with speculative execution mechanism

The invention discloses a high-performance cache fault tolerance method with a speculative execution mechanism. According to the method, a cache controller, a check code encoder, a check code decoder and an instruction submission unit check determiner are arranged in a cache mechanism; in a cache access stage, a cache controller adopts a speculative execution mechanism for data, assumes a successfully matched data field in a cache line as correct to-be-read data, and sends successfully matched cache line data to an instruction submission unit verification determiner; and in an instruction submission stage, an instruction submission unit verification determiner judges the correctness of speculative execution based on the error detection and correction information of the check code, and if an error is detected, a rollback recovery mechanism is triggered to recover the execution state. According to the method, check code coding and decoding logic is shifted out of a critical path of cache access through a speculation execution mechanism, so that the performance of a processor is improved; meanwhile, according to the method, through rollback operation, high performance is kept, and meanwhile the accuracy of data is fully guaranteed.
Owner:CHINA ACADEMY OF SPACE TECHNOLOGY

A disorderly speculative execution launch queue and electronic device

The present invention proposes an out-of-order speculative execution emission queue and electronic device, comprising a write buffer, an overflow area, a main entry area, and a first selector. The write buffer, overflow area, and main entry area are all connected to the first selector, and the write buffer and overflow area are all connected to the main entry area, and the write buffer is also connected to the overflow area. After providing a first type of instruction to the first selector, or when the first direct transmission condition is not met, if the number of valid instructions remaining in the write buffer is greater than or equal to 1, the write buffer is used to send the valid instructions therein to the overflow area and / or the main entry area according to a preset priority rule. The emission queue is divided into three areas: the write buffer, the overflow area, and the main entry area. Each area only needs to find the corresponding instruction as output according to the instruction age, which greatly reduces the logical depth of the entire emission queue age matrix, thereby reducing the difficulty of timing convergence of the emission queue.
Owner:CIX TECH (SHANGHAI) CO LTD +1

Branch prediction system and method

The invention discloses a branch prediction system and method, and the system comprises a main branch predictor, a decoupling queue, and a branch instruction stream processing unit, and the main branch predictor comprises a branch jump target buffer zone subsystem and a branch jump direction prediction subsystem. The decoupling queue generates entries including a PC list and branch data, and the branch instruction stream information processing unit includes a difficult-to-predict branch instruction buffer area and a branch instruction stream information buffer area; the main branch predictor predicts a branch instruction, and a generated prediction result is input into the decoupling queue; and the branch instruction stream processing unit receives the branch analysis information, performs identification and instruction stream analysis on the branch difficult to predict, and feeds back an analysis result to the main branch predictor. According to the method, the problems that the delay of the branch in the speculative execution dependency chain analysis method is difficult to predict, the accuracy is insufficient and the hardware resource overhead is large are solved.
Owner:NANJING YINGQI INTELLIGENT TECH CO LTD

A method of value prediction for a sequential execution processor

This invention belongs to the field of processor architecture and microarchitecture optimization technology, and relates to a value prediction method for sequential execution processors. It includes: constructing a step value prediction table; recording the most recent true result after the first execution of the target instruction, and forming step information after subsequent executions; generating a predicted value based on the most recent true result and step information when encountering the same static instruction later; writing the predicted value into the corresponding scoreboard entry and propagating it to dependent instructions during the instruction issue phase; writing the true result back to the scoreboard and comparing it with the predicted value after the target instruction is executed; maintaining the existing speculative execution result when the prediction is correct, and pausing front-end instruction fetching, preserving the scoreboard state, and performing a partial reissue recovery on the affected instructions when the prediction is incorrect, before resuming pipeline progression. This invention, while ensuring execution correctness, eliminates some waiting caused by real data dependencies in advance, improving the processor's ability to utilize potential instruction-level parallelism.
Owner:CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY

Parallel transaction processing method, system and computer storage medium for blockchain

The present application relates to a parallel transaction processing method, system and computer storage medium of blockchain, one aspect is to set X consensus nodes as master nodes, and distribute the additional workload bound to the master nodes to X copies instead of one copy, and to reduce the workload of the master nodes by shunting; another aspect is to divide the transaction request into Y partitions, and distribute them to Z master nodes to realize parallelism in the request transaction stage; both combine workflow consensus -> execution and workflow execution -> consensus -> verification, partition the request, and perform the workflow in parallel, the nodes call the consensus protocol for part of the partitioned transaction request to agree on the order of the transaction, and the normal nodes perform speculative execution of the client transaction in parallel, when there is inconsistency, only the transaction request guided by one master node restarts the execution of the consensus protocol and the normal nodes simultaneously perform speculative transactions, which greatly improves the workflow speed and response speed, and reduces the task amount of a single master node.
Owner:HUNAN TIAN HE GUO YUN TECH CO LTD

Prefetcher with improved stability and noise reduction

Prefetchers can have improved ability to make predictions in the presence of speculative execution. The prefetcher can maintain a training queue of addresses and can predict an address for a prefetch request based on the addresses in the training queue. The addresses in the training queue can be marked to distinguish addresses associated with speculatively-executed instructions from addresses associated with non-speculatively-executed instructions. If the speculation is determined to be incorrect, the prefetcher can remove the speculative addresses from the training queue, and if the speculation is determined to be correct, the prefetcher can remove the marking from the addresses associated with speculatively-executed instructions. The prefetcher can also pause prefetching during speculative execution.
Owner:SIFIVE INC

Security calculation method and system for side channel resistance based on data marking

A security calculation method and system for side channel resistance based on data marking is used to defend against a microarchitectural side-channel attack (SCA), particularly a transient-execution attack. The method is data-centric, and based on data marking. Through the data marking (page table entry (PTE) marking and instruction marking), and a delayed update mechanism dependent on the data marking, the method ensures that any subsequent instructions dependent on a memory-sensitive instruction are not executed in speculative execution, thereby preventing confidential data from being leaked when a conditional branch outcome is unknown, effectively resisting the SCA, improving the security, and minimizing the impact on central processing unit (CPU) performance.
Owner:NANHU LAB

Consistent Speculation of Pointer Authentication

In an embodiment, a processor includes hardware circuitry which may be used to authenticate instruction operands. The processor may execute instructions that perform operand authentication both speculatively and non-speculatively. During speculative execution of such instructions, the processor may execute authentication such that no differences in observable state of the processor, relative to authentication result, are detectable via a side channel. During speculative execution, a result of authentication may be deferred until speculative execution of the instruction, and additional instructions, may be completed. Upon resolution of a condition that indicates acceptance of the speculative execution, a speculative execution result may cause a processor exception and stalling of execution at the instruction to be performed.
Owner:APPLE INC

Securing computing systems against microarchitectural replay attacks

A system and method for mitigating micro-architectural replay attacks in a processing system by delaying speculative execution on the processing system of a set of processor instructions upon detection that the set of processor instructions are part of a micro-architectural replay attack by detecting repeating speculative execution of the set of processor instructions interleaved with misspeculation and squashing of the set of processor instructions.
Owner:ETA SCALE AB

Apparatus and method for controlling allocation of information into cache storage

An apparatus and method are provided for controlling allocation of information into a cache storage. The apparatus has processing circuitry for executing instructions and for allowing speculative execution of one or more of those instructions. The invention also provides a cache storage having a plurality of entries for storing information for reference by the processing circuitry, and cache control circuitry for controlling the cache storage, the cache control circuitry including a speculative allocation tracker having a plurality of tracking entries. The cache control circuitry is responsive to a speculative request associated with the speculative execution requiring allocation of identified information into a given entry of the cache storage, to allocate a tracking entry in the speculative allocation tracker for the speculative request prior to allowing allocation of the identified information into the given entry of the cache storage. The allocated tracking entry is employed to maintain recovery information sufficient to enable the given entry to be recovered to an initial state existing prior to the identified information being allocated into the given entry. The cache control circuitry is also responsive to detecting a false speculation condition with respect to the speculative request, to employ the recovery information maintained in the allocated tracking entry for the speculative request to recover the given entry in the cache storage to an initial state. Such methods can provide robust protection against speculative based cache timing side-channel attacks while mitigating performance and / or power consumption issues associated with known techniques.
Owner:ARM LTD

A method and system for speculative execution acceleration of robot control based on physical drives

This invention provides a physics-driven method and system for accelerating speculative execution in robot control, relating to the fields of artificial intelligence and robot control technology. The method first performs data-driven task criticality annotation based on an offline dataset, obtaining task criticality labels. Then, it trains a multi-task draft model with residual dynamics prediction capabilities. The draft model is a lightweight multi-task network that outputs motion trajectory prediction, residual latent dynamics prediction, and stepwise criticality score prediction. During the inference phase of robot task execution, after initial inference by the VLA model, the draft model generates draft actions and performs dual-gated verification. Based on the verification results, the draft model and the VLA model are used alternately. This invention does not require intrusive modifications to the VLA model architecture and is applicable to various VLA model architectures and robot manipulation tasks of varying complexity.
Owner:SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI

Accelerating diffusion model inference using speculative execution

A computer-implemented method, system, and computer program code for generating a data item represented by a plurality of continuous-valued elements, e.g. using a diffusion model. The method generates a draft sequence, of draft representations of the data item, and determines, e.g. in parallel, whether to accept or reject these. If a draft representation is rejected a new representation is determined. In implementations the new representation is computed deterministically as a function of the rejected draft representation.
Owner:GDM HOLDING LLC