Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

217 results about "Load instruction" patented technology

Write-after-read conflict prediction method, device and equipment

The invention provides a write-after-read conflict prediction method, device and equipment, which are applied to the technical field of data processing, and the method comprises the following steps: reading a to-be-transmitted loading instruction; the loading instruction is located in a loading queue of the out-of-order execution processor. Determining an index value of the loading instruction; the index value is used for quickly positioning a conflict record matched with the loading instruction in a historical conflict table of the out-of-order execution processor. The conflict record comprises a first storage instruction corresponding to the loading instruction. And when it is determined that the conflict record corresponding to the index value exists in the historical conflict table, setting the loading instruction to be in a suspended state. And after the execution of the first storage instruction is finished, executing the loading instruction. The problem that the execution efficiency of a processor is affected due to pipeline emptying and pause caused by write conflicts after reading can be solved.
Owner:BEIJING YIHUA CLOUD NETWORK TECH CO LTD

Method and system for implementing memory access optimization of processor with multi-channel concurrent instructions

The invention provides a processor memory access optimization implementation method and system with multi-channel concurrent instructions, and relates to the technical field of processor integrated circuits, and the method comprises the steps: obtaining a Load instruction and a Store instruction to be executed; distributing a plurality of memory access instructions to each memory access channel which is independently processed in parallel, wherein each memory access channel comprises a Load memory access channel and a Store memory access channel; an independent Load queue and an independent RAW check queue are arranged in the Load memory access channel, and an independent Store queue, an independent STD queue and an independent Store Buffer queue are arranged in the Store memory access channel; the Load instruction is sequentially subjected to an address generation stage, a TLB parallel query stage, an L1Cache access stage, a RAW check and data forward push stage and a data write-back stage in the Load memory access channel; according to the method, address and data separation processing is carried out on a Store instruction, an address part enters a Store queue, a data part enters an STD queue, then merging is carried out, multiple Store operations on the same Cache Line are merged into one-time writing, the merged data are written into a Store Buffer queue, and then the Store Buffer queue submits and caches in a unified mode.
Owner:SHANDONG LINGNENG ELECTRONIC TECH CO LTD

Memory access optimization method and device for ensemble communication operator and computing equipment

The invention provides a memory access optimization method and device for an aggregate communication operator and computing equipment, and relates to the technical field of graphics processor computing. The method comprises the following steps: identifying a memory access instruction sequence in a to-be-optimized set communication operator, wherein the memory access instruction sequence comprises a standard loading instruction and a standard storage instruction; replacing the standard loading instruction with a batch loading instruction, and replacing the standard storage instruction with a batch storage instruction; the batch loading instruction and the batch storage instruction are configured to execute batch operation on the continuous memory address space at the thread beam level and finally execute the batch loading instruction and the batch storage instruction, so that the emission number of memory access instructions can be reduced, the scheduling burden of a thread beam scheduler is relieved, and the scheduling efficiency is improved. And the execution efficiency and the bandwidth utilization rate of the ensemble communication operator in the distributed training of the graphics processor are effectively improved.
Owner:SHANGHAI BIREN TECH CO LTD

Data processing device and method, processor and chip

The invention relates to a data processing device and method, a processor and a chip, and relates to the technical field of computers, the data processing device comprises a processing module and a driving module; the processing module obtains address information of each thread in a first instruction when judging that the instruction to be processed is the first instruction, the first instruction is a multi-thread parallel loading instruction, and the address information of different threads is mutually independent; the driving module determines a parallel loading result according to address information of each thread in the first instruction, the address information is used for determining an off-chip address and an on-chip address of each thread, and the thread is used for loading data at a position indicated by the off-chip address to a position indicated by the on-chip address. According to the embodiment of the invention, multi-thread parallel execution of different data accesses can be realized, and the calculation efficiency is remarkably improved.
Owner:MOORE THREADS TECH CO LTD

Pixel processing method, electronic device, machine readable medium, and program product

The embodiment of the invention provides a primitive processing method, electronic equipment, a machine readable medium and a program product, and relates to the technical field of graphic processing.The method comprises the steps that an equation coefficient of a cutting plane equation of a target cutting plane is obtained, and vertex coordinates of vertexes in a to-be-cut primitive are obtained; the target cutting plane is any one currently enabled cutting plane. Generating a target loading instruction for the equation coefficient and the vertex coordinate; the target loading instruction is used for instructing the GPU to load the equation coefficient and the vertex coordinate into a preset register. Generating a target calculation instruction based on a preset register; the target calculation instruction is used for instructing the GPU to calculate a target distance between the vertex and the target cutting plane based on an equation coefficient in a preset register and the vertex coordinates. And executing the target loading instruction and the target calculation instruction by an original calculation unit in the GPU to obtain a target distance. Therefore, the circuit overhead and the implementation cost can be reduced.
Owner:LOONGSON TECH CORP

Predictive storage to load forwarding

A method for performing predictive store-to-load forwarding on a processor includes receiving a load instruction; executing a predictive storage-to-load forwarding process, including obtaining a predicted storage queue id from a prediction table based on an address of a load instruction, searching a value from the storage queue based on the storage queue id, and executing a parallel verification process on the predicted storage queue id; determining that the parallel verification process is successful; and in response, reading a value for the load instruction from the store queue.
Owner:GOOGLE LLC

Upgrading method and device for metadata nodes in distributed cluster, equipment and medium

The invention relates to an upgrading method and device for metadata nodes in a distributed cluster, equipment and a medium, and the method comprises the steps: responding to an upgrading instruction of a to-be-upgraded metadata node, and sending a first loading instruction to a target metadata node; the first loading instruction is used for indicating the target metadata node to load read-only segment metadata in the to-be-upgraded metadata node according to a first loading rate; whether loading of the read-only segment metadata is completed is determined, and a second loading instruction is sent to the target metadata node under the condition that loading is determined to be completed; the second loading instruction is used for indicating the target metadata node to load readable-writable segment metadata in the to-be-upgraded metadata node according to a second loading rate, and the first loading rate is smaller than the second loading rate; and after receiving a loading completion notification of the readable and writable segment metadata, upgrading the to-be-upgraded metadata node. According to the method and the device, smooth upgrading of the metadata node can be realized under the condition that the service performance is not influenced, and the experience feeling of a user is improved.
Owner:CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1

Store-to-load forwarding correctness checks using physical address proxies stored in load queue entries

A microprocessor includes a load / store unit that performs store-to-load forwarding, a PIPT L2 set-associative cache, a store queue having store entries, and a load queue having load entries. Each L2 entry is uniquely identified by a set index and a way. Each store / load entry holds, for an associated store / load instruction, a store / load physical address proxy (PAP) for a store / load physical memory line address (PMLA). The store / load PAP specifies the set index and the way of the L2 entry into which a cache line specified by the store / load PMLA is allocated. Each load entry also holds associated load instruction store-to-load forwarding information. The load / store unit compares the store PAP with the load PAP of each valid load entry whose associated load instruction is younger in program order than the store instruction and uses the comparison and associated forwarding information to check store-to-load forwarding correctness with respect to each younger load instruction.
Owner:VENTANA MICRO SYSTEMS INC

Production change control method and equipment for mixed line production and medium

The embodiment of the invention discloses a production change control method and device for mixed line production and a medium, and relates to the technical field of production control, and the method comprises the steps: collecting real-time production index data corresponding to a current production product in a current production line in real time under the triggering of a production change instruction issued by a production management system, obtaining a target process parameter set corresponding to the target product change product; according to the real-time production index data and the production change instruction, a production change state vector containing a production change stage label is determined, and the production change stage label comprises a production change preparation stage, a production change pre-adjustment stage and an execution stage; and based on the production change state vector and the target process parameter set, through a preset reinforcement learning agent, determining a time sequence production change action sequence, converting the time sequence production change action sequence into an equipment control instruction for instruction distribution, and realizing production change control. The sequential production change action sequence comprises a pre-adjustment action instruction, a parallel scheduling instruction and a parameter loading instruction.
Owner:INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD

RISC-V vector data loading instruction processing method and processor

The invention provides an RISC-V vector data loading instruction processing method and a processor, and relates to the technical field of processors. When a vector loading instruction is transmitted, memory access address calculation and memory access can be started in advance only by making a source operand ready, and waiting for an old value of a target vector register is not needed, so that transmission blocking caused by dependence of the old value is avoided; by temporarily storing loaded data in a vector data buffer area, memory access operation and subsequent merging write back are decoupled, and memory access delay and old value ready waiting time are overlapped; and when the instruction becomes the oldest to-be-written-back instruction and an old value is ready, the merging unit merges new and old data element by element according to the mask and vector configuration to ensure correct implementation of RISC-V vector extension complex semantics. On the premise of not influencing the function correctness, the memory access delay is effectively hidden, the instruction level parallelism is improved, and the processing efficiency of the vector loading instruction is remarkably improved.
Owner:YIHUA TECHNOLOGY (BEIJING) CO LTD

Processor context storage and recovery method based on shadow register

The invention discloses a processor context storage and recovery method based on a shadow register. The main purpose of the invention is to realize more bottom-layer flexible processor context switching under complex processor scenes such as a multi-privilege mode based on a shadow register structure. By introducing a shadow register structure, when the context of the processor is saved and recovered, logic registers at different positions in a saving and loading instruction are mapped into different physical register files, so that the damage of a context saving program to an original site is avoided. Compared with a classic context storage and recovery method, the method can flexibly support the context switching of the processor in different modes, simplifies the context storage of a switching program, and improves the switching efficiency while guaranteeing the context integrity of the processor.
Owner:BEIHANG UNIV

Multi-transmitting channel architecture optimization method and system based on superscalar processor

The invention provides a multi-transmitting-channel architecture optimization method and system based on a superscalar processor, and belongs to the technical field of processors, and the method comprises the steps: reading an instruction by adopting a data capturing and transmitting mechanism, judging whether a register is ready, introducing a feedforward data cache region, optimizing a data path, reading and obtaining a value of the register, and obtaining a multi-transmitting-channel architecture of the superscalar processor; the read values and instructions are stored in a transmitting queue; loading storage instruction transmission optimization is carried out on the instruction, a storage instruction is divided into a storage address and storage data, the storage data is allowed to be transmitted in advance, and table items are established in a loading storage unit; and reconstructing the transmitting queues, reconstructing the LSIQ transmitting queues into a Load instruction queue and a Store instruction queue, arbitrating the transmitting queues respectively, and waiting for execution. According to the method, the access pressure of the register file can be effectively relieved, the data feedforward efficiency is improved, RAW address conflicts are reduced, and the instruction transmitting efficiency and the overall processor performance are improved.
Owner:SHANDONG LINGNENG ELECTRONIC TECH CO LTD

Global memory disambiguation for a parallel architecture with compute slices

Techniques for checking memory operations are disclosed. A processing unit is accessed, comprising compute slices, control unit, local memory disambiguation units (LMDUs), and a global MDU (GMDU). Each slice includes an execution unit and is coupled to successor and predecessor slices. Each slice is coupled to an LMDU. Each LMDU is coupled to the GMDU. A first slice executes a first slice task. The task includes a load instruction and address. The slice issues the load to an LMDU, saving load information in a memory operation table (MOT). For a not fully serviced load instruction, the LMDU sends the load information to the GMDU, storing load information in a global MOT (GMOT). The GMOT detects address aliasing between the load address and a previously issued address saved in the GMOT. The GMOT forwards memory information from previously issued memory instructions to the MOT to satisfy the load instruction.
Owner:ASCENIUM INC

Apparatus and method for hiding vector load latency in a time-based vector coprocessor

A processor includes a time counter and a vector coprocessor for executing vector instructions for statically dispatching vector instructions with preset execution times based on a write time of a register in a coprocessor register scoreboard and a time counter provided to a vector execution pipeline. The processor also provides a method for hiding the latency of the vector load instructions.
Owner:SIMPLEX MICRO INC

Two-order loading value prediction design method and system based on path historical information

The invention discloses a two-order loading value prediction design system based on path history information. The two-order loading value prediction design system comprises a path history register, a loading instruction address prediction table, a loading address reservation station, a value prediction table and a speculation conflict detection table, the invention also discloses a two-order loading value prediction design method based on path historical information. The method comprises the following steps: S1, predicting a memory address possibly accessed by a loading instruction through a loading instruction address prediction table; s2, storing the predicted memory address into a loading address reservation station; s3, when the load assembly line is idle, entering the load assembly line in advance, and accessing the data cache by using the memory address in the loading address reservation station; and S4, when the memory access instruction reaches a renaming stage, checking whether a predicted value exists in the value prediction table or not. The value of the loading instruction is predicted through the path information, and the prediction process can be more accurate.
Owner:JIANGSU HUACHUANG MICROSYSTEM CO LTD

Supercritical unit low-load deep peak shaving system and method

PendingCN121965572ASolve the problem of lack of overall collaborative strategySolve the problem of not being able to fully utilize your capabilitiesAc networks with different sources same frequencyLoad instructionLow load
The invention provides a supercritical unit low-load deep peak regulation system and method, and relates to the field of thermal generator sets, and the supercritical unit low-load deep peak regulation system comprises an instruction analysis and pre-control module, an energy redistribution module, a combustion steady-state module, a parameter cooperation module and a state backtracking and self-gain module. According to the application, through the arranged instruction parsing and pre-control module, the power grid load instruction is parsed and converted into a set of accurate equipment action time sequence list, the instruction feature recognition unit in the instruction parsing and pre-control module performs mode recognition on the load instruction, and the strategy generation unit calls the preset strategy template according to the mode, so that the power grid load instruction is analyzed and converted into the accurate equipment action time sequence list. Action instructions which face all execution units of the whole system and are arranged in a wrong sequence on a time axis are generated, and conversion from passive response to active pre-control is achieved.
Owner:HUADIAN WEIFANG POWER GENERATION CO LTD

Page loading method, device, electronic equipment and system based on page cache

The invention relates to the technical field of resource cache management, and discloses a page cache-based page loading method and device, electronic equipment and system.The method comprises the steps that after it is detected that a network agent tool meets a preset operation condition, a network request is intercepted through a capture event monitor of the network agent tool, and the network request is loaded to the electronic equipment; if the network request does not meet the special processing condition, judging whether the network request contains a target resource request, and if so, executing a resource response operation according to a resource loading strategy corresponding to the target resource request contained in the network request to obtain a first resource; otherwise, loading a second resource required by the network request through the network based on a preset network loading instruction; and displaying the target resource obtained by loading on the current page. Therefore, by implementing the method and the device, different resource loading and caching strategies can be adopted for different types of page resources, adaptive page caching is realized, the loading efficiency and stability of the page resources are improved, and the page browsing experience is improved.
Owner:SHENZHEN GREEN CONNECTION TECH CO LTD

Widening vector load instruction

A widening vector load instruction specifies at least one address operand and two or more vector destination registers, e.g. Z1 and Z2, each for specifying a vector operand having a given vector lengt
Owner:ARM LTD

Communication method between unified loading upper computer and generator controller

The invention belongs to the technical field of generator controller ground maintenance, and particularly relates to a communication method between a unified loading upper computer and a generator controller. The method comprises the following steps: S1, acquiring a communication parameter configuration file written by a user according to the model of a generator controller, an interaction state configuration file and a software upgrading program file for generator control; s2, the upper computer sends a serial loading instruction through serial port communication; s3, after the generator controller is reset, running a serial loading program; s4, sending the software upgrading program file to a serial loading program of the generator controller through a serial port; s5, verifying the received software upgrading program file, and writing the software upgrading program file into Flash of the generator controller after verification is successful; and S6, the upper computer sends an instruction for clearing the serial loading mark, and the generator controller clears the serial loading mark. The maintenance cost is remarkably reduced, and the overall working efficiency is improved.
Owner:SHAANXI AVIATION ELECTRICAL

Low-overhead processor cache micro-architecture defense method and device and computer equipment

The invention relates to a low-overhead processor cache micro-architecture defense method and device and computer device.The low-overhead processor cache micro-architecture defense method comprises the steps that when a missing state keeping register module takes out a loading instruction waiting for data from a replay queue, a target branch mask of the loading instruction is obtained; the replay queue is used for storing an instruction which is stagnated due to miss of the cache; judging whether the loading instruction is in a speculative execution state or not according to the target branch mask; and when the loading instruction is in the speculative execution state, stopping a cache write-in operation corresponding to the loading instruction. Through the method and the device, the problem of sensitive information leakage caused by incapability of defending against the cache side channel attack is solved, defending against the cache side channel attack is realized, and sensitive information leakage is prevented.
Owner:HANGZHOU HIGH-TECH ZONE (BINJIANG) INSTITUTE OF BLOCKCHAIN & DATA SECURITY +1

Data processing method, solid state disk, system and storage medium

The invention discloses a data processing method, a solid state disk, a system and a storage medium, and relates to the technical field of solid state disk storage, and the data processing method comprises the steps that it is detected that an NVMe PCIe link configures an instruction register, and an NVMe KV instruction is loaded from a host memory; based on the NVMe KV instruction, obtaining a key value pair configured by a host according to to-be-operated data and an operation code corresponding to a host demand operation type; the KV operation corresponding to the operation code is executed according to the key value pair, and the configuration operation of the key value pair is unloaded to the host to be executed, so that the computing resource occupation amount of the SSD is reduced, and the service quality of the SSD is improved.
Owner:ZTE CORP

Method and device for verifying memory consistency of multi-core processor

The invention relates to a multi-core processor memory consistency verification method and device, the method is executed by a memory model in a processor co-simulation verification architecture, and the method comprises the following steps: monitoring a storage instruction of each processor core in an RTL, and maintaining a storage queue for each processor core based on the storage instruction; in response to a data request initiated when the ISS executes the loading instruction, speculating a set of compliance values possibly read by the loading instruction based on the storage queue and the visible relation of the storage instructions between the maintained different processor cores; comparing the actual read data of the loading instruction obtained from the RTL with the set of compliance values; and if the actually read data is in the set of compliance values, determining that the memory consistency verification of the multi-core processor is compliant. Therefore, an RISC-V multi-core processor memory consistency verification scheme which can be integrated in a front simulation environment, supports real-time accurate verification and has high coverage is provided.
Owner:CHENGDU QUNXIN MICROELECTRONICS TECHNOLOGY CO LTD

Processor and memory access method

The invention relates to a processor and a memory access method. The processor comprises a UAV instruction sending unit, a UAV register management unit, a general register unit, a UAV data caching unit and a loading storage unit; the loading storage unit is configured to receive an SIMD instruction; reading resource information of the target UAV register from the UAV register management unit according to the address information of the target UAV register; under the condition that the SIMD instruction is determined to be a loading instruction according to the operation code, reading channel logic address information of each channel from the general register unit; under the condition that the memory addresses to be accessed by the multiple channels of the SIMD instruction are judged to be continuous, calculating the memory access address of the first channel, and calculating the memory access addresses of the other channels based on the memory access address of the first channel; and generating a memory access request based on the memory access address of each channel. According to the method, the throughput rate of the loading storage instruction of the continuous address is improved, the number of ALUs calculated by the memory address is reduced, and meanwhile, the power consumption is reduced.
Owner:GLENFLY TECH CO LTD

Control system for improving load response rate of coal-fired thermal power unit based on ACE mode

The application discloses a control system for improving load response rate of a coal-fired thermal power unit under an ACE mode, which comprises a PID controller with feedforward, a first adder, a first analog AI input module, a second analog AI input module, a first analog AO output module, a second adder, a third analog AI input module, a fourth analog AI input module and a fifth analog AI input module; the output end of the fourth analog AI input module and the output end of the fifth analog AI input module are connected with the input end of the second adder, the output end of the third analog AI input module and the output end of the second adder are connected with the input end of the first adder, and the output end of the first analog AI input module, the output end of the second analog AI input module and the output end of the first adder are connected with the input end of the PID controller with feedforward; the system can realize the purpose of fast response of the coal-fired thermal power unit to load instruction change.
Owner:XIAN THERMAL POWER RES INST CO LTD

Load automatic adjustment system and method suitable for electric power spot transaction

The invention relates to the technical field of power system automation and control, in particular to an automatic load adjusting system and method suitable for power spot transaction. The invention relates to an automatic load adjustment system suitable for electric power spot transaction and a power generation spot transaction platform system, which are used for providing real-time clearing results and pre-clearing results including a plurality of time points in the future. The data acquisition and release system is used for converting the clearing result into a power value per minute of a time sequence; the centralized control center computer monitoring system is used for generating a load set value; the power station control unit is used for automatically adjusting the output of the corresponding power generation unit; and the security I area communication gateway is used for realizing data communication with the data acquisition and release system in the security II area. The electric power spot transaction platform is deeply integrated with the power plant monitoring system and the AGC system, so that the whole process automation from acquisition of clearing results to issuing of load instructions and final execution is realized, the labor intensity of operators is greatly reduced, and the overall operation efficiency is improved.
Owner:SICHUAN HUANENG BAOXINGHE HYDROPOWER CO LTD

Window-based memory dependency predictor

Methods and systems for out of order processing are disclosed herein. A disclosed method for out-of-order instruction processing includes identifying a first load instruction for processing, determining, based on the first load instruction, whether a store instruction window has been established for the first load instruction, confirming, when the store instruction window has been established, that one or more store instructions, within the store instruction window, that the first load instruction is dependent upon have been executed, and executing the first load instruction after execution of the one or more store instructions has been confirmed.
Owner:TENSTORRENT USA INC

Memory dependency prediction method and device based on branch predictor and processor architecture

The invention relates to the technical field of processors, and provides a memory dependency prediction method and device based on a branch predictor and a processor architecture, and the method comprises the steps: selecting a prediction block from an instruction fetching queue; determining a target load instruction in the prediction block; and predicting memory dependence information of the target load instruction according to global branch historical information of the branch direction predictor. According to the embodiment of the invention, the prediction precision of memory dependency prediction can be improved, the memory dependency prediction cost can be reduced, and the method is suitable for a modern processor which is high in performance, performs out-of-order execution and is provided with a decoupling front-end framework.
Owner:CHENGDU QUNXIN MICROELECTRONICS TECHNOLOGY CO LTD

Zero-delay real-time hybrid test method and system based on EtherCAT-PDO mapping and application

The invention relates to a zero-delay real-time hybrid test method and system based on EtherCAT-PDO mapping and application, and belongs to the technical field of structural engineering tests. Comprising the following steps: establishing an EtherCAT master station manager, and uniformly managing EtherCAT network resources by adopting a single example mode; the EtherCAT slave station equipment is configured to comprise a bus coupler, an analog output module and an analog input module; a straight-through zero-delay data exchange mechanism is set, and data are directly written into a PDO mapping area to realize lowest-delay transmission; a PDO data mapping mechanism is established, and real-time transmission from a numerical substructure calculation result to a physical substructure loading instruction is achieved; a multi-substructure parallel communication architecture is established, and a plurality of numerical substructures and physical substructures are supported to participate in a hybrid test at the same time. According to the method, the technical problems of data loss, overlarge delay, insufficient expansibility and the like in the existing real-time hybrid test communication method are effectively solved.
Owner:BEIJING LINGJIE CHENGCHUAN TECHNOLOGY CO LTD

Load value prediction based elimination of execution of instructions that compute constant values

A method of an aspect includes predicting a value that a load instruction of a first iteration of a loop would load and executing a plurality of instructions occurring after the load instruction in the first iteration of the loop to generate a plurality of results. Each of the plurality of results depends only on the value, one or more constant values, values derived from the value and / or the one or more constant values, or any combination thereof. The method also includes producing the plurality of results for the plurality of instructions during each of one or more iterations of the loop after the first iteration without re-executing the plurality of instructions. Other methods, apparatus, and systems are also disclosed.
Owner:INTEL CORP

Variable precision and mixed-type representation of multiple layers in a network

In one example, an apparatus comprises a plurality of execution units including at least a first type of execution unit and a second type of execution unit and logic, at least partially including hardware logic, to: expose an embedded projection operation in at least one of a load instruction or a store instruction; determine a target precision level for the projection operation; and load the projection operation at the target precision level. Other embodiments are also disclosed and claimed.
Owner:INTEL CORP