Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

108 results about "Microarchitecture" patented technology

In computer engineering, microarchitecture, also called computer organization and sometimes abbreviated as µarch or uarch, is the way a given instruction set architecture (ISA) is implemented in a particular processor. A given ISA may be implemented with different microarchitectures; implementations may vary due to different goals of a given design or due to shifts in technology.

Node task migration and scheduling system based on digital twinning

The invention provides a node task migration and scheduling system based on digital twinning, and relates to the technical field of computer system structures and data processing. Computing resource interference in a multi-tenant sharing environment is quantized by sensing a cross-tenant noise coefficient and a state synchronization complexity entropy in a node micro-architecture; extracting track features of the mobile terminal, calculating spatial discrete variance, generating a self-adaptive migration decision hysteresis factor, and converting the migration decision hysteresis factor into decision damping to inhibit invalid high-frequency reciprocating migration; constructing a digital twin sandbox before physical cutover, cooperating with a chaos scene injection engine to inject a composite fault operator into a bottom layer, and performing actuarial calculation on service continuity retention after risk adjustment by using a fidelity integrator; and finally, a bottom layer controller is linked through safety baseline comparison to execute physical flow switching. According to the method, network boundary deduction is completed on the premise that physical bandwidth is not consumed, the interruption risk caused by state hard switching is avoided, and smooth transition of stateful services is effectively guaranteed.
Owner:XIAMEN KUAIKUAI NETWORK TECH CO LTD

H2P branch prediction chip circuit architecture based on BrPerceptron

The invention belongs to the field of micro-architecture design of an integrated circuit processor. The invention provides an H2P branch prediction chip circuit architecture based on BrPerceptron, a variable operation track of a program is constructed based on a context-sensitive variable tracking mechanism, a program execution path topology is constructed based on a basic block division algorithm according to the variable operation track, multi-dimensional correlation detection is performed according to the program execution path topology to judge whether an H2P branch exists or not, and if yes, the H2P branch is predicted. And when the H2P branch instruction is judged to be the H2P branch instruction, enabling the BrPerceptron predictor to output, performing multi-stage feature fusion and outputting a prediction result, and realizing high-precision prediction for the H2P branch instruction.
Owner:SHANDONG LINGNENG ELECTRONIC TECH CO LTD

Data-credible-oriented vehicle-mounted integrated security computing system and credible construction method

The invention relates to the technical field of confidential computing, and discloses a data credibility-oriented vehicle-mounted integrated security computing system and credibility construction method, which comprises the following steps of: establishing a static trust chain by using a hardware trust root, shielding external interruption in a credible execution environment, and controlling a performance monitoring unit; synchronously collecting real-time hardware micro-architecture events of the key business algorithm to generate a runtime feature vector; according to the method, the instruction stream during operation is anchored by utilizing the physical characteristics of the micro-architecture, the hijacking attack of the control stream is accurately identified in a ciphertext environment, the hijacking attack of the control stream is accurately identified, the hijacking attack of the control stream is accurately identified, the hijacking attack is accurately identified, and the hijacking attack of the control stream is accurately identified. Strong coupling verification of a calculation result and an execution behavior is realized, and logic safety of a vehicle-mounted platform is ensured.
Owner:SHANGHAI JUPO TECH CO LTD

Neural network parallel scheduling-oriented single-instruction multi-thread processor micro-architecture device

A single-instruction multi-thread processor micro-architecture device oriented to neural network parallel scheduling comprises a front-end instruction fetching module, an instruction cache module, a decoding module, an arithmetic logic operation unit, a multiplication and division module, a memory access unit and a data cache module, and the micro-architecture device allocates a unique thread number for each thread. Threads are organized into thread groups, and in each period, the fair alternate arbiter selects an instruction from an instruction buffer of the thread group and sends the instruction to a subsequent decoding stage. The micro-architecture not only solves challenges faced by end-side equipment when processing high-performance calculation requirements such as neural network reasoning, but also provides an effective solution capable of reducing energy consumption and improving calculation efficiency through an innovative architecture design. The method is of great significance in promoting development of end-side AI application.
Owner:XI AN JIAOTONG UNIV

Branch prediction unit power consumption management method and device

The invention discloses a branch prediction unit power consumption management method and device, and the method comprises the steps: setting a multi-stage branch predictor and a low-power-consumption control logic coupled with the multi-stage branch predictor in a processor core micro-architecture, and enabling the multi-stage branch predictor to comprise three stages of predictors which are sequentially accessed in an assembly line, the hardware resource consumption and the prediction accuracy of each predictor are gradually increased from front to back; monitoring the prediction condition of each predictor in real time through low-power-consumption control logic, and closing all predictors behind the first predictor meeting the prediction requirement within a first preset time according to a front-to-back sequence; branch jump addresses are divided into high-order address fields and low-order address fields by BTBs in the multiple levels of predictors except the first-level predictor to be stored in different storage blocks, and label registers used for taking the high-order addresses as label indexes are arranged. The power consumption unit of each predictor can be finely controlled, the energy consumption is greatly reduced, system-level cooperation is not needed, and the universality is high.
Owner:NANJING YINGQI INTELLIGENT TECH CO LTD

Cooperative evaluation method for processor micro-architecture performance upper limit exploration and related device

ActiveCN121255595BResource allocationHardware monitoringSpeculative executionProcessor model
The embodiment of the application discloses a kind of collaborative evaluation method for exploring the upper limit of processor micro-architecture performance and related device, the method includes that various idealized micro-architecture components are constructed according to the collaborative process of ideal model by real processor model, processor micro-architecture performance is collaboratively evaluated, and collaborative process includes: real processor model sends instruction query and information request to ideal model, and the request carries the memory address of dynamic instruction;Ideal model executes dynamic instruction according to memory address and obtains target value in normal working state, and target value is returned as the response of the request to real processor model;Real processor model executes dynamic instruction and obtains actual value in execution phase;Target value and actual value are verified in submission phase, and verification result is obtained.Using the embodiment of the application, the theoretical performance upper limit of value prediction and other speculative execution techniques can be accurately quantified, and the real performance bottleneck of the entire processor system integrated with the technique can be identified.
Owner:UNIV OF SCI & TECH OF CHINA

Anti-quantum method and system based on stateless signature and execution isomorphism

This invention discloses a quantum-resistant method and system based on stateless signatures and execution isomorphism, belonging to the field of quantum-resistant cryptography. First, the sender generates an mKEM broadcast payload based on a modulus error rounding algorithm and signs it using a stateless hash signature algorithm. The receiver performs microsecond-level verification of the signature at the network card driver layer; if verification fails, the signature is silently discarded. After successful verification, a dedicated PQC hardware engine decapsulates the signature, implicitly outputting a pseudo-random scrap key if verification fails. The operating system executes an I / O-aware microarchitecture with isomorphic execution, forcing subsequent processes to maintain physical isomorphism in system calls, memory accesses, and peripheral bus activities, regardless of whether the key is real or scrap. This invention eliminates the risk of private key leakage caused by cloud-based state management through stateless signatures and completely eliminates distinguishable side-channel fingerprints across the entire link through physical-level execution isomorphism, achieving system-level quantum-resistant security in high-concurrency scenarios.
Owner:BEIJING LANGKONG QUANTUM TECHNOLOGY CO LTD

Low-overhead processor cache micro-architecture defense method and device and computer equipment

PendingCN121561900APlatform integrity maintainanceSpeculative executionLoad instruction
The invention relates to a low-overhead processor cache micro-architecture defense method and device and computer device.The low-overhead processor cache micro-architecture defense method comprises the steps that when a missing state keeping register module takes out a loading instruction waiting for data from a replay queue, a target branch mask of the loading instruction is obtained; the replay queue is used for storing an instruction which is stagnated due to miss of the cache; judging whether the loading instruction is in a speculative execution state or not according to the target branch mask; and when the loading instruction is in the speculative execution state, stopping a cache write-in operation corresponding to the loading instruction. Through the method and the device, the problem of sensitive information leakage caused by incapability of defending against the cache side channel attack is solved, defending against the cache side channel attack is realized, and sensitive information leakage is prevented.
Owner:HANGZHOU HIGH-TECH ZONE (BINJIANG) INSTITUTE OF BLOCKCHAIN & DATA SECURITY +1

Combinable core particle network-oriented period precision simulator design method

The invention discloses a period precise simulator design method for a combinable core particle network, and relates to the technical field of core simulation testing, and the method comprises the steps: designing a simulation frame of a combinable core particle network simulator; wherein the simulation framework comprises a two-stage core particle network configuration unit and a multi-thread parallel simulation framework; carrying out combinable topology modeling, modular routing mechanism modeling and heterogeneous router micro-architecture modeling on the combinable core particle network; performing protocol layer protocol conversion modeling, protocol layer flow control mechanism modeling, adaptation layer retransmission mechanism modeling and physical layer electrical behavior modeling on the core particle interconnection protocol interface; based on the above design, the period precision simulator for the combinable core particle network is constructed. By designing the simulation framework, more accurate actual core particle network behaviors can be obtained; a higher simulation speed is realized through a multi-thread parallel simulation framework, multi-thread parallel accelerated simulation under a large-scale network is supported, and the simulation test efficiency is improved.
Owner:SUN YAT SEN UNIV

Heterogeneous cloud container cluster scheduling model training method and scheduling method based on topology awareness

The invention belongs to the technical field of intelligent scheduling, and particularly relates to a heterogeneous cloud container cluster scheduling model training method and scheduling method based on topology awareness, and the training method comprises the steps: obtaining a cluster topological graph; obtaining a state vector time sequence of the nodes based on the node operation data, wherein each state vector comprises a performance index representing bottom layer resource contention; obtaining a queue state matrix; processing the state vector time sequence by using a feature extractor to obtain a node initial embedding vector; processing the initial embedded vectors of all the nodes by using a graph attention network to obtain a cluster global topology vector; obtaining a joint observation state based on the cluster global topology vector and the queue state matrix; inputting the joint observation state to the Actor-Critic network to generate a scheduling strategy; the accumulated rewards are maximized through a near-end strategy optimization algorithm, and network parameters are updated; and connecting the feature extractor, the graph attention network and the decision network to form a scheduling model. And the topological relation between the nodes is globally sensed, the interference attribute of the micro-architecture is captured, and accurate scheduling is realized.
Owner:CHONGQING UNIV

A RISC-V multi-core heterogeneous platform intelligent load balancing method and system

The application provides an RISC-V multi-core heterogeneous platform intelligent load balancing method and system, and relates to the technical field of resource allocation and scheduling. Micro-architecture performance data of each processing core in the RISC-V multi-core heterogeneous platform is acquired to construct a state vector; the micro-architecture performance data comprises instruction cycle number, cache miss rate at each level and memory pause proportion; the state vector is input into a pre-trained deep Q network model to generate optimal action instructions, so that the load balancing of thread resources and core capacity is realized; the optimal action instructions comprise thread migration, thread exchange and core frequency adjustment; the pre-training of the deep Q network model is performed on a parallel computing program running on the RISC-V platform; the environment is randomly disturbed before the program runs; the action is executed through a random strategy, and state, action and reward data are collected; an offline experience dataset is constructed to perform pre-training. Dynamic, cooperative and adaptive optimization of parallel computing of the RISC-V multi-core heterogeneous platform is realized.
Owner:SHANDONG UNIV

Systems and Methods for Immobilizing Extracellular Matrix Material on Organ on Chip, Multilayer Microfluidics Microdevices, and Three-Dimensional Cell Culture Systems

The presently disclosed subject matter provides an approach to address the needs for microscale control in shaping the spacial geometry and microarchitecture of 3D collagen hydrogels. For example, the disclosed subject matter provides for compositions, methods, and systems employing N-sulfosuccinimidyl-6-(4′-azido-2′-nitro-phenylamino)hexanoate (“sulfo-SANPAH”), to prevent detachment of the hydrogel from the anchoring substrate due to cell-mediated contraction.
Owner:THE TRUSTEES OF THE UNIV OF PENNSYLVANIA

A hidden PMU reverse engineering general framework based on differential analysis and simulation verification

The application provides a hidden PMU reverse engineering general framework based on differential analysis and simulation verification. The framework can efficiently analyze the function and underlying micro-architecture behavior of Intel processor hidden PMU events. By observing the count change of the entire event space during the execution of a carefully designed instruction fragment, the underlying micro-architecture behavior related to these events is inferred. Based on these initial inferences, a logically equivalent simulator is implemented to model these events. Through the differential analysis between the simulator count and the real CPU count, the simulator is gradually optimized to be consistent with the real CPU behavior, enabling us to reveal the micro-architecture features behind the hidden PMU events. The application has wide application prospects in the fields of performance monitoring, micro-architecture research and security detection, not only improving the hardware resource utilization efficiency of performance counters, but also significantly expanding the ability range of processor performance analysis.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Method and system for verifying RISC-V processor and storage medium

PendingCN121957999Aaccurate associationresolve the disconnectDetecting faulty computer hardwareBiological modelsDesign phaseLogisim
The invention discloses a method and system for verifying an RISC-V processor and a computer readable storage medium. The method comprises the following steps: analyzing an RISC-V official manual and a micro-architecture user manual, and generating a static instruction-resource mapping table; and generating a first test instruction sequence for each instruction in the static instruction-resource mapping table, executing the first test instruction sequence by using an RISC-V processor simulation model constructed by a software simulator, and dynamically calibrating the static instruction-resource mapping table in the execution process. And traversing the instruction and resource mapping relationship in the calibrated static instruction-resource mapping table, analyzing RISC-V processor verification logic, and if a conflict is found, outputting a conflict report so as to complete defect repair before RTL code writing. According to the application, hardware behaviors can be accurately associated, the problem of disjunction of semantics and micro-architecture is solved, and conflict prediction in an architecture design stage and end-to-end automatic verification from instruction semantics to micro-architecture behaviors are realized.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Game CPU-oriented model identification method and system

InactiveCN121116623AResource allocationVideo gamesAlgorithmComputer performance
The invention discloses a game CPU-oriented model identification method and system, and belongs to the technical field of computer performance optimization, and the method comprises the following steps: S1, collecting a real-time performance vector from a CPU core performance monitoring unit; S2, obtaining a predictive load vector from a game engine; s3, adjusting the dynamic learning rate based on the change degree of the predictive load vectors at the current moment and the previous moment; s4, calculating an instantaneous matching degree for each candidate model in a preset candidate CPU micro-architecture model set based on the real-time performance vector, the predictive load vector and a preset reference performance matrix associated with the candidate model; and S5, updating and generating a dynamic confidence score for each candidate model in combination with the confidence, the instantaneous matching degree and the dynamic learning rate of the previous moment, and winning a precious time window for predictive performance optimization of a game engine.
Owner:CHENGDU QISHENGFENG INTERNET TECHNOLOGY CO LTD

A sacrificial layer-constrained elastomer thermo-stretched neural probe for inhibiting necking fracture, its preparation method and application.

This invention relates to the field of flexible sensing fibers, disclosing an elastomer thermally stretched neural probe with sacrificial layer constraint to suppress necking fracture, its preparation method, and its application. The method includes: hot-pressing an elastomer material with a sacrificial layer to form a preform; then hot-stretching it in a low-temperature polymer drawing tower while simultaneously feeding in conductive wires to obtain thermally stretched fibers; finally, removing the sacrificial layer to obtain the elastomer thermally stretched neural probe. This invention introduces a high-viscosity, high-storage-modulus sacrificial layer outside the core layer to form a hierarchical constraint-type microarchitecture preform with pre-formed channels. During the hot-stretching process, the sacrificial layer, as the main load-bearing component, bears most of the axial traction stress and uniformly transfers the deformation stress to the viscous core layer through interfacial mechanical coupling. This stress transfer mechanism effectively protects the fragile core layer, allowing it to thin synchronously with the outer shell without fracture, thereby achieving precise molding of refined conductive units.
Owner:DONGHUA UNIV

Systems and methods for immobilizing extracellular matrix material on organ on chip, multilayer microfluidics microdevices, and three-dimensional cell culture systems

The presently disclosed subject matter provides an approach to address the needs for microscale control in shaping the spacial geometry and microarchitecture of 3D collagen hydrogels. For example, the disclosed subject matter provides for compositions, methods, and systems employing N-sulfosuccinimidyl-6-(4′-azido-2′-nitro-phenylamino)hexanoate (“sulfo-SANPAH”), to prevent detachment of the hydrogel from the anchoring substrate due to cell-mediated contraction.
Owner:THE TRUSTEES OF THE UNIV OF PENNSYLVANIA

Graph processor modeling and analyzing method and device

PendingCN122021503AProcessor architectures/configurationCAD circuit designSoftware emulationSystem-level simulation
The invention discloses a graph processor modeling and analyzing method and device, and belongs to the field of computer system structure aided design. The method comprises the following steps: establishing a unified calculation model decoupled from a storage architecture model; full-link micro-architecture behavior data is obtained through FPGA actual measurement, and a bandwidth constraint model is constructed to determine the optimal parallelism degree of a calculation unit; carrying out closed-loop calibration on bus delay and internal logic parameters of the simulation model by utilizing measured data; and on the calibrated reference model, keeping the optimal parallelism unchanged, and carrying out space evaluation and performance quantitative analysis on various storage architecture schemes. The device comprises an upper computer, an FPGA hardware platform and a communication interface assembly, wherein system-level simulation software and hardware control software are integrated on the upper computer. According to the method, by constructing a software and hardware closed-loop cooperation mechanism, the problems that the degree of parallelism is difficult to reasonably determine and the simulation precision of pure software is insufficient are solved, and high-precision and high-efficiency evaluation of the calculation subsystem and the storage subsystem of the graph processor is realized.
Owner:BEIJING TECH & BUSINESS UNIV

Processor-based system including a processing unit for dynamically reconfiguring micro-architectural features of the processing unit in response to workload being processed on the processing unit

Aspects disclosed in the detailed description include a processing unit for dynamically reconfiguring micro-architectural features of the processing unit in response to a workload being processed on the processing unit and a processing unit control unit configured to receive a plurality of signals from the processing unit. The plurality of signals are indicia of the workload being processed on the processing unit. In response, the processing unit control unit determines whether performance, power consumption, or both, of the processing unit may be improved by modifying a micro-architectural feature. In response to determining that performance or power consumption of the processing unit may be improved by modifying micro-architectural features, the processing unit control unit triggers the processing unit to modify one or more of its micro-architectural features. The processing unit, in response to being triggered by the processing unit control unit, modifies one or more of its micro-architectural features.
Owner:QUALCOMM INC

Edge computing network cooperative scheduling system for security response sinking of end part of high-speed rail platform

The invention relates to the technical field of electric digital data processing, and discloses an edge computing network collaborative scheduling system for high-speed rail platform end security response sinking, which comprises a processor core used for processing a security task; the data acquisition module acquires a video data stream; the scheduling distribution module calculates a bit rate change parameter and predicts a load, and issues a storage isolation control word containing a task identifier and a spatial index when an emergency condition is met; the memory management module intercepts the addressing request and modifies a page table entry memory attribute bit to configure a private address field; the cache consistency filtering module recognizes the memory access state, the hardware level shields a cache line state sniffing request generated by a consistency protocol, through silent isolation of the physical level, micro-architecture level performance degradation caused by the inter-core consistency protocol is eliminated, and it is ensured that the security algorithm processing time delay can be predicted in the complex interference environment.
Owner:HUNAN YOULIANG ELECTRONIC TECH CO LTD

Power embedded kernel micro-architecture optimization method for low-carbon mobile terminal

The invention relates to a power embedded kernel micro-architecture optimization method for a low-carbon mobile terminal, and the method comprises the following steps: S1, obtaining multi-source energy consumption and scene data, carrying out the preprocessing, and obtaining a standardized time series data stream; s2, performing joint feature engineering and scene label generation based on the standardized time sequence data stream, and obtaining a feature vector matrix with a scene label; s3, performing short-term prediction on the load, the temperature, the energy consumption and the QoE based on a short-term prediction model according to the feature vector matrix with the scene label; s4, on the basis of prediction, solving an energy efficiency optimal control quantity meeting a preset constraint, and obtaining a resource scheduling decision and a system configuration parameter; and S5, based on the resource scheduling decision and the system configuration parameters, optimizing the memory access mode and the cache strategy based on the graph neural network, and obtaining a dynamic memory partitioning scheme, a cache replacement strategy and a prefetching decision. The energy consumption of the mobile terminal can be obviously reduced.
Owner:FUJIAN YIRONG INFORMATION TECH +1

Generating iteration transfer information for code execution with a compute slice microarchitecture

A processor core is accessed. The core is configured to execute instructions associated with an instruction set architecture (ISA). The core comprises a plurality of compute slices, a plurality of barrier register files, and a control unit. Each compute slice includes at least one arithmetic logic unit (ALU), a local register file, and is coupled to a successor compute slice and a predecessor compute slice by a barrier register file. Code associated with the ISA is evaluated, where the code includes a first loop. The evaluating includes generating iteration transfer information associated with the first loop. Each slice task within a plurality of slice tasks associated with the first loop is distributed to a compute slice. The processor core executes the plurality of slice tasks. Data forwarding between successive compute slices is based on the plurality of barrier register files and the iteration transfer information.
Owner:ASCENIUM INC

Program control flow protection system and method based on stack data confidentiality and integrity

The application belongs to the technical field of information security of embedded systems, and specifically discloses a program control flow protection system and method based on stack data confidentiality and integrity. The application generates and manages a key stream while the special hardware is accessing the memory at the micro-architecture level, and realizes the fast encryption and decryption calculation of the key data of the stack through the XOR operation of the key stream and the stack data. The stack data confidentiality protection supports the burst transmission property of the AHB-Lite bus which depends on the stack data transmission, and reduces the execution overhead of the program control flow protection method. On the basis of data encryption, the stack data integrity verification architecture based on the MAC code calculates and checks the MAC code when the stack data is written and read, realizes the real-time detection of the tampering of the key data of the stack, and discovers the program control flow anomaly in time. The application realizes the protection of the program control flow from the perspective of the stack data protection through the parallel encryption and decryption and integrity verification of the stack data.
Owner:SHANDONG UNIV OF SCI & TECH

An edge algorithm network coordination scheduling system for high-speed rail platform end security response sinking

ActiveCN121681072BInstant computing power allocationGuaranteed transmission bandwidthProgram initiation/switchingResource allocationDigital dataData stream
The application relates to the technical field of electric digital data processing, and discloses an edge algorithm network cooperative scheduling system for high-speed rail platform end safety response sinking, which comprises a processor core used for processing a safety task; a data acquisition module used for acquiring a video data stream; a scheduling and distribution module used for calculating a bit rate change parameter and predicting a load, and used for issuing a storage isolation control word containing a task identifier and a space index when a burst condition is met; a memory management module used for intercepting an addressing request and modifying a page table item memory attribute bit to configure a private address domain; and a cache consistency filtering module used for identifying a memory access state, and used for shielding a cache line state sniffing request generated by a consistency protocol at a hardware level; the application eliminates micro-architecture level performance attenuation caused by inter-core consistency protocols through physical layer silence isolation, and ensures the predictability of safety algorithm processing delay in a complex interference environment.
Owner:HUNAN YOULIANG ELECTRONIC TECH CO LTD

Job-level parallel performance comprehensive measurement method for heterogeneous system

PendingCN121560715AHardware monitoringMachine learningSoftware emulationAmdahl's law
The invention discloses a job-level parallel performance comprehensive measurement method for a heterogeneous system, and relates to the technical field of performance measurement. Comprising the following steps: obtaining static characteristics, micro-architecture characteristics and job-level operation characteristics of software, and extracting corresponding performance characteristic data; selecting a machine learning model, and setting software simulation data to train and optimize the model to obtain a software performance model; testing the software to generate a throughput curve, and analyzing the throughput curve by using the expandability model to obtain an expandability index; constructing a speed-up ratio model based on the Amdall law; and constructing a measurement model about the performance model, the expandability model and the speed-up ratio model, inputting to-be-predicted software parameters into the measurement model, and weighting generated data to obtain a software comprehensive performance measurement index. According to the method, comprehensive and accurate evaluation on the parallel performance of the software operation in the heterogeneous system is realized.
Owner:COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI

Stream processing-oriented multi-level cache collaborative CGRA processor architecture

The invention relates to the technical field of computer system structures, and discloses a multi-level cache collaborative CGRA processor architecture for stream processing, which comprises a global controller, a software management memory subsystem and a processing element array, and is characterized in that the global controller is connected with the software management memory subsystem and the processing element array through a special bus; and the global controller is used for analyzing the instruction to generate a control signal, and cooperatively scheduling the external storage device, the special cache library and the private storage of the processing element to construct a three-level cache collaborative system. An intermediate result is written into a special cache library through a stream write-back channel so as to reduce access to an external memory device, and an asynchronous execution protocol is executed by utilizing a double-layer configuration storage space in a processing element micro-architecture so as to realize parallel configuration loading and task execution. A configurable data caching mechanism or a roundabout routing path is configured by using a static delay planning method to solve the problem of data synchronization, so that the flow processing calculation energy efficiency is improved.
Owner:CHONGQING UNIV

RISC-V branch predictor closed loop verification method and system

The invention relates to the technical field of processor verification, and provides an RISC-V branch predictor closed-loop verification method and system, and the method comprises the steps: providing a plurality of branch predictor micro-architecture-oriented test modes, and each test mode corresponds to a branch behavior feature and has parameterized configuration; receiving a random seed, selecting a target test mode from the plurality of test modes based on the random seed, and determining parameters for the target test mode; and according to the target test mode and the parameters thereof, generating a corresponding RISC-V assembly test program so as to perform closed-loop verification on the RISC-V branch predictor. According to the method and the device, efficient closed-loop verification is realized through a micro-architecture-oriented parameterized test mode and a random seed driving mechanism, the problem that the test excitation randomness is too high or the simulation time is too long in the prior art is solved, and the verification efficiency is effectively improved.
Owner:YIHUA TECHNOLOGY (BEIJING) CO LTD

Parameter adjustment method and apparatus, and electronic device and computer program product

The embodiments of the present disclosure relate to a parameter adjustment method, a parameter adjustment apparatus, an electronic device, and a computer program product. The method comprises: during the process of a chip executing a service, reading at least one target parameter of the chip, wherein the at least one target parameter includes one or more of the following: a microarchitecture parameter, or a system-on-chip parameter. The method further comprises: acquiring at least one target service performance metric of the service. The method further comprises: during the process of the chip executing the service, adjusting the at least one target parameter on the basis of the at least one target service performance metric. In this way, during the process of a chip executing a service, a microarchitecture parameter and / or a system-on-chip parameter of the chip can be dynamically adjusted on the basis of a target service performance metric of the service, so that the adjusted microarchitecture parameter and / or system-on-chip parameter can help improve the target service performance metric. Thus, the performance of the service can be improved as the chip operates.
Owner:HUAWEI TECH CO LTD

A mainstream CPU hidden PMU event search and utilization method based on machine learning

The application provides a mainstream CPU hidden PMU event search and utilization method based on machine learning. By traversing the entire event space, recording the event count changes of the CPU when executing all valid instructions, searching for hidden PMU events, and generating a feature matrix. Then, combined with the machine learning clustering algorithm, the data is clustered and analyzed to extract the common features of the hidden events, thereby reducing the number of redundant hidden events. In the selection of clustering algorithm, due to the differences in the hidden PMU feature matrix of different processors, the method adopts two commonly used algorithms of DBSCAN and K-Means++, and adjusts the parameters according to the specific situation to optimize the clustering effect. Then, through the evaluation of the contour coefficient index, the optimal clustering algorithm is selected to obtain the best effect. Finally, in order to further prove the effectiveness of the hidden PMU event, the application uses the hidden PMU to perform transient execution attack detection and construct a micro-architecture side channel attack, which shows the utilization potential and security threat of the hidden PMU event.
Owner:BEIJING UNIV OF POSTS & TELECOMM

CPU microarchitecture performance evaluation methods, bare-metal performance evaluation tools, electronic devices and computer program products

This application relates to the field of computer technology, providing a CPU microarchitecture performance evaluation method, a bare-metal performance evaluation tool, an electronic device, and a computer program product. The method, applied to the bare-metal performance evaluation tool, includes: configuring at least one event to be monitored and target test code based on the performance evaluation requirements of the CPU microarchitecture to be evaluated; controlling a virtual CPU to execute the target test code and collecting event execution data corresponding to all events to be monitored during code execution; the virtual CPU being created based on the CPU microarchitecture to be evaluated; and determining the performance evaluation information of the CPU microarchitecture based on the event execution data. This solution fills the gap in the lack of CPU microarchitecture performance evaluation methods in the pre-silicon stage, enabling the evaluation of CPU microarchitecture performance without the need for tape-out.
Owner:GUANGDONG LEAPFIVE TECH CO LTD