Method and device for accelerating startup of CPU core in R & D verification environment
The CPU core startup is accelerated by using a functional abstraction model to replace UNCORE and DRAM, addressing the inefficiencies in the boot up process and reducing development time.
Patent Information
- Application Number
- CN202411048243.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-07-31
AI Technical Summary
In the prior art, the startup process of the CPU core in the R&D verification environment is limited by the specification adaptation and path specification modification of the UNCORE part, resulting in delayed startup and affecting R&D efficiency.
A startup acceleration device that uses functional abstract simulation, including storage model, preinstaller model and initialization model, replaces UNCORE and DRAM, defines the hardware circuit model through a hardware description language to achieve rapid startup of the performance core.
It accelerates the startup speed of the CPU core in the R&D verification environment, reduces the dependence on UNCORE specification adaptation, and improves the efficiency of R&D verification.
Smart Images

Figure CN119597353B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the technical field of CPU cores, and particularly to a method and device for accelerating the startup of a CPU core in a research and development verification environment. Background Art
[0002] The CPU is the most important component in a computer system and is used to control the entire computer to complete the work specified by a specific program. In the scenario of high-performance CPU research and development, the hardware system of the CPU can generally be simply classified into two parts: the performance core (CORE) and the non-core (UNCORE). Among them, the CORE is the main functional part, responsible for fetching instructions, decoding, executing, and other scheduling and computing tasks; the UNCORE includes the control core and various peripherals of the on-chip system (SOC) for managing functions such as the power supply, clock reset, peripheral data transfer, and event communication of the CPU.
[0003] The CORE startup process (boot up) generally includes the startup of the control core, the configuration of the peripheral system, and the boot and loading startup of the performance core. Against the background of the increasing area of the CPU and the increasing number of peripherals to be managed, and in order to accommodate more diverse application scenarios, the industry has gradually tended to package the CORE and UNCORE parts separately, so that the same CORE DIE can be matched with different peripheral parts, thereby deriving CPUs of different specifications and improving the diversification of the product line. Then, for the relatively long boot up process of the CPU, the CORE R & D department can invest more energy in the CORE boot up itself in the early stage, ignoring the configuration and initialization of the peripheral system, thereby improving the iteration efficiency in the early stage of the product.
[0004] However, the existing CORE boot up process generally adapts to a complete CPU. Most of the content is not concerned in the CORE design stage, but a large amount of energy has to be invested in sorting out the specifications of the peripheral UNCORE for adaptation modifications, or the startup is delayed due to being blocked by the progress of modifying the access specifications between the CORE and the UNCORE. Therefore, how to improve the startup speed of the CPU core in the research and development verification environment has become an urgent technical problem to be solved at present. Summary of the Invention
[0005] The purpose of the embodiments of this specification is to provide a method and device for accelerating the startup of a CPU core in a research and development verification environment to improve the startup speed of the CPU core in the research and development verification environment.
[0006] To achieve the above object, on the one hand, the embodiments of this specification provide a device for accelerating the startup of a CPU core in a research and development verification environment, including an initialization model, a pre-installer model, and a storage model;
[0007] The storage model is used to store program files so that when the performance core initiates an instruction data request with non-cache attributes from its front end, it can obtain target instruction data from the program files in the storage model and decode it;
[0008] The initialization model is used to decode the start clock reset instruction data obtained from the program files in the storage model, and perform a start clock reset on the performance core based on the decoded start clock reset instruction data;
[0009] The preloader model is used to preload the program file preloader in the storage model into the cache of the performance core so that when the performance core initiates an instruction data request with cache attributes from its front end, it can obtain target instruction data from the program files in the cache and decode it.
[0010] In the start-up acceleration device of the CPU core in the R & D verification environment according to the embodiment of this specification, obtaining and decoding target instruction data from the program files in the storage model includes:
[0011] If the data interface of the performance core is not defined, the target instruction data is obtained and decoded from the program files in the storage model in sequence through the cache and the bypass interface connecting the cache and the storage model.
[0012] In the start-up acceleration device of the CPU core in the R & D verification environment according to the embodiment of this specification, obtaining and decoding target instruction data from the program files in the storage model includes:
[0013] If the data interface of the performance core is defined, the instruction data request is protocol-converted and then sent to the data interface to obtain and decode target instruction data from the program files in the storage model through the data interface.
[0014] In the start-up acceleration device of the CPU core in the R & D verification environment according to the embodiment of this specification, the cache of the performance core includes the outermost cache of the performance core.
[0015] In the start-up acceleration device of the CPU core in the R & D verification environment according to the embodiment of this specification, the initialization model, the preloader model, and the storage model are all hardware circuit models defined based on a hardware description language.
[0016] In the start-up acceleration device of the CPU core in the R & D verification environment according to the embodiment of this specification, the preloader model is pre-constructed based on the physical structure of the cache.
[0017] On the other hand, the embodiment of this specification also provides a start-up acceleration method for the CPU core in the R & D verification environment based on the above start-up acceleration device, including:
[0018] Move the program file to the storage model;
[0019] Preload the program file in the storage model into the cache of the performance core;
[0020] Decode the start clock reset instruction data obtained from the program file in the storage model, and perform a start clock reset on the performance core based on the decoded start clock reset instruction data;
[0021] When the performance core initiates an instruction data request for cache attributes from its front end, obtain target instruction data from the program file in the cache and decode it; or,
[0022] When the performance core initiates an instruction data request for non-cache attributes from its front end, obtain target instruction data from the program file in the storage model and decode it.
[0023] In the start-up acceleration method of the CPU core in the R & D verification environment according to the embodiments of the present specification, obtaining target instruction data from the program file in the storage model and decoding it includes:
[0024] If the performance core does not define a data interface, sequentially obtain target instruction data from the program file in the storage model through the cache and the bypass interface connecting the cache and the storage model and decode it.
[0025] In the start-up acceleration method of the CPU core in the R & D verification environment according to the embodiments of the present specification, obtaining target instruction data from the program file in the storage model and decoding it includes:
[0026] If the performance core has defined a data interface, perform protocol conversion on the instruction data request and send it to the data interface to obtain target instruction data from the program file in the storage model through the data interface and decode it.
[0027] On the other hand, the embodiments of the present specification also provide a computer device, including a memory, a processor, and a computer program stored on the memory. When the computer program is run by the processor, it executes the instructions of the above method.
[0028] On the other hand, the embodiments of the present specification also provide a computer storage medium, on which a computer program is stored. When the computer program is run by the processor of a computer device, it executes the instructions of the above method.
[0029] On the other hand, the embodiments of the present specification also provide a computer program product, which includes a computer program. When the computer program is run by the processor of a computer device, it executes the instructions of the above method.
[0030] As can be seen from the technical solutions provided in the embodiments of this specification above, in the R & D and verification stage of the performance core (CORE), the non-core part (UNCORE) and memory (DRAM) in the CPU are replaced by the start-up acceleration device simulated by function abstraction, so that the boot-up of the CORE can be transformed from the dependence on the UNCORE and DRAM to the dependence on the start-up acceleration device simulated by function abstraction, thus avoiding the impact brought by the lag in the specification adaptation between the CORE and the UNCORE, and further accelerating the start-up speed of the CORE's boot-up. Brief Description of the Drawings
[0031] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings. In the drawings:
[0032] Figure 1 Shows a schematic diagram of the start-up acceleration principle of the CPU core in the prior art in the R & D and verification environment;
[0033] Figure 2 Shows a schematic diagram of the principle of the start-up acceleration device of the CPU core in some embodiments of this specification in the R & D and verification environment;
[0034] Figure 3 Shows Figure 2 A schematic diagram of the connection relationship between the outermost cache and the storage model of the performance core in the shown embodiment;
[0035] Figure 4 Shows a flowchart of the start-up acceleration method of the CPU core in some embodiments of this specification in the R & D and verification environment;
[0036] Figure 5 Shows a structural block diagram of a computer device in some embodiments of this specification.
[0037]
Explanation of the Reference Numerals in the Drawings
[0038] 10. Start-up acceleration device;
[0039] 11. Storage model;
[0040] 12. Preloader model;
[0041] 13. Initialization model;
[0042] 14. Bypass interface;
[0043] 15. Data interface;
[0044] 20. Performance core;
[0045] 30. Program file;
[0046] 502. Computer device;
[0047] 504. Processor;
[0048] 506. Memory;
[0049] 508. Driving mechanism;
[0050] 510. Input / output interface;
[0051] 512. Input device;
[0052] 514. Output device;
[0053] 516. Presentation device;
[0054] 518. Graphical user interface;
[0055] 520. Network interface;
[0056] 522. Communication link;
[0057] 524. Communication bus. Detailed implementation
[0058] In order to enable those skilled in the art of this technology to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this specification without creative efforts shall fall within the scope of protection of this specification.
[0059] Figure 1The traditional CORE startup process is shown. In this CORE startup process, the program is pre-moved to a Dynamic Random Access Memory (DRAM). The Manage Processer (MP) obtains its own program from the DRAM and performs power-on reset on the CORE and the Level 3 Cache (L3C). The CORE initiates an instruction data (program) transfer request from the FrontEnd. The instruction data starts from the DRAM and sequentially passes through the L3C and the Level 2 Cache (L2C) to reach the FrontEnd, and instruction decoding is completed in the FrontEnd. The content of the startup program includes not only the initialization of the CORE itself but also the initialization of the peripherals in the UNCORE.
[0060] It can be seen that the traditional CORE startup process depends on the normal functions and clear specifications of the UNCORE part, and is premised on the data path connection between the CORE and the UNCORE meeting the expected functions. In other words, the existing CORE boot up process needs to be adapted to the complete CPU. Most of the content does not need to be concerned about during the CORE design stage, but a large amount of effort has to be invested in sorting out the specifications of the peripheral UNCORE for adaptation modifications, or the startup is delayed due to the progress of the communication path specification modification between the CORE and the UNCORE. This will affect the CORE boot up progress and the early debugging work.
[0061] In view of this, the embodiments of this specification provide a startup acceleration device and method for a CPU core in a research and development verification environment to improve the startup speed of the CPU core in the research and development verification environment.
[0062] Figure 2The schematic diagram of the principle of the startup acceleration device of the CPU core in the R & D verification environment in some embodiments of this specification is shown; the startup acceleration device 10 is a functional abstract simulation of UNCORE and DRAM in the CPU to complete the data and process support during the CORE startup process, so as to avoid the impact of the UNCORE part on the CORE startup process. The startup acceleration device 10 may include a storage model (MemModel) 11, a preloader model (Preloader) 12, and an initialization model (Init) 13. Among them, the storage model 11 is used to store the program file 30, so that when the performance core 20 (i.e., the CORE of the CPU) initiates a non-cachable instruction data request from its front end (FrontEnd), it can obtain the target instruction data (here it refers to the instruction data required for the startup of the performance core 20) from the program file (Program) 30 in the storage model 11 and decode it; the initialization model 13 is used to decode the power-on reset instruction data obtained from the program file 30 in the storage model 11, and perform a power-on reset on the performance core 20 based on the decoded power-on reset instruction data; the preloader model 12 is used to preload the program file 30 in the storage model 11 into the outermost cache of the performance core 20, so that when the performance core 20 initiates a cacheable instruction data request from its FrontEnd, it can obtain the target instruction data from the program file 30 in the outermost cache and decode it. Among them, FrontEnd represents the functional part in the CORE except for the outermost cache in this article.
[0063] In the embodiments of this specification, during the R & D verification stage of the CORE, the startup acceleration device of functional abstract simulation is used to replace UNCORE and DRAM in the CPU. Thus, the boot up of the CORE can be transformed from relying on UNCORE and DRAM to relying on the startup acceleration device of functional abstract simulation (i.e., simple functional abstraction), so as to avoid the impact brought by the lag in the specification adaptation between the CORE and UNCORE (that is, remove the initialization content of the UNCORE part peripherals from the CORE's boot up, and finally achieve the purpose of completely shielding the impact of UNCORE changes), and then the startup speed of the CORE's boot up can be accelerated. In this way, a CORE-level fast debugging method is provided for CORE R & D, so that there is no need to invest a lot of effort in providing the clear specifications of UNCORE during the process of adapting different UNCOREs for R & D verification of the CORE.
[0064] In some embodiments of this specification, the storage model, the preloader model, and the initialization model can all be hardware circuit models defined based on a Hardware Description Language (HDL). For example, in an exemplary embodiment of this specification, the above-mentioned storage model, preloader model, and initialization model can be defined based on languages such as Systemverilog and Verilog.
[0065] Combined with Figure 3 As shown, in some embodiments of this specification, the storage model 11 in the startup acceleration device 10 can be connected to the outermost cache of the performance core 20 through a bypass interface (bypass interface) 14 and a data interface 15. Among them, the data interface 15 refers to the top-level data communication interface between the CORE and the UNCORE. In some embodiments of this specification, when the data interface 15 of the performance core 20 is not defined, the storage model 11 can communicate with the outermost cache of the performance core 20 through the bypass interface 14; thus, in the case where the data interface 15 of the performance core 20 is not defined, through the bypass interface 14, it can be ensured that the storage model 11 can still work properly. In this case, obtaining and decoding target instruction data from the program file in the storage model 11 includes: if the data interface 15 of the performance core 20 is not defined, the target instruction data can be obtained and decoded from the program file in the storage model 11 in sequence through the outermost cache and the bypass interface 14 connecting the outermost cache and the storage model 11.
[0066] In some embodiments of this specification, when the data interface 15 of the performance core 20 is defined (such as defined as a Cbox interface, etc.), the storage model 11 can communicate with the outermost cache of the performance core 20 through the data interface 15 of the performance core 20; thus, startup simulation of a specific UNCORE part can be achieved. In this case, obtaining and decoding target instruction data from the program file in the storage model 11 includes: if the data interface 15 of the performance core 20 is defined, the instruction data request can be protocol-converted and then sent to the data interface 15 to obtain and decode the target instruction data from the program file in the storage model 11 through the data interface 15.
[0067] In some embodiments of this specification, the defined data interface 15 can mean that the signal line name, the number of signal lines, and the signal behavior (i.e., communication protocol) of the data interface 15 are defined. Correspondingly, the undefined data interface 15 can mean that at least one of the signal line name, the number of signal lines, and the signal behavior of the data interface 15 is undefined.
[0068] In addition, both the preloader model and the initialization model can communicate with the CORE through a simulated hardware data communication interface (signal line).
[0069] In some other embodiments of this specification, considering that the CPU has multiple architectures (such as x86, ARM, RISC-V, etc.), between the CORE and the UNCORE under different CPU architectures, when the data interface of the CORE is not defined, it is possible to perform data communication through other interfaces with similar functions to the bypass interface. Therefore, in some embodiments of this specification, when the data interface of the CORE is not defined, the data communication between the CORE and the UNCORE is realized through the bypass interface, which is only an exemplary illustration and should not be understood as the sole limitation of the embodiments of this specification.
[0070] In some embodiments of this specification, the preloader model can be pre-constructed based on the physical structure of the outermost cache of the performance core; when the preloader model preloads the program file in the storage model into this outermost cache, it can directly load the data at a certain physical address originally stored in the storage model into the corresponding storage unit position inside the outermost cache, and the mapping position of this physical address in the outermost cache is strongly related to the physical structure of the outermost cache; in this way, the effect of simulating the performance core loading its own outermost cache can be achieved. Therefore, as long as the physical structure of the outermost cache of the CORE does not change significantly, even when the data interface 15 of the CORE is in an undefined state, the FrontEnd of the CORE can still directly hit the outermost cache and obtain the data required for startup from it, thereby saving the time-consuming overhead of communication between the CORE and the outside.
[0071] For example, the preloader model defined based on the Systemverilog language can specify the values of specific circuit-level signals at a specified moment, that is, the preloader model can complete the mapping from the physical address to the specific circuit level, so that the original "loading the program file to the specific physical address" can be changed to "loading the program file to the circuit-level signal of the outermost cache (i.e., the physical unit of the outermost cache)".
[0072] In some embodiments of this specification, the initialization model 13 is a functional abstract simulation of the MP in the UNCORE to replace the MP to perform power-on reset on the CORE.
[0073] It should be noted that the outermost cache of the above-mentioned performance core is only for illustrative purposes, so that the startup process of the CPU core in the R & D verification environment can cover a longer communication path as much as possible. For example, in an exemplary embodiment of this specification, if the performance core has a level 1 cache (L1C) and a level 2 cache (L2C), the outermost cache of the performance core is L2C; in another exemplary embodiment of this specification, if the performance core only has L1C, the outermost cache of the performance core is L1C. Therefore, in the embodiments of this specification, the outermost cache of the performance core depends on the specific situation of the hierarchical or layered architecture of the cache of the performance core.
[0074] In other embodiments of this specification, according to actual needs, the outermost cache of the above-mentioned performance core can also be replaced with any designated cache of the performance core (not necessarily its outermost cache).
[0075] For the convenience of description, when describing the above device, it is divided into various units according to functions and described separately. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0076] Based on the above-mentioned startup acceleration device of the CPU core in the R & D verification environment, the embodiments of this specification provide a startup acceleration method for the CPU core in the R & D verification environment. Refer to Figure 4 As shown, in some embodiments of this specification, the startup acceleration method for the CPU core in the R & D verification environment may include the following steps:
[0077] Step 401, Move the program file to the storage model.
[0078] In some embodiments of this specification, the program file can be moved from an external storage (such as a hard disk, optical disc, USB flash drive, etc.) to the storage model.
[0079] Step 402, Preload the program file in the storage model into the outermost cache of the performance core.
[0080] In some embodiments of this specification, the program file in the storage model can be preloaded into the outermost cache of the performance core through a preloader model, so that the performance core can obtain it as needed during subsequent startup.
[0081] Step 403, Decode the power-on reset instruction data obtained from the program file in the storage model, and perform a power-on reset on the performance core based on the decoded power-on reset instruction data.
[0082] In some embodiments of the present specification, the open clock reset instruction data obtained from the program file in the storage model can be decoded through an initialization model, and the performance core can be reset based on the decoded open clock reset instruction data. Among them, the open clock reset is to perform a reset operation on the clock of the performance core (such as hardware reset, software reset, and power-on reset).
[0083] Step 404: When the performance core initiates an instruction data request for cache attributes from its front end, obtain and decode the target instruction data from the program file in the outermost cache; or, when the performance core initiates an instruction data request for non-cache attributes from its front end, obtain and decode the target instruction data from the program file in the storage model.
[0084] In some embodiments of the present specification, when the performance core initiates an instruction data request for non-cache attributes (i.e., non-cacheable attributes) from its front end, obtaining and decoding the target instruction data from the program file in the storage model includes: if the data interface of the performance core is not defined, sequentially obtain and decode the target instruction data from the program file in the storage model through the outermost cache and the bypass interface connecting the outermost cache and the storage model.
[0085] In some embodiments of the present specification, when the performance core initiates an instruction data request for non-cache attributes (i.e., non-cacheable attributes) from its front end, obtaining and decoding the target instruction data from the program file in the storage model further includes: if the data interface of the performance core is defined, perform protocol conversion on the instruction data request and then send it to the data interface to obtain and decode the target instruction data from the program file in the storage model through the data interface.
[0086] In some embodiments of the present specification, when the performance core initiates an instruction data request for cache attributes (i.e., cacheable attributes) from its front end, if the target instruction data is not stored in the outermost cache, the target instruction data can be obtained and decoded from the program file in the storage model, that is, if the target instruction data is not stored in the outermost cache, the target instruction data can also be sequentially obtained and decoded from the program file in the storage model through the outermost cache and the bypass interface connecting the outermost cache and the storage model.
[0087] Although the process flow described above includes multiple operations that occur in a specific order, it should be clearly understood that these processes can include more or fewer operations, and these operations can be executed sequentially or in parallel (such as using a parallel processor or a multi-threaded environment).
[0088] It should be noted that in the embodiments of this specification, the user information that may be involved (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data that have been authorized and consented to by the user and fully authorized by all parties. That is, the acquisition, transmission, storage, use, processing, etc. of data in the technical solution of this application all comply with the relevant provisions of national laws and regulations.
[0089] The embodiments of this specification also provide a computer device. As Figure 5 shown, in some embodiments of this specification, the computer device 502 may include one or more processors 504, such as one or more central processing units (CPUs) or graphics processing units (GPUs), and each processing unit may implement one or more hardware threads. The computer device 502 may also include any memory 506, which is used to store any kind of information such as code, settings, data, etc. In a specific embodiment, a computer program that can be run on the memory 506 and on the processor 504. When the computer program is run by the processor 504, it can execute the instructions of the startup acceleration method of the CPU core in the R & D verification environment described in any of the above embodiments. Non-limitingly, for example, the memory 506 may include any one or a combination of the following: any type of RAM, any type of ROM, flash memory devices, hard disks, optical discs, etc. More generally, any memory can use any technology to store information. Further, any memory can provide volatile or non-volatile retention of information. Further, any memory can represent a fixed or removable component of the computer device 502. In one case, when the processor 504 executes the associated instructions stored in any memory or combination of memories, the computer device 502 can perform any operation of the associated instructions. The computer device 502 also includes one or more drive mechanisms 508 for interacting with any memory, such as a hard disk drive mechanism, an optical disc drive mechanism, etc.
[0090] The computer device 502 may also include an input / output interface 510 (I / O), which is used to receive various inputs (via the input device 512) and to provide various outputs (via the output device 514). A specific output mechanism may include a presentation device 516 and an associated graphical user interface 518 (GUI). In other embodiments, the input / output interface 510 (I / O), the input device 512, and the output device 514 may not be included, and it only serves as a computer device in the network. The computer device 502 may also include one or more network interfaces 520, which are used to exchange data with other devices via one or more communication links 522. One or more communication buses 524 couple the components described above together.
[0091] The communication link 522 can be implemented in any manner, for example, through a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication link 522 can include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc. governed by any protocol or combination of protocols.
[0092] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), computer-readable storage media, and computer program products according to some embodiments of this specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processors to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processors generate means for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0093] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processors to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0094] These computer program instructions can also be loaded onto a computer or other programmable data processors, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, so that the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0095] In a typical configuration, a computer device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0096] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0097] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computer device. As defined in this specification, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0098] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, the embodiments of this specification can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0099] The embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The embodiments of this specification can also be practiced in a distributed computing environment where tasks are performed by remote processors connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0100] It should also be understood that in the embodiments of this specification, the term "and / or" is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0101] Each embodiment in this specification is described in a progressive manner, with each embodiment highlighting the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0102] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of this specification. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0103] The above are only the embodiments of this application and are not used to limit this application. For those skilled in the art, various changes and modifications can be made to this application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.
Claims
1. A startup acceleration device for a CPU core in a research and development verification environment, characterized in that It includes an initialization model, a preloader model, and a storage model; the initialization model, the preloader model, and the storage model are all hardware circuit models defined based on a hardware description language; the preloader model is pre-constructed based on the physical structure of the cache. The storage model is used to store program files so that when the performance core initiates an instruction data request with non-cache attributes from its front end, it can obtain target instruction data from the program files in the storage model and decode it. The initialization model is used to decode the start clock reset instruction data obtained from the program files in the storage model and perform a start clock reset on the performance core based on the decoded start clock reset instruction data. The preloader model is used to preload the program files at specific physical addresses in the storage model to the storage unit positions corresponding to the outermost cache of the performance core, so that when the performance core initiates an instruction data request with cache attributes from its front end, it can obtain target instruction data from the program files in the outermost cache and decode it; wherein, the mapping position of the specific physical address in the outermost cache is related to the physical structure of the outermost cache.
2. The startup acceleration device for the CPU core in the R & D verification environment according to claim 1, characterized in that Obtaining target instruction data from the program files in the storage model and decoding it includes: If the data interface of the performance core is not defined, sequentially obtain target instruction data from the program files in the storage model and decode it through the cache and the bypass interface connecting the cache and the storage model.
3. The startup acceleration device for the CPU core in the R & D verification environment according to claim 1, wherein Obtaining target instruction data from the program files in the storage model and decoding it includes: If the data interface of the performance core is already defined, perform protocol conversion on the instruction data request and send it to the data interface to obtain target instruction data from the program files in the storage model and decode it through the data interface.
4. A startup acceleration method for a CPU core of the startup acceleration device according to claim 1 in a research and development verification environment, characterized in that, It includes: Moving the program files to the storage model; Preloading the program files in the storage model into the cache of the performance core; Decoding the start clock reset instruction data obtained from the program files in the storage model and performing a start clock reset on the performance core based on the decoded start clock reset instruction data; When the performance core initiates an instruction data request with cache attributes from its front end, obtaining target instruction data from the program files in the cache and decoding it; Or, When the performance core initiates an instruction data request with non-cache attributes from its front end, obtaining target instruction data from the program files in the storage model and decoding it.
5. The startup acceleration method according to claim 4, wherein Obtaining target instruction data from the program files in the storage model and decoding it includes: If the data interface of the performance core is not defined, sequentially obtain target instruction data from the program files in the storage model and decode it through the cache and the bypass interface connecting the cache and the storage model.
6. The startup acceleration method according to claim 4, wherein Obtaining target instruction data from the program files in the storage model and decoding it includes: If the data interface of the performance core is already defined, perform protocol conversion on the instruction data request and send it to the data interface to obtain target instruction data from the program files in the storage model and decode it through the data interface.
7. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, When the computer program is run by the processor, it executes the instructions of the method according to any one of claims 4-6.
8. A computer storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor of the computer device, it executes the instructions of the method according to any one of claims 4-6.
9. A computer program product, characterized in that, The computer program product includes a computer program which, when run by the processor of the computer device, executes the instructions of the method according to any one of claims 4-6.
Citation Information
Patent Citations
Server-based microkernel operating system deployment method and operating system
CN113127077A
Systems and methods for recording instruction sequences in a microprocessor having a dynamically decoupleable extended instruction pipeline
US20070074012A1