Chip optimization processing method and system and chip

By implementing the system on chip on the chip and obtaining and applying behavior prediction rules to optimize processor performance, the problem that hardware accelerators cannot cover multiple software scenarios and are cost-effective, and efficient optimization for different target programs is achieved.

CN120068790AActive Publication Date: 2025-05-30BEIJING YUNYAO XINDAO TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510124735.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-30
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

In processor optimization, the use of hardware accelerators cannot cover multiple software scenarios and is costly.

Method used

By implementing the system on chip on the chip, behavior prediction rules obtained based on historical operation information analysis, and taking pre-emptive measures when running the target software to optimize performance.

Benefits of technology

The processor performance optimization for different target programs is achieved, avoiding the problem that the hardware accelerator cannot cover other software, and at the same time reducing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068790A_ABST
    Figure CN120068790A_ABST
Patent Text Reader

Abstract

The invention discloses a chip optimization processing method and system and a chip, and the method comprises the steps that a system-on-chip located on the chip obtains a rule, the rule reflects behavior prediction when target software is operated on the system-on-chip, and the target software is obtained according to the behavior prediction; the rule is obtained through analysis according to operation information collected when the system-on-chip operates the target software in history, and the operation information is generated when the system-on-chip operates the target software; when the system-on-chip runs the target software, measures are taken in advance according to the rules, and the measures taken in advance estimate future operations performed by the system-on-chip and make preparations for the operations. By means of the method and device, the problems that when processor optimization is conducted on specific software, a hardware accelerator is adopted, so that other software cannot be covered, and the cost is high are solved, and therefore processor performance optimization can be achieved for different target programs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of chips, and in particular, to a method, a system, and a chip for optimizing chip processing. Background Art

[0002] Traditional processors are mainly oriented towards general computing scenarios, aiming to achieve relatively balanced performance for flexible and variable programs. However, this inevitably leads to different performance fluctuations due to different programs. In fact, in some scenarios, the same program runs for a long time, including but not limited to dedicated programs such as cloud computing, search, recommendation, advertising, and EDA that run on servers for a long time.

[0003] In order to further improve the software running efficiency, the most straightforward approach is to enhance the performance of the processor. However, in actual applications, designing dedicated hardware for a program incurs high costs. In other processing methods, a hardware-specific accelerator (Domain Specific Accelerator, abbreviated as DSA) can be used. The hardware accelerator designs a dedicated hardware architecture for a specific program to achieve optimal performance. However, the disadvantage is that it cannot cover a rich range of scenarios and the cost is also not low. Summary of the Invention

[0004] Embodiments of the present application provide a method, a system, and a chip for optimizing chip processing, so as to at least solve the problems of inability to cover other software and high cost caused by using a hardware accelerator when optimizing a processor for a specific software.

[0005] According to an aspect of the present application, there is provided a method for optimizing chip processing, including: a system-on-chip located on the chip obtains a rule, where the rule reflects a behavior prediction when the system-on-chip runs a target software, the rule is obtained by analyzing operation information collected when the system-on-chip historically runs the target software, and the operation information is generated when the system-on-chip runs the target software; when the system-on-chip runs the target software, it takes measures in advance according to the rule, where the measures are determined according to the rule, and the measures taken in advance estimate the operations to be performed by the system-on-chip in the future and prepare for the operations.

[0006] Further, it also includes: the analysis software receives the operation information, wherein the operation information is obtained through an interface pre-configured on the chip, and after the operation information is read out from the interface, it is saved in a storage space outside the chip; the analysis software determines the rule after analyzing the operation information; the analysis software sends the rule to the system on chip; or, the system on chip obtains the rule including: the analysis software receives the operation information; the analysis software obtains parameters for generating the rule based on the operation information; the system on chip receives the parameters and uses the parameters in the algorithm running on the system on chip to generate the rule; wherein the algorithm is used to generate the rule.

[0007] Furthermore, the computing resources for running the analysis software and the system on chip for running the target software are different computing resources.

[0008] Further, in the case where the chip includes multiple cores, the core on the chip includes two parts, the first part of the two parts is used to run the analysis software, and the second part of the two parts runs the target software as a system on chip; or, the analysis software runs on a computing device, and the computing device is different from the computing device where the chip is located; or, the analysis software runs in the system on chip, and the system on chip uses different time domain resources to time-share the analysis software and the target software.

[0009] Furthermore, the target software runs on multiple chips, wherein the multiple chips are respectively located on different computing devices; the analysis software receives the operating information generated by the system-on-chip on each of the multiple chips when the target software runs; the analysis software analyzes the operating information generated by all the collected system-on-chips to obtain the rules; and the analysis software passes the rules to the system-on-chip on each of the multiple chips.

[0010] According to another aspect of the present application, a chip is also provided, comprising a central processing unit and a cache; wherein the system on chip on the chip is used to execute the following method steps: obtaining rules, wherein the rules reflect the behavior prediction of the system on chip when running target software, and the rules are obtained by analyzing the operation information collected when the system on chip historically ran the target software, and the operation information is generated by the system on chip running the target software; when running the target software, taking measures in advance according to the rules, wherein the measures are determined according to the rules, and the measures taken in advance predict the operations to be performed by the system on chip in the future and prepare for the operations.

[0011] Furthermore, it also includes: an interface, wherein the interface is one or more, and the interface is used to send the operating information to the outside of the system on chip, and to receive rules from the analysis software or receive parameters from the analysis software; wherein the parameters are used in the algorithm running on the system on chip to generate the rules, and the algorithm is used to generate the rules.

[0012] According to another aspect of the present application, a chip optimization processing system is also provided, comprising the above-mentioned chip and analysis software, wherein the analysis software is used to execute the following method steps: receiving operation information, wherein the operation information is obtained through an interface pre-configured on the chip, and after the operation information is read out from the interface, it is saved in a storage space outside the chip; determining the rules after analyzing the operation information; sending the rules to the on-chip system; or, obtaining parameters for generating the rules based on the operation information, and sending the parameters to the on-chip system; wherein the parameters are used in the algorithm running on the on-chip system to generate the rules; and the algorithm is used to generate the rules.

[0013] Furthermore, the computing resources for running the analysis software and the system on chip for running the target software are different computing resources.

[0014] Further, in the case where the chip includes multiple cores, the core on the chip includes two parts, the first part of the two parts is used to run the analysis software, and the second part of the two parts runs the target software as a system on chip; or, the analysis software runs on a computing device, and the computing device is different from the computing device where the chip is located; or, the analysis software runs in the system on chip, and the system on chip uses different time domain resources to time-share the analysis software and the target software.

[0015] In an embodiment of the present application, a system-on-chip acquisition rule located on a chip is used, wherein the rule reflects a prediction of the behavior of the system-on-chip when running the target software, and the rule is obtained by analyzing the operation information collected when the system-on-chip historically ran the target software, and the operation information is generated by the system-on-chip running the target software; when the system-on-chip runs the target software, the system-on-chip takes measures in advance according to the rules, wherein the measures are determined according to the rules, and the measures taken in advance estimate the operations to be performed by the system-on-chip in the future and prepare for the operations. This application solves the problem of not being able to cover other software and the high cost caused by using hardware accelerators when optimizing the processor for specific software, so that the processor performance can be optimized for different target programs. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings, which are a part of this cost application, are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the accompanying drawings:

[0017] Figure 1 is a schematic diagram of the computer software and hardware architecture in the related art;

[0018] Figure 2 is a schematic diagram of the optimization of the target software operation according to an embodiment of this application Figure 1 ;

[0019] Figure 3 is a schematic diagram of the optimization of the target software operation according to an embodiment of this application Figure 2 ;

[0020] Figure 4 is a schematic diagram of the optimization of the target software operation according to an embodiment of this application Figure 3 ;

[0021] Figure 5 is a schematic diagram of the interaction timing sequence according to an embodiment of this application Figure 1 ;

[0022] Figure 6 is a schematic diagram of the interaction timing sequence according to an embodiment of this application Figure 2 ;

[0023] Figure 7 is a schematic diagram of the interaction timing sequence according to an embodiment of this application Figure 3 ;

[0024] Figure 8 is an interaction flow chart according to an embodiment of this application;

[0025] Figure 9 is a schematic diagram of the interaction process according to an embodiment of this application; and,

[0026] Figure 10 is a flow chart of the chip optimization processing method according to an embodiment of this application. Detailed implementation manners

[0027] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The following will refer to the accompanying drawings and combine with the embodiments to detail this application.

[0028] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0029] In related technologies, in order to improve performance, a processor usually uses more resources. For example, to increase the hit rate of the cache, the common approach of traditional processor manufacturers is mainly to increase the area of the cache and increase the amount of cached data to improve performance. This approach increases costs and also increases the size of the processor. In the following embodiments, the historical running information of the target software or target program is mainly processed and analyzed to predict future running trajectories. In this way, instead of trading area for performance, performance is improved through prediction while saving area.

[0030] First, the technical terms involved in the following embodiments will be described.

[0031] System on Chip

[0032] In the following embodiments, the system in a chip (or also referred to as a processor) is called a system on chip. For a system on chip, it may include a central computing unit, a cache, etc. In some literature, the central processing unit (CPU) can be understood as a system on chip, and in other literature, the central processing unit can also be understood as the central computing unit on the system on chip. To avoid this ambiguity, the system on chip and the central computing unit are used in the following embodiments.

[0033] Offline software or offline program

[0034] In the following embodiments, an online program represents real-time operation. For example, if the target software is an online software, it will run in real time for high performance; the analysis software is an offline program, such as it can run offline separately in terms of timing or on other machines. In some scenarios, the analysis software can also be an online program for online analysis for high performance. It should be noted that if there are multiple cores, the cores can be divided into two parts, one part runs the online program and the other part runs the offline program. In one example, the processing running inside the chip is hardware real-time processing, but if it runs in real time at the operating system layer on the system on chip, it is also considered an online program. Offline is more of a separation in terms of time and space relationship. For example, what does not need to run in real time is called offline, and what can run on other systems is called offline. In scenarios that require high-performance execution, online implementation will be selected.

[0035] Figure 1 It is based on the schematic diagram of the computer software and hardware architecture in related technologies, as Figure 1 shown, the application or program runs on the operating system, and then runs on the chip through the compiler. In Figure 1 the application or program, the operating system, and the compiler are all referred to as the software part, and the system on chip running in the chip and inside the chip is called the hardware part.

[0036] i-cache and d-cache

[0037] The i-cache (instruction cache) and d-cache (data cache) play different roles respectively to improve the execution efficiency of the system-on-chip and the overall system performance. The i-cache is a high-speed cache specifically used to store the instructions that the system-on-chip is about to execute. When the system-on-chip needs to execute an instruction, it first checks whether the instruction has been cached in the i-cache. If the instruction has been cached in the i-cache (referred to as a cache hit), the system-on-chip can directly read the instruction from the i-cache without waiting to fetch it from the memory. This method significantly improves the access speed of the instructions, thus accelerating the execution efficiency of the program. The d-cache is a high-speed cache specifically used to store the data that the system-on-chip has recently used, such as variables, arrays, etc. When the system-on-chip needs to access this data, it checks whether the data has been cached in the d-cache. If the data has been cached in the d-cache, the system-on-chip can directly read it from the d-cache, thus avoiding the latency caused by frequently reading data from the memory. This method improves the access speed of the data, thereby enhancing the overall performance of the program.

[0038] In one example, the system-on-chip may also adopt a multi-level cache. In this case, the i-cache and d-cache can be part of the multi-level cache system integrated within the system-on-chip. For example, a typical system-on-chip may include three levels of caches: L1, L2, and L3. Among them, the L1 cache is the cache closest to the core of the system-on-chip and is usually divided into two parts: i-cache and d-cache. When the system-on-chip executes a program, it first looks for the required instructions and data in the L1 cache. If the required data or instructions are not found in the L1 cache, the system-on-chip will further look for them in the L2 or L3 cache.

[0039] In the following embodiments, in order to improve the performance of the system-on-chip, the operation of the system-on-chip is predicted, so that instructions and / or data can be prefetched in advance. Such prefetching may include: configuring the instructions that will be needed in the future in the i-cache, and / or, configuring the data that will be needed in the future in the d-cache.

[0040] Overview of the paging mechanism

[0041] Paging is a memory space partitioning mechanism with a finer granularity than segmentation. After enabling the paging mechanism, the processor must convert the linear address into a physical address through the page management mechanism. The paging mechanism divides both the linear address space and the physical address space into fixed-size pages (usually 4KB), and maintains a set of page translation table structures through the Memory Management Unit (MMU) to complete the mapping and conversion of pages in the linear address space to pages in the physical address space.

[0042] Replacement algorithm

[0043] The Central Processing Unit (CPU) replacement algorithm refers to the algorithm used to manage the allocation of CPU time slices, mainly used to determine which programs will be preferentially allocated to the CPU when multiple programs are running simultaneously. These algorithms consider factors such as program priority, CPU load, I / O operation frequency and size. Several replacement algorithms in related technologies are introduced below.

[0044] Optimal Page Replacement Algorithm: This is a theoretical algorithm that can select the page that will not be accessed for the longest time in the future for replacement, thereby minimizing the page fault rate. However, due to its complexity and large computational amount, it is rarely used in actual systems.

[0045] First-In-First-Out (FIFO) algorithm: This algorithm replaces pages in the order they enter memory, and the page that enters first is replaced first. This algorithm is simple to implement, but may lead to the "Belady phenomenon", that is, an increase in the allocated pages may cause an increase in the page fault rate.

[0046] Least Recently Used (LRU) algorithm: This algorithm selects the page that has been least recently used for replacement. The LRU algorithm can better reflect the usage trend of pages, but requires more management overhead.

[0047] Clock algorithm: This is an approximate implementation of the LRU algorithm, which decides which pages should be replaced by simulating the ticking of a clock. This algorithm reduces the management overhead, but its performance is not as good as that of the LRU in some cases.

[0048] The technical solutions in the following embodiments can be applied to the replacement algorithm.

[0049] Translation Lookaside Buffer

[0050] The Translation Lookaside Buffer (TLB), also translated as page table cache or translation bypass cache, is a cache on the chip used by the memory management unit to improve the translation speed from virtual address to physical address. All desktop and server chips (such as x86) use TLB. TLB has a fixed number of slots to store the tagged page table entries that map virtual addresses to physical addresses. It is a typical content-addressable memory (CAM). Its search key is the virtual memory address, and its search result is the physical address. If the requested virtual address exists in the TLB, the CAM will give a very fast matching result, and then the obtained physical address can be used to access the memory. If the requested virtual address is not in the TLB, the tagged page table will be used for virtual-to-physical address translation, and the access speed of the tagged page table is much slower than that of the TLB. In some systems, the tagged page table is allowed to be swapped to secondary memory, so the virtual-to-physical address translation may take a very long time.

[0051] The technical solutions in the following embodiments can also be applied to the TLB.

[0052] Target software

[0053] The target software is software running on top of the operating system. In one example, the target software runs online in real time; the analysis software can be offline, or it can also be online in high-performance demand scenarios. The target software can be software with various functions. For example, in a big data system, the target software can be software for data processing; in a video system, the target software can be software for image processing, etc. In the following embodiments, the target software can be software that runs for a long time. For example, the license plate recognition and warning software in a video surveillance system needs to run for a long time. Optimization measures in the following embodiments can be adopted for these software, which can make the operation of the target software more efficient. To distinguish it from the target software, in the following embodiments, the software for chip information analysis processing (chip information offline processing program) is called log analysis software (abbreviated as analysis software), and this analysis software can also be called the first software (hereinafter, offline software all refers to the first software), and the target software can also be called the second software.

[0054] To solve the problems in the related art, in the following embodiments, a chip optimization processing method is provided. Figure 10 It is a flowchart of the chip optimization processing method according to the embodiments of the present application, as Figure 10 shown. The steps included in the method in Figure 10 will be described below.

[0055] Step S102, the system - on - chip located on the chip obtains a rule, where the rule reflects the behavior prediction when the system - on - chip runs the target software, and the rule is obtained by analyzing the operation information collected when the system - on - chip has run the target software historically, and the operation information is generated when the system - on - chip runs the target software.

[0056] The rule used in this step can be analyzed in various ways. As an optional implementation, the operation information can be analyzed by analyzing software. The analyzing software can process the rule through the following steps: The analyzing software receives the operation information, where the operation information is obtained through an interface pre - configured on the chip. After the operation information is read out from the interface, it is saved in a storage space (e.g., memory) outside the chip; the analyzing software analyzes the operation information and then determines the rule; the analyzing software sends the rule to the system - on - chip.

[0057] For the rule, it can be determined whether the rule needs to be further optimized according to the running situation of the system - on - chip. That is, as an optional implementation, after obtaining the rule according to the operation information, the system - on - chip runs the target software according to the rule and records the performance data when the system - on - chip runs the target software according to the rule; when the performance data meets the expectation, it is determined that the rule is a qualified rule; when the performance data does not meet the expectation, the system - on - chip continues to run the target software to collect the operation information, the analyzing software regenerates the rule according to all the collected operation information, the system - on - chip runs the target software according to the regenerated rule, records the performance data again and judges whether it meets the expectation, and so on, until the performance data meets the expectation.

[0058] Step S104, when the system - on - chip runs the target software, it takes measures in advance according to the rule, where the measures are determined according to the rule, and the measures taken in advance estimate the future operations of the system - on - chip and prepare for the operations.

[0059] Step S106, different rules are adopted when the system - on - chip runs different target software.

[0060] Through the above steps, the problems of being unable to cover other software and high cost caused by using a hardware accelerator when optimizing the processor for a specific software are solved, so that the processor performance can be optimized for different target programs.

[0061] In the above steps, there are many measures that the system - on - chip can take according to the rule. The following is an example for illustration.

[0062] Example 1. The system-on-chip determines the cache to be used in the future when running the target software according to the rule, and prepares the cache to be used in the future in advance. The cache includes at least one of the following: instruction cache, data cache.

[0063] Example 2. The system-on-chip determines the instruction branches to be used in the future when running the target software according to the rule, and prepares the instruction branches to be used in the future in advance.

[0064] Example 3. The system-on-chip determines the translation lookaside buffer to be used in the future when running the target software according to the rule, and prepares the translation lookaside buffer to be used in the future in advance.

[0065] It should be noted that the above 3 examples can be used alone or together. Of course, other measures can be taken.

[0066] For the analysis software, the computing resources for running the analysis software and the system-on-chip for running the target software can be different computing resources. For example, in the case where the chip includes multiple cores, the cores on the chip include two parts. The first part of the two parts is used to run the analysis software, and the second part of the two parts is used as the system-on-chip to run the target software. Another example is that the analysis software runs on a computing device, and the computing device is different from the computing device where the chip is located. Another example is that the target software runs on multiple chips, where the multiple chips are respectively located on different computing devices; the analysis software receives the running information generated by the system-on-chip on each of the multiple chips running the target software; the analysis software analyzes all the running information collected by the system-on-chip to obtain the rule; the analysis software transmits the rule to the system-on-chip on each of the multiple chips.

[0067] In another example, the system-on-chip can run an algorithm, and the algorithm can generate a rule. At this time, the system-on-chip obtaining the rule includes: the analysis software receives the running information; the analysis software obtains the parameters for generating the rule according to the running information; the system-on-chip receives the parameters and uses the parameters in the algorithm running on the system-on-chip to generate the rule; where the algorithm is used to generate the rule. In this example, the algorithm can be a neural network model, and the parameters can be the parameters of the neural network model or the weights of each layer.

[0068] The following will be described with reference to the accompanying drawings.

[0069] Figure 2 is a schematic diagram of optimizing the running of the target software according to an embodiment of the present application,Figure 2 The software side in Figure 1 corresponds to the software part in Figure 2 The hardware side in Figure 1 corresponds to the hardware part in Figure 2 The chip information offline processing program in Figure 2 corresponds to the above-mentioned analysis software, and the accelerated program running in real time in

[0070] corresponds to the above-mentioned target software. There are multiple cores on the hardware side, and these cores can include two parts: the first part of the cores includes core 10, core 11 to core 1M, and the second part of the cores includes core 20, core 21 to core 2N. The second part of the cores is used to run the accelerated program, and this part of the cores is called the program acceleration cores. The first part of the cores is used to run the analysis software, and this part of the cores is called the information processing cores.

[0071] The prediction work is performed by the chip information offline processing program (i.e., the analysis software) (it should be noted that in one example, in some high-performance demand scenarios, it can also be online program real-time analysis and real-time rule configuration, that is, the analysis software can be offline or online). By analyzing, the characteristics of the on-chip system when running the target software are obtained. Through these characteristics, rules (i.e., the operation rules generated on the software side) are obtained. The rule predicts the future behavior of the on-chip system when running the target software. After obtaining the rule, the rule is sent to the program acceleration cores through the memory, and the program acceleration cores perform processing according to the rule when running the target software.

[0072] In Figure 2 the analysis software is run by a part of the cores. In another implementation manner, it can also be executed by an independent host or computing resource. Figure 3 is a schematic diagram of optimizing the operation of the target software according to the embodiments of the present application Figure 2 As Figure 3 shown, in Figure 3 an independent host is used instead of Figure 2In the first part of the kernel (i.e., the information processing kernel), it should be noted that the independent host here means that the chip running the analysis software and the chip running the target software are not the same chip. This independent host can also be understood as the computing resources independent of the chip running the target software. This computing resource can be a single or cluster host, or cloud computing resources. In Figure 3 the target software (i.e., the program to be accelerated) runs through the chip, and the chip may include multiple cores, that is, Figure 3 cores 0, 1 to K in Figure 2 are used to run the target software. During the running process, the chip operation information on the hardware side is collected, and the software offline processing program (i.e., the analysis software) generates operation rules according to the operation information (the operation rules here are the same as those in

[0073] After generating the operation rules, they are sent to the system-on-chip on the chip, and the system-on-chip processes them according to the operation rules when running the target software. Figure 4 is a schematic diagram of optimizing the operation of the target software according to the embodiments of the present application Figure 3 , as Figure 4 shown, here the processors run distributively, each processor is used to run the target software, and each processor sends the operation information to the analysis software. After summarizing the operation information, the analysis software analyzes the operation information to obtain the operation rules (i.e., the guiding rules in Figure 4 ), and sends the guiding rules to the processors. It should be noted that the operation rules sent to each processor can be the same, or different guiding rules can be sent according to the attributes of the processors.

[0074] It can be seen from the above implementation manners that the software program can be divided into two parts. One part is the program to be accelerated running in real time (i.e., the target software or target program), and the other part is the software program for offline processing of the information generated by the real-time operation of the system-on-chip (i.e., the analysis software or analysis program).

[0075] In the above Figures 2 to 4 , in terms of structural division, the offline analysis program and the program to be accelerated running in real time are respectively run on the software side. On the hardware side, the operation information of the system-on-chip generated by running the target program is transmitted to the offline analysis program on the software side. The offline analysis program can run on some information processing cores of the same server like Figure 2 , or it can run on other servers or PC hosts. The memory layer is an information interaction mechanism, which can be like Figure 2 this kind of internal interaction through the memory, or it can be like Figure 3 or Figure 4Information is exchanged on other hosts through a network or other forms. In another example, the analysis software may also be run on the system-on-chip, and the system-on-chip uses different time domain resources to time-share the analysis software and the target software. Regardless of how the analysis software and the target software use different computing resources, the analysis software only needs to obtain the running information through a storage device outside the chip, and then perform analysis, and finally optimize the operation of the target software on the system-on-chip through rules. As an example, the external storage device can be a memory outside the chip, and the chip is connected to the memory. The following is an example of memory.

[0076] Memory can be a CPU off-chip storage medium or on-chip storage medium. In addition to storing the instructions and data required by the program, it also stores the CPU key information generated by the acceleration core (that is, the core that runs the target program). The software will use this key information to calculate and generate new operating rules. The rules will be sent to the memory and configured to the program acceleration core (that is, the core that runs the target program) to guide the CPU operating behavior and improve system performance.

[0077] Above Figures 2 to 4 Several optional architectures are introduced, and several optional software and hardware interaction sequences are described below. The system-on-chip runs the target software in the first period of time, and sends the operation information obtained by running the target software in the first period of time to the analysis software. The analysis software analyzes the operation information in the second period of time to obtain a first rule, and passes the first rule to the system-on-chip. The system-on-chip runs the target software according to the first rule in the third period of time, and sends the operation information generated in the third period of time to the analysis software. The analysis software analyzes all the collected operation information to generate a second rule and passes it to the system-on-chip, and so on. Alternatively, in the Nth period of time, the system-on-chip runs the target software according to the rule sent by the analysis software before, and in the Nth period of time, the analysis software analyzes the operation information obtained before to obtain a new rule; in the N+1th period of time, the system-on-chip runs the target software according to the new rule, and in the N+1th period of time, the analysis software summarizes the operation information of the Nth period of time and analyzes it again to obtain a new rule, and so on. Alternatively, after the system on chip runs the target software for a period of time, the analysis software analyzes the operation information obtained during the period to obtain rules, and sends the rules to the system on chip; thereafter, the system on chip always uses the rules.

[0078] The above-mentioned timing sequences are described below with reference to the accompanying drawings.

[0079] Figure 5 This is a schematic diagram of the interaction timing according to an embodiment of the present application.Figure 1 , in Figure 5 , the timing diagram of "intermittent processing" is shown, as Figure 5 shown. The hardware operation refers to the system-on-chip running the target software, and during the operation, information generated by the chip running the target software is collected (in Figure 5 it is hardware information or key information). After the key information is collected, the analysis software analyzes it to generate rules, and then sends the rules to the system-on-chip for hardware operation. In Figure 5 , within the L0 to L1 segment, the hardware operates based on default rules, and then the information collected within the L0 to L1 segment is sent to the analysis software. The analysis software analyzes it within the L1 to L2 segment to generate the first rule (i.e., the rule for L0 - L1). The hardware operation within the L1 to L2 segment is still based on the default rules. Within the L2 to L3 segment, the hardware operates based on the L0 - L1 rule (since this rule is optimized, the hardware in this segment will have an acceleration effect during operation, which is called accelerated operation based on the L0 - L1 rule). The key information generated within the L2 to L3 segment is collected and sent to the analysis software. The analysis software can analyze the key information collected within the L2 to L3 segment to generate new rules (which can be called the second rule), or it can also analyze the key information collected within the L0 to L3 segment to generate new rules (which can be called the third rule). The hardware still operates based on the first rule within the L2 to L3 segment and based on the second rule or the third rule within the L3 to L4 segment, and so on, cycling until the generated rules meet the expectations.

[0080] Figure 6 is the interaction timing schematic according to the embodiment of the present application Figure 2 , in Figure 6 , the timing diagram of "pipelined processing" is shown, as Figure 6 shown. The system-on-chip runs the target software from L0 to L1, collects the information generated within the L0 to L1 segment. The analysis software analyzes the hardware information within the L1 to L2 segment to obtain rules. At the same time, the system-on-chip still operates based on default rules within the L1 to L2 segment; the system-on-chip accelerates the operation based on the rules generated from L0 to L1 within the L2 to L3 segment. The analysis software processes the information from L0 to L2 within the L2 to L3 segment to obtain the operation rules. The system-on-chip accelerates the operation based on the rules generated from L0 to L2 within the L3 to L4 segment. The analysis software processes the information from L0 to L3 within the L3 to L4 segment to obtain new operation rules; the system-on-chip accelerates the operation based on the rules from L0 to L3 after the L4 segment. The analysis software processes the information generated from L0 to L4 after the L4 segment to obtain rules, and so on, until the obtained operation rules meet the expectations.

[0081] Figure 7 is the interaction timing schematic according to the embodiment of the present application Figure 3 , inFigure 7 The timing diagram of "one-time processing" is shown in Figure 7 As shown, the system-on-chip runs the target software from L0 to L1, collects information at L1, then analyzes the software from L0 to L1 to process the key information to obtain the running rules, and the system-on-chip accelerates the operation based on the L0-L1 rules after L2. In Figure 7 only one rule is generated, so it is called one-time processing.

[0082] In summary, in Figure 5 the hardware runs continuously. When key information is collected or a data request from the software is received, key information data will be generated and the data will be transmitted to the software side for processing. When a new running rule is calculated and generated, the rule will be sent to the hardware side to guide the subsequent execution and prediction of the hardware. The hardware operation and software processing are asynchronous. The software processing can be executed at intervals, and the results will be sent to the hardware as soon as they are generated. During this period, the hardware will continuously run the acceleration program. Essentially, as long as a new rule comes, it will make the hardware operation more efficient, but if no new rule comes, it will still execute according to the previous rules without blocking the hardware operation. In Figure 6 the hardware and software still execute asynchronously, but at this time the system will perform extremely efficient pipelined execution interaction. While the hardware runs continuously, the software side will also execute in real time according to the full bandwidth. The software side continuously requests or reads hardware information and updates the rules for the hardware at a high frequency. This design is suitable for scenarios where the dedicated program for hardware acceleration changes relatively frequently. In Figure 7 the offline processing software only needs to process the hardware information once within a long period of time and configure it for the hardware. This design is suitable for scenarios where the same type of dedicated program runs for a long time.

[0083] Figures 5 to 7 Regardless of how the timing is processed in Figure 8 is the interaction flowchart according to the embodiment of the present application, as shown in Figure 8As shown, the interaction process includes the following steps: First, the program to be accelerated (i.e., the target software) starts. After the target software starts, it runs on the chip. During the running process, information generated when the chip runs the target software (i.e., key information) is collected. This collection process can also start after receiving a request message from the analysis software. It should be noted that there is a cache inside the chip, which is used to store this key information. However, the size of the cache is limited. Therefore, it is necessary to grab the key information in time and save it in memory or other storage spaces, so that a relatively large amount of key information can be saved. Then, the collected key information is analyzed by the analysis software to generate rules. Since these rules guide the operation of the hardware of the chip, these rules are also called hardware guidance rules. These hardware guidance rules are sent to the chip and used by the system-on-chip on the chip when running the target software. Briefly speaking, Figure 8 The following steps are reflected: 1. Start; 2. The program to be accelerated starts; 3. The chip hardware runs the program to be accelerated; 4. The hardware generates key information, and if generated, it is transmitted to the software; 5. Run the information processing software program (i.e., the analysis software); 6. The analysis software generates rules and writes them to the system-on-chip of the chip to achieve acceleration.

[0084] Figure 9 It is a schematic diagram of the interaction process according to an embodiment of the present application. As Figure 9 shown, the analysis software analyzes the information from the system-on-chip through software analysis algorithms. Since the information obtained from the system-on-chip is obtained according to different target software, the rules generated by different target software may be different. Therefore, corresponding rules can be generated for the information generated by the operation of different target software on the system-on-chip. That is, this is a customized processing for the target software. The system-on-chip includes a central processing unit and a cache. The central processing unit is used to execute the target software, and the cache is used to save the instructions and data corresponding to the target software. Then, the information generated by the system-on-chip when executing the target software (the so-called real-time information is to continuously grab information from the cache and save it in a larger storage space such as memory). After the analysis software obtains rules based on the information, the central processing unit can obtain the operations to be performed in the future according to the rules, so as to perform pre-processing (therefore, the rules can also be understood as a prediction of future behaviors). In this way, the efficiency of the system-on-chip can be improved. In Figure 9Among them, the central processing unit executes the program to be accelerated. The system-on-chip stores the key information generated in real time (i.e., real-time information) in the memory. The analysis software will develop customized analysis algorithms according to different dedicated programs. The input of the algorithm is the real-time information generated by the hardware. The software will generate the future trajectory guidance rules required for hardware execution and simultaneously issue the location execution rules that can be located. When the system-on-chip executes the target software, the search rule for the address in the memory has a higher information density, occupies a smaller area, can reduce the search time complexity, and avoid the bandwidth loss and increased search latency caused by directly reading the execution rules. The system-on-chip will issue a search request according to the accuracy of the internal prediction and generate the information required for the search. The search information and the search rule are matched and searched to find the segment address pointer of the corresponding rule. After reading the hardware execution rule, the information is extracted from it, and the prediction information is input to the system-on-chip for subsequent prediction behavior.

[0085] In the above embodiment, an interface can be pre-configured in the processor (or chip), and the information generated during the hardware execution of the target software can be obtained through this interface. That is, the hardware side will expose the information interface, and the software side reads the information and performs offline processing. By using the flexibility of the software, specific optimizations for different programs can be achieved to reach the optimal performance of each program. Traditional processors do not perform specific optimizations for each program, or rather, the processor itself cannot be specifically optimized. Only software personnel can optimize the program or compiler to adapt to the processor to achieve better performance. In the above embodiment, the processor can adapt to the program. The specific method is that the hardware remains unchanged because the cost of modifying the hardware is very high. However, by making full use of the flexibility of the software processing algorithm, different rules can be designed according to the running program (this rule can be understood as the rule dedicated to the target software or target program), and program-specific rules are generated to guide the hardware operation and improve the performance.

[0086] In an alternative embodiment, the software-side offline algorithm will analyze and process the real-time information provided by the processor hardware. The processed rules can be discrete segments independent of timing or long historical segments with strong temporal and spatial locality. These rules can predict the future running behavior of the processor. However, to avoid occupying too much storage space, a compression algorithm can also be used to compress the rules, so that more rules can be stored with less storage space, realizing the prediction architecture design with long historical information and guidance information.

[0087] In the above embodiments, software-hardware co-design such as software offline analysis is adopted, which requires using the information generated by running the target program on the system-on-chip (this information is used for future prediction and is therefore also called historical information). Since the cache of the processor for historical information is usually only stored in the on-chip cache (RAM or SRAM), there is less historical information, making it impossible to perform analysis and processing over a long historical range. For example, dynamic branch prediction algorithms such as Branch History Table (abbreviated as BHT), GShare, and TAgged GEometric history (abbreviated as TAGE) mainly store historical information on SRAM, and the size of SRAM determines the length of historical information that can be stored. Therefore, the traditional approach generally improves performance by increasing the area of the processor. The above embodiments use on-chip or off-chip memory to store long historical information as the prediction data source. The processor hardware exposes an information interface to the offline analysis software, and the offline analysis software uses its flexibility to perform specific analysis on the dedicated program (i.e., the target software), providing rules for guiding the future operation of the system-on-chip. The hardware performs high-performance prediction and execution based on the rules.

[0088] Through the above embodiments, the offline software uses the long historical information provided by the hardware for analysis and processing, ensuring prediction accuracy and improving the execution performance of the processor with prediction accuracy. The above embodiments ensure the flexibility of the software-hardware system without changing the hardware, saving costs.

[0089] In the above embodiments, based on the upload of hardware information and the offline processing on the software side to give hardware execution rules, the following aspects can be accelerated: branch prediction; i-cache, d-cache prefetch to improve cache hit rate; TLB prefetch to improve TLB hit rate and avoid directly reading the page table. In the above embodiments, by predicting the time node and data itself of multi-core data update, the performance and power consumption losses caused by repeated updates due to the broadcast mechanism in traditional multi-core cache coherence are solved. The future prediction trajectory can guide the prefetch behavior in the local host cache in the distributed cluster, thereby effectively reducing the impact of network latency on system performance and improving performance.

[0090] The computer program corresponding to the above analysis software can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate computer-implemented processing. Thus, the instructions executed on the computer or other programmable device provide for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1Steps for specifying functions in one or more boxes can be implemented by different modules corresponding to different steps. Such a device or system is provided in this embodiment. The system or device is used to implement the functions of the method in the above embodiment. Each module in the system or device corresponds to each step in the method and has been described in the method, so it will not be elaborated here.

[0091] In one embodiment, a chip is further provided, including a central processing unit and a cache; wherein, the system on chip on the chip is used to execute the following method steps: obtaining a rule, wherein the rule reflects the behavior prediction when the system on chip runs the target software, and the rule is obtained by analyzing the operation information collected when the system on chip historically runs the target software, and the operation information is generated when the system on chip runs the target software; when running the target software, taking measures in advance according to the rule, wherein the measures are determined according to the rule, and the measures taken in advance estimate the future operations of the system on chip and prepare for the operations.

[0092] As an optional implementation manner, it further includes: an interface, wherein the interface is one or more, and the interface is used to send the operation information to the outside of the system on chip, and receive the rule from the analysis software or receive the parameter from the analysis software; wherein, the parameter is used in the algorithm running on the system on chip to generate the rule, and the algorithm is used to generate the rule.

[0093] For the above chip, in one embodiment, a chip optimization processing system is further provided. The system includes the above chip and analysis software. The analysis software is used to execute the following method steps: receiving operation information, wherein the operation information is obtained through an interface pre-configured on the chip, and after the operation information is read out from the interface, it is stored in a storage space outside the chip; determining the rule after analyzing the operation information; sending the rule to the system on chip; or obtaining a parameter for generating the rule according to the operation information and sending the parameter to the system on chip; wherein, the parameter is used in the algorithm running on the system on chip to generate the rule; the algorithm is used to generate the rule. Other steps executed by the analysis software have been described in the above embodiments and will not be elaborated here one by one. For example, the computing resources for running the analysis software and the system on chip for running the target software are different computing resources.

[0094] Optionally, when multiple cores are included on the chip, the cores on the chip include two parts. The first part of the two parts is used to run the analysis software, and the second part of the two parts runs the target software as a system-on-chip. Alternatively, the analysis software runs on a computing device different from the computing device where the chip is located. Alternatively, the analysis software runs in the system-on-chip, and the system-on-chip runs the analysis software and the target software by sharing time domain resources.

[0095] Through the above embodiments, the problems of being unable to cover other software and high cost caused by using a hardware accelerator when optimizing a processor for a specific software are solved, so that the processor performance can be optimized for different target programs.

[0096] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A chip optimization processing method, characterized in that: include: A system-on-chip acquisition rule located on a chip, wherein the rule reflects a prediction of the behavior of the system-on-chip when running the target software, and the rule is obtained by analyzing the operation information collected when the system-on-chip historically ran the target software, and the operation information is generated by the system-on-chip running the target software; When the system on chip runs the target software, it takes measures in advance according to the rules, wherein the measures are determined according to the rules, and the measures taken in advance predict the future operations of the system on chip and prepare for the operations.

2. The method according to claim 1, characterized in that: The method further comprises: the analysis software receives the operation information, wherein the operation information is obtained through an interface pre-configured on the chip, and after the operation information is read from the interface, it is stored in a storage space outside the chip; the analysis software determines the rule after analyzing the operation information; the analysis software sends the rule to the system on chip; or The on-chip system acquires the rules, including: the analysis software receives the operation information; the analysis software obtains parameters for generating the rules according to the operation information; the on-chip system receives the parameters and uses the parameters in an algorithm running on the on-chip system to generate the rules; wherein the algorithm is used to generate the rules.

3. The method according to claim 2, characterized in that The computing resources for running the analysis software and the system on chip for running the target software are different computing resources.

4. The method according to claim 3, characterized in that: In the case where the chip includes multiple cores, the core on the chip includes two parts, a first part of the two parts is used to run the analysis software, and a second part of the two parts runs the target software as a system on chip; or The analysis software is run on a computing device that is different from the computing device where the chip is located; or The analysis software runs in the system on chip, and the system on chip uses different time domain resources to time-share the analysis software and the target software.

5. The method according to claim 4, characterized in that The target software runs on multiple chips, wherein the multiple chips are respectively located on different computing devices; the analysis software receives the operation information generated by the system-on-chip on each of the multiple chips when the target software runs; the analysis software analyzes the collected operation information generated by all the system-on-chips to obtain the rules; and the analysis software passes the rules to the system-on-chip on each of the multiple chips.

6. A chip comprising a central processing unit and a cache; wherein: The system on chip on the chip is used to perform the following method steps: Acquiring rules, wherein the rules reflect the prediction of the behavior of the system-on-chip when running the target software, and the rules are obtained by analyzing the operation information collected when the system-on-chip ran the target software in the past, and the operation information is generated by the system-on-chip running the target software; When the target software is running, measures are taken in advance according to the rules, wherein the measures are determined according to the rules, and the measures taken in advance predict the future operations of the system on chip and prepare for the operations.

7. The chip according to claim 6, characterized in that: It also includes: an interface, wherein the interface is one or more, and the interface is used to send the operating information to the outside of the system on chip, and to receive rules from the analysis software or receive parameters from the analysis software; wherein the parameters are used in the algorithm running on the system on chip to generate the rules, and the algorithm is used to generate the rules.

8. A chip optimization processing system, characterized in that: The method comprises the chip and analysis software according to claim 6 or 7, wherein the analysis software is used to perform the following method steps: Receive operation information, wherein the operation information is obtained through an interface pre-configured on the chip, and after the operation information is read from the interface, it is stored in a storage space outside the chip; determine the rule after analyzing the operation information; send the rule to the system on chip; or, obtain parameters for generating the rule based on the operation information, and send the parameters to the system on chip; wherein the parameters are used in an algorithm running on the system on chip to generate the rule; the algorithm is used to generate the rule.

9. The system according to claim 8, characterized in that The computing resources for running the analysis software and the system on chip for running the target software are different computing resources.

10. The system according to claim 9, characterized in that In the case where the chip includes multiple cores, the core on the chip includes two parts, a first part of the two parts is used to run the analysis software, and a second part of the two parts runs the target software as a system on chip; or The analysis software is run on a computing device that is different from the computing device where the chip is located; or The analysis software runs in the system on chip, and the system on chip uses different time domain resources to time-share the analysis software and the target software.

Citation Information

Patent Citations

  • Context-based memory indirect branch target prediction

    CN114647447A

  • System on chip

    CN117951073A

  • System on chip and operation method thereof

    CN118193448A

  • Constrained metric optimization of a system on chip

    US10607039B1

  • Microprocessor with multiple operating modes dynamically configurable by a device driver based on currently running applications

    US20100011198A1

Cited By

  • Chip optimization processing method and system, and chip

    WO2026158326A1