Method for caching command data in a coprocessor and electronic device

CN122884871APending Publication Date: 2026-10-09GETONG INTELLIGENT TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611384190.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-09-08
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

[0003]然而,IRAM的存储容量和数据供给能力有限,随着报文传输速率的提高,IRAM难以及时提供协处理器所需的全部命令数据

Benefits of technology

本申请提供一种协处理器中命令数据的缓存方法及电子设备。该方法获取进入协处理器的命令数据,维护全局累计值、最近运行值和运行间隔值,结合权值及补偿系数计算命令热度值;根据协处理器一次处理的多条命令数据构建候选命令组合并编码为染色体,基于报文组合出现次数和输出地址相对距离计算适应值;通过遗传优化获得优化后的候选命令组合;结合组合适应信息和命令热度信息确定缓存优先级,并按时间段更新外部缓存器。从而可提高缓存命中率,降低命令读取延迟,提升报文处理速率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122884871A_ABST
    Figure CN122884871A_ABST
Patent Text Reader

Abstract

The application provides a cache method of command data in a coprocessor and an electronic device. The method obtains command data entering the coprocessor, maintains a global cumulative value, a recent running value and a running interval value, calculates a command heat value in combination with a weight value and a compensation coefficient; constructs a candidate command combination according to a plurality of command data processed by the coprocessor once and encodes the candidate command combination as a chromosome, calculates an adaptive value based on the number of occurrences of the message combination and the relative distance of the output address, obtains an optimized candidate command combination through genetic optimization, determines a cache priority in combination with the adaptive information of the combination and the command heat information, and updates an external cache according to time periods. Thus, the cache hit rate can be improved, the command reading delay can be reduced, and the message processing rate can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of data communication and cache control, and in particular to a method and electronic device for caching command data in a coprocessor. Background Technology

[0002] Currently, in network devices such as switches, coprocessors control the scheduling, forwarding, and output of packet data by reading command data. One existing solution stores command data in internal random access memory (IRAM) and reads the required command data from the IRAM when the coprocessor needs to perform a corresponding operation. Another existing solution uses an external cache to store a portion of the command data in advance and reads the command data from the external cache when the coprocessor needs it.

[0003] However, IRAM has limited storage capacity and data supply capability. As message transmission rates increase, IRAM struggles to provide all the command data required by the coprocessor in a timely manner. While external buffers offer higher read speeds, their storage space is also limited, unable to store all command data. When the command data required by the coprocessor is not stored in the external buffer, a cache miss occurs, forcing the coprocessor to read command data from a slower storage location, thus increasing message processing latency and reducing message transmission rates. Existing caching strategies typically select cached content based solely on fixed priority or simple access frequency, failing to simultaneously reflect the time of the most recent use of command data, the frequency of command data calls, and the combined relationships between multiple command data items. Furthermore, the relationship between command combinations and message output results is usually non-linear and discontinuous, making it difficult for conventional analysis methods to accurately identify command combinations that improve cache hit rates. Summary of the Invention

[0004] This application provides a method and electronic device for caching command data in a coprocessor, which can improve cache hit rate, reduce command read latency, and increase message processing speed.

[0005] In a first aspect, embodiments of this application provide a method for caching command data in a coprocessor, including: Obtain command data to enter the coprocessor; For each type of command data, maintain a global cumulative value, a recent run value, and a run interval value; wherein, the global cumulative value is used to represent the cumulative number of runs of the command data; the recent run value is used to represent the most recent run position of the command data; and the run interval value is used to represent the interval between two consecutive runs of the command data. Based on the most recent running value, the running interval value, the weight corresponding to the command data, and the compensation coefficient, the command popularity value corresponding to the command data is calculated; the command popularity value is used to characterize the probability that the corresponding command data will be called again by the coprocessor in the near future. Based on the multiple command data used by the coprocessor in a single processing step, multiple candidate command combinations are constructed, and each candidate command combination is encoded into a corresponding chromosome. For each chromosome, the occurrence count of the message combination corresponding to the chromosome and the relative distance of the output address of the message combination are obtained, and the fitness value corresponding to the chromosome is calculated based on the occurrence count and the relative distance of the output address. Based on the fitness value corresponding to each chromosome, genetic optimization is performed on multiple chromosomes to obtain an optimized combination of candidate commands. Based on the optimized candidate command combinations and the command popularity value corresponding to the command data contained in the candidate command combinations, the cache priority of each candidate command combination is determined, and the candidate command combinations that meet the cache conditions are written into the external cache of the coprocessor according to different time periods.

[0006] Secondly, embodiments of this application provide an electronic device, including: The communication interface is used to receive message data to be processed and to send processed message data. The coprocessor is used to schedule, forward, and output message data received from the communication interface based on command data; An external cache is used to store optimized combinations of candidate commands; Memory, used to store one or more programs; processor; A communication bus is used to connect the communication interface, the coprocessor, the external buffer, the memory, and the processor. When one or more programs are executed by the processor, the methods in the first aspect and various possible implementations described above are implemented, and the optimized candidate command combination is written to the external cache so that the coprocessor reads the candidate command combination from the external cache.

[0007] The technical solutions provided in the embodiments of this specification may include the following beneficial effects: This application provides a method and electronic device for caching command data in a coprocessor. The method acquires command data entering the coprocessor, maintains a global cumulative value, a recent execution value, and an execution interval value, and calculates a command popularity value by combining weights and compensation coefficients. It constructs candidate command combinations based on multiple command data processed by the coprocessor at one time and encodes them as chromosomes, calculating fitness values ​​based on the frequency of occurrence of message combinations and the relative distance to output addresses. It obtains optimized candidate command combinations through genetic optimization. Finally, it determines cache priority by combining combination fitness information and command popularity information, and updates the external cache according to time periods. This improves cache hit rate, reduces command read latency, and increases message processing speed.

[0008] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description

[0009] Figure 1 A flowchart illustrating a method for caching command data in a coprocessor, provided as an embodiment of this application; Figure 2 A flowchart illustrating another method for caching command data in a coprocessor provided in an embodiment of this application; Figure 3 A schematic diagram of the most recent running value and running interval value provided in an embodiment of this application; Figure 4 A flowchart illustrating another method for caching command data in a coprocessor provided in an embodiment of this application; Figure 5 A flowchart illustrating another method for caching command data in a coprocessor provided in an embodiment of this application; Figure 6 A flowchart illustrating another method for caching command data in a coprocessor provided in an embodiment of this application; Figure 7 A flowchart illustrating another method for caching command data in a coprocessor provided in an embodiment of this application; Figure 8 A flowchart illustrating another method for caching command data in a coprocessor provided in an embodiment of this application; Figure 9 A flowchart illustrating another method for caching command data in a coprocessor provided in an embodiment of this application; Figure 10 A flowchart illustrating another method for caching command data in a coprocessor provided in an embodiment of this application; Figure 11 This is a schematic diagram of the hardware architecture of an electronic device provided in an embodiment of this application. Detailed Implementation

[0010] The exemplary embodiments will now be described in detail. When the description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification; they are merely exemplary embodiments of apparatuses and methods consistent with some aspects of this specification.

[0011] In high-speed data exchange scenarios, the coprocessor's scheduling, forwarding, and output of message data all rely on corresponding command data. For the coprocessor, command data is used not only to indicate the processing order of message data but also to control the specific operations of message data at different processing stages. Therefore, whether command data can be provided in a timely and continuous manner directly affects the coprocessor's message processing efficiency. When the supply rate of command data is lower than the coprocessor's processing speed, even if message data has arrived at the coprocessor, the coprocessor may wait due to the lack of corresponding command data, thus affecting the overall message data transmission rate.

[0012] From the perspective of command data storage, one existing solution stores command data in internal random access memory (IRAM). While IRAM can provide command data to the coprocessor, its storage capacity and data supply capability are limited. As message transmission rates increase, the amount of command data entering the coprocessor per unit time increases, and the types and quantities of command data the coprocessor needs to read also increase. In this situation, IRAM struggles to provide all the command data required by the coprocessor in a timely manner, easily creating a command data supply bottleneck.

[0013] Another existing approach involves setting up a cache outside the coprocessor to pre-store some command data. The external cache offers high read speeds; when the command data required by the coprocessor is already stored in the external cache, the coprocessor can quickly read that data, thus reducing waiting time. However, the storage space of the external cache is also limited, and it cannot store all command data simultaneously. Therefore, whether the external cache can improve the efficiency of command data provisioning depends not only on the read speed of the external cache itself, but also on whether the command data stored in the external cache is the command data that the coprocessor actually needs in the near future.

[0014] A cache miss event occurs when the command data required by the coprocessor is not stored in the external cache. At this time, the coprocessor needs to read the corresponding command data from a slower storage location. This process increases the command data read time, preventing the coprocessor from completing the corresponding message processing operation in a timely manner while waiting for the command data. As the number of cache miss events increases, the command data read latency accumulates, eventually leading to increased message processing latency and a decrease in message transmission rate. Therefore, improving the cache hit rate of the external cache is key to improving the coprocessor's message processing efficiency.

[0015] Furthermore, existing caching strategies typically select command data to be cached based on fixed priority or simple access frequency. Fixed priority fails to reflect the actual changes in command data calls over different time periods, while simple access frequency mainly reflects the cumulative number of times command data has been called, failing to simultaneously reflect the time of the most recent use of command data and the interval between two adjacent calls. For example, if a command data point has a high historical access count but has not been called for a long time recently, the likelihood of it being called again in the short term is likely low; conversely, if a command data point has not the highest cumulative access count but has been called consecutively recently, it may require higher priority caching in the short term. Therefore, relying solely on fixed priority or simple access frequency makes it difficult to accurately determine the likelihood of command data being called again by the coprocessor in the near future.

[0016] Furthermore, coprocessors typically require multiple command data points during a single processing run, and these commands may have combined relationships. For a single command, its "command heat value" reflects the likelihood of it being invoked again; however, for command combinations formed by multiple commands, simply sorting them based on their heat values ​​is insufficient to accurately evaluate their actual effectiveness. Different command combinations may have varying impacts on the output address of the message combination, and the relationship between command combinations and message output results may be non-linear and discontinuous. Therefore, conventional linear analysis methods are insufficient to accurately identify command combinations that improve the external cache hit rate.

[0017] Therefore, the core technical problem this application needs to solve is not simply expanding the storage space of the external buffer, but rather, given the limited storage space of the external buffer, how to more accurately select the command data and command combinations that should be written to the external buffer. Specifically, this application needs to solve the following problems: how to quantify the probability of command data being called again in the near future based on both the recent usage time and the frequency of invocation of the command data; how to evaluate the merits of candidate command combinations based on the impact of candidate command combinations formed by multiple command data on the message output results; how to select optimized candidate command combinations from multiple candidate command combinations; and how to dynamically update the candidate command combinations in the external buffer at different time periods based on the optimization results and command popularity values.

[0018] To address the aforementioned issues, the core improvement approach of this application is as follows: First, for each type of command data, maintain a global cumulative value, a recent execution value, and an execution interval value. Combine this with the weights and compensation coefficients corresponding to the command data to calculate a command popularity value, enabling the command popularity value to simultaneously reflect the recent call status and frequency of command data. Second, introduce a Genetic Algorithm (GA) to screen multiple candidate command combinations: each candidate command combination is encoded as a chromosome, and its impact on message output is quantitatively evaluated using fitness values. Candidate command combinations are constructed based on multiple command data used by the coprocessor in a single processing step, and these combinations are encoded as chromosomes. The fitness value / fitness of the chromosomes is calculated based on the frequency of occurrence of the message combination and the relative distance to the output address. The candidate command combinations are then quantitatively evaluated based on the message output results. Third, genetic optimization processes are used to select, crossover, mutate, and iterate multiple chromosomes to obtain optimized candidate command combinations. Finally, cache priorities are determined based on the optimized candidate command combinations and the command popularity values ​​corresponding to the command data contained within each candidate command combination. Candidate command combinations that meet the cache conditions are then written into an external cache according to different time periods.

[0019] Based on the core improvement approach described above, the recent invocation probability of a single command, the combined correlation between multiple command data, and the impact of candidate command combinations on message output can all contribute to determining the cached content in the external buffer. Specifically, regarding the non-linear and discontinuous relationship between command combinations and message output, genetic optimization does not require modeling or differentiation of this relationship. It only evaluates candidate command combinations based on fitness values, approximating the optimization result through iterative search of selection, crossover, and mutation. Simultaneously, mutation allows the search to escape locally optimal candidate command combinations, reducing the risk of optimization results concentrating in a localized range. Therefore, the optimized candidate command combinations reflect both the recent invocation of a single command and the actual impact of multiple command combinations on message output. When external buffer storage space is limited, the external buffer can prioritize storing candidate command combinations that are more likely to be called by the coprocessor recently and have a smaller impact on message output, thereby improving the external buffer's cache hit rate, reducing command data read latency caused by coprocessor cache misses, and increasing message data processing and transmission rates.

[0020] Optionally, Figure 1 This is a flowchart illustrating a method for caching command data in a coprocessor, as provided in an embodiment of this application. See also... Figure 1 The method includes: Step 100: Obtain the command data for entering the coprocessor.

[0021] Optionally, after packet data enters the switch, the coprocessor regulates the packet data according to command data. The command data entering the coprocessor may include commands for controlling packet scheduling, forwarding, or output.

[0022] Step 101: For each type of command data, maintain the global cumulative value, the most recently run value, and the run interval value.

[0023] Among them, the global cumulative value is used to represent the cumulative number of command data runs; the most recent run value is used to represent the most recent run position of the command data; and the run interval value is used to represent the interval between two consecutive runs of the command data.

[0024] Optionally, the most recently executed value is the logical sequence number position, that is, the global cumulative value corresponding to the most recent execution of the command data. Furthermore, the interval between two adjacent executions of the command data represents the logical count interval, that is, the number of command data between two adjacent executions, expressed as the difference of the global cumulative values.

[0025] Optionally, in the following example, the global cumulative value is represented by an S value, the most recently run value by an O value, and the run interval value by a D value. From the start of system operation, the cumulative number of command data entering the coprocessor can be recorded, and a corresponding most recently run value and run interval value can be maintained for each type of command data.

[0026] Step 102: Calculate the command popularity value corresponding to the command data based on the most recent running value, running interval value, weight corresponding to the command data, and compensation coefficient.

[0027] The command heat value is used to characterize the likelihood that the corresponding command data will be called again by the coprocessor in the near future.

[0028] Optionally, a larger recent execution value indicates that the most recent execution time of the corresponding command data is closer to the current time; a smaller execution interval value indicates that the time interval between two consecutive executions of the corresponding command data is shorter, and the call is more frequent. The command popularity value combines the above execution information with the weights and compensation coefficients corresponding to the command data to evaluate the activity level of the command data.

[0029] Step 103: Based on the multiple command data used by the coprocessor in one processing step, construct multiple candidate command combinations and encode each candidate command combination into a corresponding chromosome.

[0030] Optionally, the coprocessor's data processing bit width can be 128 bits, and the bit width of a single command data can be 32 bits or 64 bits. Therefore, the coprocessor can use multiple command data in a single processing step. By constructing candidate command combinations from the multiple command data used in a single processing step and encoding these candidate command combinations as chromosomes, subsequent genetic optimization processing can be performed on a unit basis of command combinations.

[0031] Step 104: For each chromosome, obtain the occurrence count of the message combination corresponding to the chromosome and the relative distance of the output address of the message combination, and calculate the fitness value corresponding to the chromosome based on the occurrence count and the relative distance of the output address.

[0032] Optionally, candidate command combinations can affect message output results. The frequency of occurrence of message combinations reflects the relative occurrence of the corresponding output scenario, while the relative distance to the output address reflects the degree of influence of candidate command combinations on the message output address. By calculating the fitness value using these two types of information, candidate command combinations can be evaluated in reverse from the message output results.

[0033] Optionally, the output address is the output destination identifier (e.g., output port number or destination address identifier) ​​after the message is processed by the coprocessor; the relative distance of the output address is the deviation between the output address that the message should arrive at in sequence and the actual output address.

[0034] Step 105: Based on the fitness value of each chromosome, perform genetic optimization on multiple chromosomes to obtain optimized candidate command combinations.

[0035] Optionally, since the relationship between command combinations and output message data can be non-linear and discontinuous, this embodiment uses genetic optimization to screen multiple candidate command combinations. Genetic optimization may include processes such as generating an initial population, selection, crossover, mutation, and iterative convergence.

[0036] Optionally, fitness value serves as the basis for genetic optimization, forming an iterative closed loop. Specifically, chromosomes are sorted according to their fitness values ​​and a selection process is performed; the higher the fitness value ranking, the higher the probability of selection. Using the chromosome with the highest fitness value as a control, parent chromosomes are selected and crossover is performed based on the fitness value difference and ratio. The mutation probability is determined based on the fitness value ranking, and mutation is performed to form a new generation population. If the termination condition is not met, the fitness value is recalculated until the difference between the highest fitness values ​​of two adjacent generations is less than a preset threshold. The chromosome with the highest fitness value is then selected for decoding, resulting in the optimized candidate command combination.

[0037] Step 106: Based on the optimized candidate command combinations and the command heat value corresponding to the command data contained in the candidate command combinations, determine the cache priority of each candidate command combination, and write the candidate command combinations that meet the cache conditions into the external cache of the coprocessor according to different time periods.

[0038] Optionally, the invocation status of command data and candidate command combinations may change over different time periods. By determining the cache priority based on the optimized candidate command combinations and command popularity values, and updating the external cache according to different time periods, the external cache can store different candidate command combinations at different times.

[0039] In this embodiment of the invention, the recent execution status and adjacent execution interval of command data are converted into command popularity values; multiple command data are then constructed into candidate command combinations and encoded as chromosomes; fitness values ​​are calculated based on the occurrence frequency of message combinations and the relative distance of output addresses; genetic optimization is then performed on the chromosomes; finally, cache priority is determined based on the optimized candidate command combinations and command popularity values, and the external cache is updated according to different time periods. Therefore, the external cache can store candidate command combinations that are more likely to be invoked recently, thereby helping to improve the cache hit rate of the external cache and increase the processing speed of message data.

[0040] Optionally, in the above example, step 102 requires calculating the command heat value based on the most recent running value and the running interval value. Therefore, the accuracy of maintaining the global cumulative value, the most recent running value, and the running interval value in step 101 directly affects the calculation result of the command heat value. Especially when target command data repeatedly enters the coprocessor, if the most recent running value is updated before calculating the running interval value, the running interval value may be incorrectly calculated as zero. Based on this, this embodiment performs parameter maintenance on target command data entering the coprocessor for the first time and for non-first-time entries, and limits the order between calculating the running interval value and updating the most recent running value, thereby providing accurate running statistics parameters for subsequent command heat value calculation. Specifically, in Figure 1 On this basis, Figure 2 A flowchart illustrating another method for caching command data in a coprocessor provided in this application embodiment. See also... Figure 2 Step 100 includes: Step 100-1: After obtaining the data for each command, update the global cumulative value.

[0041] Optionally, the global cumulative value is used to represent the cumulative number of command data that have entered the coprocessor since the system started running.

[0042] Step 100-2: When the target command data first enters the coprocessor, initialize the most recent running value and running interval value of the target command data.

[0043] Optionally, when the target command data first enters the coprocessor, the current global cumulative value is used as the most recent running value of the target command data, and the running interval value of the target command data is initialized to infinity.

[0044] Step 100-3: When the target command data enters the coprocessor for the first time, calculate the running interval value based on the current global cumulative value and the most recently run value that has been stored.

[0045] Optionally, when the target command data enters the coprocessor for the first time, the difference between the current global cumulative value and the most recently stored run value is used as the run interval value. The smaller the run interval value, the fewer command data items are stored between two consecutive runs of the target command data.

[0046] Step 100-4: After calculating the running interval value, update the most recent running value of the target command data.

[0047] Optionally, after calculating the running interval value, the current global cumulative value is updated to the most recent running value of the target command data. By calculating the running interval value first and then updating the most recent running value, the current global cumulative value can be prevented from prematurely overwriting the previous running position.

[0048] Optionally, Figure 3 This is a schematic diagram of a recent running value and a running interval value provided in an embodiment of this application. Figure 3 This includes a cache indication area, a data queue, and a final status table. The cache indication area represents the command data pages currently participating in the runtime statistics; the data queue represents the command data sequentially entering the coprocessor; and the final status table represents the most recent run value and run interval value for each command data page after all command data in the data queue has been processed.

[0049] Figure 3 The page number in the code identifies the command data. For example, page number 1 represents command data 1, page number 2 represents command data 2, page number 3 represents command data 3, page number 4 represents command data 4, and page number 5 represents command data 5. The page number is only used to distinguish different types of command data and does not indicate the order in which the command data enters the coprocessor.

[0050] Figure 3 The letter "O" in the table represents the most recently run value, which indicates the global cumulative value corresponding to the most recent run of the command data for the corresponding page number. Figure 3 The letter "D" in the code represents the execution interval value, used to indicate the interval between two consecutive executions of command data for a corresponding page number. When the command data for the corresponding page number enters the coprocessor for the first time, D is infinite (D=∞); when the command data for the corresponding page number does not enter the coprocessor for the first time, D is calculated based on the current global cumulative value and the most recent execution value before the update.

[0051] Figure 3 The final state in the data queue refers to the most recent running value O and the running interval value D corresponding to each page number after all command data in the data queue has completed the update of running statistics parameters. The final state can reflect the most recent running position and adjacent running interval of each type of command data after the entire data queue has been processed.

[0052] Taking the sequence of command data 1, command data 2, command data 3, command data 4, command data 3, command data 1, and command data 5 entering the coprocessor as an example, after each command data is acquired, the global cumulative value is updated sequentially to 1, 2, 3, 4, 4, 4, and 5. The corresponding update process is shown in Table 1 below: Table 1: Examples of recent run values ​​and run interval values ​​updates

[0053] As can be seen from the above example, when command data 1 first enters the coprocessor, the most recently run value is 1 and the run interval value is infinite (D=∞). When command data 1 enters the coprocessor again, the current global cumulative value is still 4, the stored most recently run value is 1, so the run interval value is 3, and then the most recently run value is updated to 4.

[0054] When command data 3 first enters the coprocessor, the most recently run value is 3 and the run interval value is infinite. When command data 3 enters the coprocessor again, the current global cumulative value is 4, the stored most recently run value is 3, so the run interval value is 1, and then the most recently run value is updated to 4.

[0055] Command data 2, command data 4, and command data 5 each entered the coprocessor only once in the above example. Therefore, the execution interval values ​​of the three are all infinite, and the most recent execution values ​​are 2, 4, and 5, respectively.

[0056] After the data queue is processed as described above, the final status of each page number is shown in Table 2 below: Table 2: Example of the final state after data queue processing is completed

[0057] In this context, page number 1 corresponds to a final state of O equal to 4 and D equal to 3, indicating that the global cumulative value of command data 1 during its most recent run is 4, and the interval between two consecutive runs is 3. Page number 3 corresponds to a final state of O equal to 4 and D equal to 1, indicating that the global cumulative value of command data 3 during its most recent run is 4, and the number of new command data entering the coprocessor between two consecutive runs is 1. Page numbers 2, 4, and 5 all correspond to D being infinite (D=∞), indicating that command data 2, 4, and 5 are each run only once in the data queue. These most recent run values ​​and run interval values ​​will serve as input parameters for calculating the command heat value in step 102. The most recent run value reflects the position of the most recent run of the command data, and the run interval value reflects the interval between two consecutive runs of the command data.

[0058] This embodiment updates the global cumulative value after acquiring each command data entry, ensuring a unified cumulative reference for the execution order of each command data entry. It initializes the most recent execution value and execution interval value when the target command data first enters the coprocessor, preventing the first execution of command data from lacking historical execution positions. Furthermore, it calculates the execution interval value based on the current global cumulative value and the stored most recent execution value before updating the most recent execution value on subsequent entries, preventing the most recent execution value from being prematurely overwritten. Therefore, it can accurately distinguish between the first execution and repeated executions of command data, and accurately obtain execution statistics parameters reflecting the most recent execution position and adjacent execution intervals, providing a data foundation for subsequently distinguishing between recent and frequent calls.

[0059] Optionally, relying solely on the cumulative number of accesses to command data is insufficient to reflect the time when the command data was most recently invoked; similarly, relying solely on the time of the most recent invocation is insufficient to reflect whether the command data is frequently invoked. Furthermore, different command data may have varying degrees of importance within the system, and relying solely on the execution interval may be affected by changes in command distribution over the statistical period. Therefore, this embodiment uses the most recent execution value, execution interval value, system weight, and compensation coefficient together to calculate the command popularity value. This allows the command popularity value to simultaneously reflect the recent invocation status, frequency of invocation, and system weight of command data, thereby more accurately characterizing the likelihood of command data being invoked again by the coprocessor in the near future.

[0060] In one possible implementation, the command popularity value corresponding to the nth type of command data The following relationship must be satisfied: in, This represents the weight of the nth command data in the system; This represents the most recently executed value of the nth command data; This represents the execution interval value for the nth type of command data; This represents the compensation coefficient corresponding to the nth type of command data.

[0061] Optionally, compensation coefficient This is used to compensate for the discrepancy that arises when command heat values ​​are calculated solely based on the most recent running values ​​and running interval values. The compensation coefficient can be determined using a statistical re-averaging method.

[0062] In one possible implementation, assuming the number of command data items executed within a certain period is m, the execution count of each type of command data item is counted to obtain the execution count of the nth type of command data item. Then the compensation coefficient corresponding to the nth type of command data The following relationship must be satisfied: Optionally, based on the command popularity values ​​of various command data, the appropriate cache ratio for each type of command data in the external cache can be determined. For example, command data with higher popularity values ​​can correspond to a higher cache ratio, while command data with lower popularity values ​​can correspond to a lower cache ratio.

[0063] This embodiment uses the most recent execution value to reflect the location of the most recent execution of command data, and the execution interval value to reflect the interval between two adjacent executions of command data, so that the command popularity value takes into account both recentity and frequency. System weights are used to distinguish the importance of different command data in the system, avoiding the use of the same standard for caching all command data. A compensation coefficient, combined with the actual number of executions within the statistical period, corrects statistical biases arising from relying solely on the most recent execution value and execution interval value. Therefore, the command popularity value can more accurately indicate the likelihood of the corresponding command data being called again in the near future, providing a quantitative basis for subsequently determining cached objects in the external cache.

[0064] Optionally, the command popularity value mainly reflects the likelihood of a single command being invoked. However, a coprocessor typically uses multiple commands in a single processing step, and the popularity of a single command cannot directly reflect the actual effect of combining multiple commands into candidate commands. Candidate command combinations may change the actual output address of the message combination, and different relative distances between output addresses correspond to different degrees of influence. Based on this, this embodiment calculates the fitness value corresponding to the chromosome based on the frequency of occurrence of the message combination and the relative distance between output addresses, converting the influence of candidate command combinations on the message output results into a quantitative evaluation index that can be used for genetic optimization processing.

[0065] Because the combination of command and data can affect message output, for example, if the message output is at address 1... Next, the messages at address 2 should be output in sequence. However, due to the influence of candidate command combinations, the actual output was the message at address n. Then, the corresponding candidate command combination can be evaluated based on the address distance of the affected output message.

[0066] In one possible implementation, the fitness value corresponding to the z-th chromosome... The following relationship must be satisfied: in, This indicates the message combination corresponding to the candidate command combination; Indicates the number of message combinations included in the statistics; This indicates the number of times the s-th message combination occurs; Indicates the first The relative distance of the output address after a combination of messages is affected by a combination of candidate commands.

[0067] Optionally, since the relationship between candidate command combinations and message output results may be non-linear and discontinuous, relying solely on the command popularity value of a single command or a simple fixed sorting rule is insufficient to select suitable combinations for writing to the external buffer from multiple candidate command combinations. Therefore, this embodiment converts multiple candidate command combinations into an initial population, selects chromosomes with fitness values ​​that meet the requirements through selection processing, forms a new generation population through crossover and mutation processing, and recalculates fitness values ​​when preset termination conditions are not met, thus forming an iterative screening process for candidate command combinations.

[0068] Optionally, in Figure 1 On this basis, Figure 4 A flowchart illustrating another method for caching command data in a coprocessor provided in this application embodiment. See also... Figure 4 Step 105 includes: Step 105-1: Generate an initial population based on the combination of multiple candidate commands.

[0069] In this context, each chromosome in the initial population represents a candidate command combination.

[0070] Step 105-2: Based on the fitness values ​​of each chromosome, perform selection processing on the initial population to form the selected chromosomes.

[0071] Step 105-3: Perform crossover processing on the selected chromosomes to form crossover chromosomes.

[0072] Step 105-4: Perform mutation treatment on the crossed chromosomes to form a new generation of population.

[0073] Step 105-5: Determine whether the preset termination conditions are met based on the new generation population.

[0074] If the new generation of the population does not meet the preset termination condition, the process returns to step 104. If the preset termination condition is met, the process proceeds to steps 105-6.

[0075] Step 105-6: Output the optimized candidate command combination.

[0076] Optionally, Figure 4 The genetic optimization process shown may include: converting the offset relationship between command combinations and message data into a quantization function of chromosome fitness values; initializing the population; calculating the fitness value of each chromosome in the population; determining whether the termination condition is met; performing selection, crossover, and mutation when the termination condition is not met, and generating a new generation of population; and outputting the optimal solution when the termination condition is met.

[0077] This embodiment generates an initial population based on multiple candidate command combinations, ensuring that each chromosome can represent a candidate command combination. Selection is performed based on fitness values, allowing chromosomes with suitable fitness values ​​to proceed to subsequent processing. Crossover and mutation processes form a new generation population, enabling new combinations of candidate commands to emerge during iteration. When a preset termination condition is not met, the fitness value is recalculated, allowing the new generation population to continue receiving evaluation of message output results. Thus, multiple candidate command combinations can undergo multiple rounds of quantitative evaluation and combination adjustments, ultimately outputting an optimized candidate command combination.

[0078] Optionally, if a completely randomly generated chromosome is used directly as the initial population, the quality and coverage of candidate command combinations in the initial population may be unstable, thus affecting the starting state of subsequent selection, crossover, and mutation processes. Therefore, in this embodiment, after generating an initial chromosome of a preset population size, gene segments in the current chromosome are modified according to a preset flipping probability, and chromosomes are selected for retention based on a comparison of the fitness values ​​of the new chromosome and the current chromosome, as well as the current temperature. Through the above initialization iteration, the coverage of candidate command combinations can be expanded while chromosomes entering the initial population are screened.

[0079] Optionally, in Figure 4 On this basis, Figure 5 A flowchart illustrating another method for caching command data in a coprocessor provided in this application embodiment. See also... Figure 5 Step 105-1 includes: Step 105-1-1: Generate the initial chromosomes for the preset population size.

[0080] Step 105-1-2: According to the preset flipping probability, change at least one gene segment in the current chromosome to form a new chromosome.

[0081] Optionally, each gene segment in the current chromosome can be flipped and mutated according to a beta probability distribution model to obtain a new chromosome y.

[0082] In one possible implementation, the flipping probability of gene o The following relationship must be satisfied: in, Let be the flipping probability of gene o, and rd, a, and t are all random numbers following a uniform distribution.

[0083] Step 105-1-3: Based on the comparison of the fitness values ​​of the new chromosome and the current chromosome, and the current temperature, select the chromosome to retain.

[0084] The current temperature is used to regulate the range of different candidate command combinations retained in the initial population during the initialization phase. Furthermore, this current temperature decreases as the number of initialization iterations increases, so that the initial population covers more diverse candidate command combinations in the early stages of initialization and concentrates on candidate command combinations with higher fitness values ​​in the later stages.

[0085] Optionally, the fitness difference between the new chromosome and the current chromosome is calculated and denoted as . .when When the value is greater than 0, the new chromosome is retained. When the value is not greater than 0, the chromosome to be retained can be determined based on the current temperature. Specifically, when... When not greater than 0, The probability of retaining the old chromosome is , where T represents the current temperature.

[0086] Then determine whether the preset initialization termination condition is met.

[0087] Optionally, the preset initialization termination condition can be reaching the maximum number of initialization iterations. When the algorithm reaches the maximum number of initialization iterations, the resulting new chromosome y is used as the final chromosome of the initialization phase, completing the initialization process of the corresponding initial chromosome.

[0088] Step 105-1-4: If the preset initialization termination condition is not met, return to step 105-1-2.

[0089] Alternatively, the current temperature can be updated according to the following relationship: Where T' represents the updated current temperature; a represents the predetermined cooling rate.

[0090] This embodiment modifies gene segments in the current chromosome according to a preset flipping probability, enabling the initial chromosome to form new chromosomes corresponding to different candidate command combinations. By selecting chromosomes to retain based on fitness comparison results, candidate command combinations with better fitness values ​​are preserved. Introducing the current temperature adjusts the range of different candidate command combinations retained in the initial population, preventing the initialization phase from being entirely limited by a single fitness comparison result. Repeating the gene segment modification process when the preset initialization termination condition is not met allows the initial population to cover more candidate command combinations. Therefore, the diversity of candidate command combinations in the initial population can be improved, providing an initialized and screened population for subsequent genetic optimization processing.

[0091] Optionally, the initial population includes multiple chromosomes corresponding to different candidate command combinations. If chromosomes entering the next generation are selected completely randomly, candidate command combinations with higher fitness values ​​may not be effectively retained; if only the chromosome with the highest fitness value is selected, candidate command combinations in the population may be concentrated too early. Based on this, this embodiment sorts the chromosomes in the initial population according to their fitness values, and selects the chromosomes to enter the next generation based on the Zip distribution model and the sorting results, thus correlating the selection probability of a chromosome with its fitness ranking.

[0092] Optionally, in Figure 4 On this basis, Figure 6 A flowchart illustrating another method for caching command data in a coprocessor provided in this application embodiment. See also... Figure 6 Step 105-2 includes: Step 105-2-1: Sort the multiple chromosomes in the initial population according to the fitness values ​​of each chromosome.

[0093] Optionally, all chromosomes in the initial population can be sorted in ascending order according to their fitness values, and the ranking of the chromosomes can be determined based on their positions in the sorting results.

[0094] Step 105-2-2: Based on the Ziv distribution model and the sorting results of each chromosome, select the chromosomes to enter the next generation population.

[0095] In one possible implementation, the probability of chromosome Z being selected to enter the next generation population... Z The following relationship must be satisfied: Z in, This indicates the fitness ranking of chromosome Z; represents the fitness ranking of the i-th chromosome; skew represents the skewness of the Zif distribution model; k represents the number of chromosomes involved in the selection.

[0096] This embodiment sorts chromosomes in the initial population according to their fitness values, giving each chromosome a ranking corresponding to the evaluation results of candidate command combinations. By selecting chromosomes based on the Zip distribution model and the sorting results, chromosomes with different rankings have different selection probabilities. Thus, higher-ranked chromosomes have a higher chance of retention, while chromosomes with other rankings can still enter the next generation population according to their corresponding probabilities. This approach selects better candidate command combinations while avoiding complete concentration of candidate command combinations in the initial population.

[0097] Optionally, the selected chromosomes obtained from the selection process correspond to different candidate command combinations. If only the selected chromosomes are retained without combination adjustments, it is difficult to generate new candidate command combinations. Based on this, this embodiment selects the chromosome with the highest fitness value from the previous generation population as the control chromosome, selects the father chromosome based on the fitness value difference between the selected chromosome and the control chromosome, selects the mother chromosome based on the fitness value ratio between the selected chromosome and the control chromosome, and cross-references the gene segments of the father and mother chromosomes to form new crossover chromosomes.

[0098] Optionally, in Figure 4 On this basis, Figure 7 A flowchart illustrating another method for caching command data in a coprocessor provided in this application embodiment. See also... Figure 7 Step 105-3 includes: Step 105-3-1: Select the chromosome with the highest fitness value from the previous generation population as the control chromosome.

[0099] Optionally, the previous generation population is the population that performs the selection process in this instance; in addition, in the first iteration, the previous generation population is the initial population; in subsequent iterations, the previous generation population is the new generation population obtained from the previous iteration.

[0100] In one possible implementation, the fitness value of the control chromosome x' satisfies the following relationship: Step 105-3-2: Select the father chromosome based on the fitness value difference between each selected chromosome and the control chromosome.

[0101] Optionally, calculate the fitness difference between all chromosomes in the current population and the control chromosome. The fitness difference between chromosome Z and the control chromosome x' is denoted as... .

[0102] In one possible implementation, the probability that chromosome Z is chosen as the parent chromosome... The following relationship must be satisfied: in This represents the probability that chromosome Z will be selected as the parent chromosome; e is the natural base; and c is a constant greater than 0, used to limit the minimum threshold for selection as the parent chromosome.

[0103] Step 105-3-3: Select the maternal chromosome based on the fitness ratio between each selected chromosome and the control chromosome.

[0104] Optionally, the remaining chromosomes that were not selected as the parent chromosome are randomly arranged, and then the parent chromosome is selected based on the beta probability distribution model according to the arrangement order.

[0105] Optionally, the probability of chromosome Z being selected as the mother chromosome. The following relationship must be satisfied: in, This represents the fitness value of chromosome Z; This represents the fitness value of the control chromosome; a and t are both random numbers following a preset distribution. This represents the beta function.

[0106] Step 105-3-4: Cross over the gene segments of the father's chromosome and the mother's chromosome to form a crossover chromosome.

[0107] Optionally, after selecting the father and mother chromosomes, a crossover operation can be performed on gene segments in the father and mother chromosomes to obtain a new crossover chromosome.

[0108] This embodiment uses the chromosome with the highest fitness value as the control chromosome, providing a unified evaluation reference for the selection of both paternal and maternal chromosomes. By selecting the paternal chromosome based on the fitness value difference, it ensures that selected chromosomes with small differences from the control chromosome have a corresponding paternal selection criterion. Similarly, by selecting the maternal chromosome based on the fitness value ratio, it links the selection of the maternal chromosome to the fitness level of the selected chromosomes. Furthermore, by cross-linking gene segments from the paternal and maternal chromosomes, the crossed-over chromosome can simultaneously carry gene segments from both selected chromosomes. Thus, new candidate command combinations can be formed while retaining the characteristics of existing candidate command combinations.

[0109] Optionally, after selection and crossover, chromosomes in the population may gradually become similar, leading to a narrowing of the search range for candidate command combinations. To generate new combinations based on existing candidate command combinations, this embodiment determines the mutation probability based on the fitness ranking of the crossover chromosome in the current population, and modifies at least one gene segment in the crossover chromosome according to this mutation probability. By correlating the mutation probability with the fitness ranking, the gene segment modification process can be combined with the relative evaluation result of the crossover chromosome in the current population.

[0110] Optionally, in Figure 4 On this basis, Figure 8 A flowchart illustrating another method for caching command data in a coprocessor provided in this application embodiment. See also... Figure 8 Step 105-4 includes: Step 105-4-1: Determine the mutation probability corresponding to the crossing chromosome based on the fitness ranking of the crossing chromosome in the current population.

[0111] Alternatively, the Gombertz model can be used to represent the mutation probability of a single chromosome.

[0112] In one possible implementation, the mutation probability of chromosome Z The following relationship must be satisfied: in, This represents the fitness ranking of chromosome Z in the current population; a, b, and c are all random numbers that follow a uniform distribution.

[0113] Step 105-4-2: According to the mutation probability, change at least one gene segment in the crossing chromosome.

[0114] Optionally, based on the mutation probability determined in step 105-4-1, one or more gene segments in the crossing chromosome are altered to obtain the mutated chromosome.

[0115] This embodiment determines the mutation probability based on the fitness ranking of the crossing chromosome in the current population, giving crossing chromosomes with different rankings corresponding mutation possibilities. By altering at least one gene segment in the crossing chromosome according to the mutation probability, the candidate command combinations corresponding to the crossing chromosome can produce new variations. Therefore, the risk of excessive chromosome similarity in the population after selection and crossing over can be reduced, and new candidate command combinations can be introduced for the next generation of population.

[0116] Optionally, genetic optimization requires multiple rounds of iteration to screen candidate command combinations. However, without a clear termination condition, the algorithm may continue to perform selection, crossover, and mutation processes even when it can no longer generate better chromosomes, resulting in wasted processing resources. Therefore, this embodiment sets termination conditions such as a preset maximum number of iterations, the absence of better chromosomes in the new generation population, and the difference between the highest fitness values ​​of two adjacent generations being less than a preset threshold. These conditions are used to determine whether the genetic optimization process has reached a state where optimized candidate command combinations can be output.

[0117] Therefore, for step 105-5, when a preset termination condition occurs, it is determined that the new generation population meets the preset termination condition. The preset termination scenarios include at least one of the following: Genetic optimization processing reaches the preset maximum number of iterations; There are no chromosomes in the new generation with fitness values ​​superior to those in the previous generation. The difference in the highest fitness value between two adjacent generations of population is less than a preset threshold.

[0118] Optionally, the preset maximum number of iterations and preset threshold can be configured according to the scale of the genetic optimization process, the number of candidate command combinations, or the processing latency requirements. When there are no chromosomes in the new generation with a fitness value better than the previous generation, it indicates that the current iteration did not produce new offspring with better fitness values.

[0119] This embodiment limits the number of iterations in the genetic optimization process by setting a maximum number of iterations, thus preventing the process from continuing indefinitely. It terminates the process when no chromosome in the new generation has a fitness value superior to that of the previous generation, preventing further iterations when no better candidate command combinations can be generated. Furthermore, it terminates the process when the difference in the highest fitness value between two adjacent generations is less than a preset threshold, preventing the continued consumption of processing resources when fitness changes are already minimal. Therefore, a clear balance can be achieved between the optimization level of candidate command combinations and the execution overhead of the genetic optimization process.

[0120] Optionally, the genetic optimization process, after meeting the preset termination condition, yields optimization results in chromosome form, while the coprocessor's external buffer actually stores candidate command combinations. Therefore, it is necessary to convert chromosomes that meet the output conditions into command data that can be recognized and executed by the coprocessor. Based on this, in this embodiment, when the new generation population meets the preset termination condition, chromosomes whose fitness values ​​meet the preset output condition are selected, decoded into command identifier sequences, and optimized candidate command combinations are determined based on the command identifier sequences, thereby completing the conversion from genetic optimization results to actual command data.

[0121] Optionally, in Figure 4 On this basis, Figure 9 A flowchart illustrating another method for caching command data in a coprocessor provided in this application embodiment. See also... Figure 9 Steps 105-6 include: Step 105-6-1: When the new generation population meets the preset termination condition, select the chromosome whose fitness value meets the preset output condition.

[0122] Optionally, the preset output condition can be the highest fitness value, or the fitness value being within a preset output range. One or more chromosomes can be selected.

[0123] Step 105-6-2: Decode the selected chromosome into a command identifier sequence.

[0124] Optionally, gene segments in the chromosome correspond to command data in the candidate command combinations. After decoding the selected chromosome, a command identifier sequence composed of command identifiers can be obtained.

[0125] Step 105-6-3: Determine the optimized candidate command combination based on the command identifier sequence.

[0126] Optionally, based on each command identifier in the command identifier sequence, multiple corresponding command data are determined, and the determined multiple command data are used as an optimized candidate command combination.

[0127] This embodiment selects chromosomes whose fitness values ​​meet preset output conditions, ensuring the output chromosomes correspond to the evaluation results of genetic optimization. By decoding the selected chromosomes into command identifier sequences, gene segments within the chromosomes can be converted into specific command identifiers. Furthermore, by determining optimized candidate command combinations based on the command identifier sequences, the genetic optimization results can be transformed into multiple command data that the coprocessor can actually invoke. This avoids the optimization results remaining solely at the chromosome encoding level, allowing the optimized candidate command combinations to enter the subsequent cache priority determination process.

[0128] Optionally, the external cache has limited storage space and cannot store all optimized candidate command combinations simultaneously. Furthermore, the command data invocation behavior may change over different time periods, making it difficult for a fixed cache content to consistently match the actual needs of the coprocessor. Therefore, this embodiment determines the overall cache priority based on combination adaptation information and combination popularity information, and selects candidate command combinations to be written based on the available capacity of the external cache. After writing candidate command combinations within the current time period, the embodiment also adjusts the candidate command combinations in the external cache according to the updated overall cache priority when entering the next time period.

[0129] Optionally, in Figure 1 On this basis, Figure 10 A flowchart illustrating another method for caching command data in a coprocessor provided in this application embodiment. See also... Figure 10 Step 106 includes: Step 106-1: Determine the overall cache priority of each optimized candidate command combination based on the combination adaptation information and combination popularity information.

[0130] Among them, the combined adaptation information includes the adaptation value, which corresponds to the optimized candidate command combination; the combined popularity information includes multiple command popularity values; the command popularity values ​​correspond one-to-one with the command data; and the command data is the command data contained in the optimized candidate command combination.

[0131] Optionally, the combination adaptation information can reflect the impact of the optimized candidate command combination on the message combination output result, and the combination heat information can reflect the probability that each command data contained in the optimized candidate command combination will be called again by the coprocessor in the near future.

[0132] Step 106-2: Select the candidate command combination to be written based on the available capacity of the external cache.

[0133] Optionally, when the available capacity of the external cache is insufficient to store all optimized candidate command combinations, the candidate command combinations to be written are selected based on the overall cache priority. Candidate command combinations with higher overall cache priority are given priority.

[0134] Step 106-3: Within the current time period, combine the candidate commands to be written and write them to the external cache.

[0135] Optionally, within the current time period, the candidate command combinations to be written selected in step 106-2 are stored in an external cache. When the coprocessor needs the corresponding candidate command combinations, it can read them from the external cache.

[0136] Step 106-4: When entering the next time period, adjust the candidate command combination stored in the external cache according to the updated comprehensive cache priority.

[0137] Optionally, after entering the next time period, the fitness value and command popularity value of the candidate command combination can be retrieved again, the updated comprehensive cache priority can be determined based on the updated combination fitness information and combination popularity information, and the candidate command combination stored in the external cache can be adjusted based on the updated comprehensive cache priority.

[0138] Optionally, the above method can be used to combine the candidate command combinations selected by genetic optimization with the command popularity value to identify the candidate command combinations with high access frequency, and cache different candidate command combinations in an external cache based on different time periods.

[0139] This embodiment incorporates adaptive information to introduce the impact of optimized candidate command combinations on message output results, and incorporates combination popularity information to introduce the likelihood of each command data in the candidate command combination being called again in the near future. This allows the comprehensive cache priority to consider both the output performance at the command combination level and the call popularity at the individual command data level. By selecting candidate command combinations to be written based on the available capacity of the external cache, the limited cache space prioritizes storing combinations that meet the comprehensive cache priority conditions. By adjusting the external cache according to the updated comprehensive cache priority in the next time period, the cache content can keep up with changes in command data and candidate command combination call status. Therefore, the likelihood of the coprocessor hitting the required candidate command combination from the external cache can be increased, command data read latency caused by cache misses can be reduced, and the message data processing rate can be improved.

[0140] Alternatively, to better illustrate the above embodiments, a complete data flow example is provided below.

[0141] During system operation, command data is sequentially entered into the coprocessor. After each command data is acquired, the global cumulative value is updated. For target command data entering the coprocessor for the first time, the most recently executed value and execution interval value of the target command data are initialized; for target command data that has entered the coprocessor before, the execution interval value is calculated based on the current global cumulative value and the stored most recently executed value, and the most recently executed value is updated after the calculation is completed.

[0142] For each type of command data, a command heat value is calculated based on the most recent execution value, execution interval value, system weight, and compensation coefficient. The command heat value is used to characterize the likelihood that the corresponding command data will be invoked again by the coprocessor in the near future.

[0143] Based on the multiple command data used in a single processing run by the coprocessor, multiple candidate command combinations are constructed. For example, when the coprocessor's data processing bit width is 128 bits and the bit width of a single command data is 32 bits, multiple 32-bit command data in a single processing run can be constructed into a candidate command combination; when the bit width of a single command data is 64 bits, multiple 64-bit command data in a single processing run can be constructed into a candidate command combination. After encoding the candidate command combinations into chromosomes, each chromosome in the initial population represents one candidate command combination.

[0144] For each chromosome, the occurrence count and relative distance of the corresponding message combination are obtained, and the fitness value is calculated. Then, based on the fitness value, a selection process is performed on the initial population, a crossover process is performed on the selected chromosomes, and a mutation process is performed on the crossed-out chromosomes to generate a new generation population. If the new generation population does not meet the preset termination conditions, the fitness value is recalculated; if the new generation population meets the preset termination conditions, chromosomes with fitness values ​​that meet the preset output conditions are selected, the selected chromosomes are decoded into command identifier sequences, and optimized candidate command combinations are determined based on the command identifier sequences.

[0145] Finally, based on the fitness value corresponding to the optimized candidate command combination and the command heat value corresponding to each command data in the optimized candidate command combination, the overall cache priority is determined. Candidate command combinations to be written are selected based on the available capacity of the external cache and written to the external cache within the current time period. Upon entering the next time period, the candidate command combinations stored in the external cache are adjusted according to the updated overall cache priority.

[0146] Through the aforementioned data flow process, the runtime statistics of command data are first converted into command popularity values. Candidate command combinations are then encoded via chromosomes and undergo genetic optimization. The optimized candidate command combinations, combined with the command popularity values, participate in cache priority determination. Finally, candidate command combinations are written to or adjusted in the external cache according to different time periods. Therefore, the high-speed read capability of the external cache, combined with command popularity values, fitness values, and genetic optimization, helps to improve the rate at which command data enters the command processing module and increases the cache hit rate of the external cache.

[0147] To perform the steps and corresponding technical effects of the above examples, the following provides possible implementation methods for electronic devices. Specifically, Figure 11 This is a schematic diagram of the hardware architecture of an electronic device provided in an embodiment of this application. See also... Figure 11 The electronic device 30 includes a processor 300, a memory 301, a coprocessor 302, an external register 303, a communication interface 304, and a communication bus 305. The processor 300, memory 301, coprocessor 302, external register 303, and communication interface 304 are connected via the communication bus 305.

[0148] Communication interface 304 is used to receive message data to be processed and send the processed message data to the target network device. Coprocessor 302 is used to schedule, forward, and output the message data received by communication interface 304 according to command data. External buffer 303 is used to store optimized candidate command combinations. When coprocessor 302 needs to perform message processing operations, it can preferentially read the corresponding candidate command combination from external buffer 303.

[0149] The memory 301 is used to store data generated during the execution of computer programs and methods. When the computer program is executed by the processor 300, it implements the command data caching method described above in the coprocessor. The memory 301 can also store global cumulative values, recent running values, running interval values, command heat values, candidate command combinations, chromosomes, fitness values, genetic optimization results, and comprehensive cache priorities.

[0150] Optionally, the command acquisition module, statistics maintenance module, heat calculation module, combinatorial coding module, fitness calculation module, genetic optimization module, and cache control module can be program modules stored in memory 301. When the processor 300 executes the above program modules, it implements the corresponding command acquisition, statistics maintenance, heat calculation, combinatorial coding, fitness calculation, genetic optimization, and cache control functions, respectively.

[0151] In one possible implementation, processor 300 acquires command data entering coprocessor 302 via communication bus 305 and updates the corresponding global cumulative value, most recently run value, and run interval value. After calculating the command heat value, processor 300 constructs candidate command combinations and performs genetic optimization processing. After the cache control module determines the comprehensive cache priority, processor 300 writes candidate command combinations that meet the cache conditions into external cache 303 via communication bus 305.

[0152] After receiving a message processing task, the coprocessor 302 prioritizes accessing the external buffer 303. When the external buffer 303 stores candidate command combinations required by the coprocessor 302, the coprocessor 302 reads the candidate command combinations from the external buffer 303 and processes the message data according to the candidate command combinations. When entering the next time period, the processor 300 adjusts the candidate command combinations stored in the external buffer 303 according to the updated integrated cache priority.

[0153] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0154] For the electronic device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The electronic device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0155] The foregoing has described exemplary embodiments of this specification. It should be understood that in some cases, the modules described in this specification may be divided in a manner different from that in the embodiments, and the described actions or steps may be performed in a different order than that in the embodiments, while still achieving the desired result. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0156] The above are merely preferred embodiments of this specification and are not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification shall be included within the scope of protection of this specification.

Claims

1. A method for caching command data in a coprocessor, characterized in that, include: Obtain command data to enter the coprocessor; For each type of command data, maintain a global cumulative value, a recent run value, and a run interval value; wherein, the global cumulative value is used to represent the cumulative number of runs of the command data; the recent run value is used to represent the most recent run position of the command data; and the run interval value is used to represent the interval between two consecutive runs of the command data. Based on the recent running value, the running interval value, the weight corresponding to the command data, and the compensation coefficient, the command heat value corresponding to the command data is calculated; the command heat value is used to characterize the probability that the corresponding command data will be called again by the coprocessor in the near future; Based on the multiple command data used by the coprocessor in a single processing step, multiple candidate command combinations are constructed, and each candidate command combination is encoded into a corresponding chromosome. For each chromosome, the occurrence count of the message combination corresponding to the chromosome and the relative distance of the output address of the message combination are obtained, and the fitness value corresponding to the chromosome is calculated based on the occurrence count and the relative distance of the output address. Based on the fitness value corresponding to each chromosome, genetic optimization is performed on multiple chromosomes to obtain an optimized combination of candidate commands. Based on the optimized candidate command combinations and the command popularity value corresponding to the command data contained in the candidate command combinations, the cache priority of each candidate command combination is determined, and the candidate command combinations that meet the cache conditions are written into the external cache of the coprocessor according to different time periods.

2. The method for caching command data in a coprocessor according to claim 1, characterized in that, The step of maintaining a global cumulative value, a recent run value, and a run interval value for each type of command data includes: After obtaining each command data entry, update the global cumulative value; When the target command data first enters the coprocessor, the most recent running value and running interval value of the target command data are initialized; When the target command data enters the coprocessor for the first time, the running interval value is calculated based on the current global cumulative value and the stored most recent running value. After calculating the running interval value, update the most recent running value of the target command data.

3. The method for caching command data in a coprocessor according to claim 1, characterized in that, The step of performing genetic optimization on multiple chromosomes to obtain optimized candidate command combinations includes: An initial population is generated based on a plurality of the candidate command combinations; wherein each chromosome in the initial population represents a candidate command combination. Based on the fitness value corresponding to each chromosome, a selection process is performed on the initial population to form the selected chromosomes; The selected chromosomes are subjected to crossover processing to form crossover chromosomes; Mutation processing is performed on the crossed chromosomes to form a new generation of population; Based on the new generation population, determine whether it meets the preset termination conditions; When the new generation population does not meet the preset termination condition, return to the steps of obtaining the occurrence count of the message combination corresponding to the chromosome and the relative distance of the output address of the message combination for each chromosome, and calculating the fitness value corresponding to the chromosome based on the occurrence count and the relative distance of the output address; When the new generation of population meets the preset termination condition, the optimized candidate command combination is output.

4. The method for caching command data in a coprocessor according to claim 3, characterized in that, The step of generating an initial population based on a combination of multiple candidate commands includes: Generate an initial chromosome with a preset population size; According to a preset flipping probability, at least one gene segment in the current chromosome is changed to form a new chromosome; Based on the comparison of the fitness values ​​of the new chromosome and the current chromosome, and the current temperature, a chromosome is selected to be retained; wherein, the current temperature is used to adjust the range of different candidate command combinations retained in the initial population during the initialization phase. If the preset initialization termination condition is not met, return to the step of changing at least one gene segment in the current chromosome according to the preset flipping probability to form a new chromosome.

5. The method for caching command data in a coprocessor according to claim 3, characterized in that, The step of performing selection processing on the initial population according to the fitness value corresponding to each chromosome to form the selected chromosomes includes: The chromosomes in the initial population are sorted according to the fitness values ​​corresponding to each chromosome. Based on the Zif distribution model and the sorting results of each chromosome, the chromosomes selected to enter the next generation population are chosen.

6. The method for caching command data in a coprocessor according to claim 3, characterized in that, The step of performing crossover processing on the selected chromosomes to form crossover chromosomes includes: Select the chromosome with the highest fitness value from the previous generation population as the control chromosome; The parent chromosome is selected based on the fitness difference between each selected chromosome and the control chromosome. The maternal chromosome is selected based on the fitness ratio between each selected chromosome and the control chromosome. Gene segments from the father's chromosome and the mother's chromosome are crossed to form a crossover chromosome.

7. The method for caching command data in a coprocessor according to claim 3, characterized in that, The step of performing mutation processing on the crossed chromosomes to form a new generation population includes: The mutation probability corresponding to the crossover chromosome is determined based on the fitness ranking of the crossover chromosome in the current population. According to the mutation probability, at least one gene segment in the crossing chromosome is altered.

8. The method for caching command data in a coprocessor according to claim 3, characterized in that, The step of outputting the optimized candidate command combination when the new generation population meets the preset termination condition includes: When the new generation population meets the preset termination condition, select the chromosome whose fitness value meets the preset output condition; Decode the selected chromosome into a command identifier sequence; The optimized candidate command combination is determined based on the command identifier sequence.

9. The method for caching command data in a coprocessor according to claim 1, characterized in that, The step of determining the cache priority of each candidate command combination based on the optimized candidate command combination and the command popularity value corresponding to the command data contained in the candidate command combination, and writing the candidate command combinations that meet the cache conditions into the external cache of the coprocessor according to different time periods, includes: Based on the combination adaptation information and combination popularity information, the comprehensive cache priority of each optimized candidate command combination is determined; wherein, the combination adaptation information includes the adaptation value, which corresponds to the optimized candidate command combination; the combination popularity information includes multiple command popularity values; the command popularity value corresponds one-to-one with the command data; the command data is the command data contained in the optimized candidate command combination; Based on the available capacity of the external buffer, select the candidate command combination to be written; Within the current time period, the candidate commands to be written will be combined and written to the external cache; When entering the next time period, the candidate command combinations stored in the external cache are adjusted according to the updated comprehensive cache priority.

10. An electronic device, characterized in that, include: The communication interface is used to receive message data to be processed and to send processed message data. A coprocessor is used to schedule, forward, and output message data received by the communication interface according to command data; An external cache is used to store optimized combinations of candidate commands; Memory, used to store one or more programs; processor; A communication bus is used to connect the communication interface, the coprocessor, the external buffer, the memory, and the processor. When one or more programs are executed by the processor, the method for caching command data in the coprocessor as described in any one of claims 1 to 9 is implemented, and the optimized candidate command combination is written to the external cache so that the coprocessor reads the candidate command combination from the external cache.