Multiplexing Method for On-Chip Information and Related Devices
By introducing multiplexing methods of multi-stage buffers and polling arbitrators in the system on chip, the problem of debugging information and tracking information competing for bus bandwidth is solved, debugging and tracking efficiency is improved, and system complexity and cost are reduced.
Patent Information
- Application Number
- CN202411807068.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-12-10
AI Technical Summary
In existing systems on chip, debugging information and tracking information compete for bus bandwidth, resulting in transmission delay and waste of resources, and increasing independent buses improves information transmission independence and reliability but increases chip area and cost.
Multiplexing of on-chip information is achieved by introducing multi-stage buffers and polling arbitrators in the system-on-chip system. The specific steps include: the information of each module is stored in a primary FIFO buffer, transmitted to the secondary FIFO buffer through a polling arbitrator, and finally processed by an on-chip logic analyzer.
It effectively improves the debugging and tracking efficiency of the system on chip, reduces system complexity and cost, and improves system flexibility by dynamically adjusting the information transmission sequence.
Smart Images

Figure CN119271622B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular, to a method for multiplexing on-chip information and related devices. Background Art
[0002] A System on Chip (SoC) is an integrated circuit that integrates multiple functional modules, and these modules work together to achieve a complete system function. The system on chip integrates tasks that are usually completed by multiple independent modules onto the same circuit, thereby improving performance, reducing power consumption and cost, and reducing the overall system size.
[0003] In related technologies, it is necessary to manage and transmit debug information (Debug) and trace information (Trace) generated by each module in the system on chip. Debug information and trace information usually need to be transmitted at the same time, so they will compete for the bandwidth resources of the same bus. This competition may cause transmission delays for some information, affecting the real-time performance of debugging and tracing. Since the transmission requirements for debug and trace information may vary greatly at different times, the bus bandwidth is not fully utilized during some time periods, resulting in resource waste.
[0004] Therefore, there is an urgent need to design a technical solution to solve at least one of the above technical problems. Summary of the Invention
[0005] In view of the technical problems existing in the prior art, this application provides a method for multiplexing on-chip information and related devices, which realizes the multiplexing of on-chip information through multi-level buffers, polling arbitrators, and priority polling mechanisms in the system on chip, effectively improving the debugging and tracing efficiency of the system on chip, while reducing system complexity and cost.
[0006] In a first aspect, an embodiment of this application provides a method for multiplexing on-chip information. The method is applied to a system on chip including multiple modules, and the method includes:
[0007] Obtain on-chip information generated by each module in the system on chip, and store it in a first-in, first-out (FIFO) buffer of the first level that matches each module; the on-chip information includes at least: debug information and / or trace information;
[0008] According to a shared scheduling rule, transfer the on-chip information from their respective currently cached first-level FIFO buffers to a second-level FIFO buffer through a first-level polling arbiter; the second-level FIFO buffer is a common storage area shared by multiple modules;
[0009] Transfer the on-chip information from the second-level FIFO buffer to an on-chip logic analyzer through a second-level polling arbiter;
[0010] The on-chip information is processed by the on-chip logic analyzer to achieve multiplexing processing of the on-chip information generated by multiple modules.
[0011] In a second aspect, an embodiment of the present application provides a multiplexing device for on-chip information. The device is applied to an on-chip system including multiple modules, and the device at least includes the following units:
[0012] A first-in first-out (FIFO) buffer, configured to acquire and store the on-chip information generated by each module in the on-chip system; the on-chip information at least includes: debugging information and / or tracing information;
[0013] A first-level polling arbiter, configured to transfer the on-chip information from their respective currently cached first-level FIFO buffers to a second-level FIFO buffer according to a shared scheduling rule; the second-level FIFO buffer is a common storage area shared by multiple modules;
[0014] A second-level polling arbiter, configured to transfer the on-chip information from the second-level FIFO buffer to the on-chip logic analyzer according to a static scheduling rule;
[0015] The on-chip logic analyzer is configured to process the on-chip information to achieve multiplexing processing of the on-chip information generated by multiple modules.
[0016] In a third aspect, an embodiment of the present application provides an electronic device, and the electronic device includes:
[0017] At least one processor, a memory, and an input / output unit;
[0018] Wherein, the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the multiplexing method for on-chip information in the first aspect.
[0019] In a fourth aspect, a computer-readable storage medium is provided, which includes instructions. When the instructions are run on a computer, the computer is made to execute the multiplexing method for on-chip information in the first aspect.
[0020] The beneficial effects of this application are as follows: A method for multiplexing on-chip information and related devices are provided. In this technical solution, first, on-chip information generated by each module in the system-on-chip is obtained and stored in a first-in, first-out (FIFO) buffer that matches each module; the on-chip information includes at least: debugging information and / or tracing information. Then, according to the shared scheduling rule, the on-chip information is respectively transmitted from the current first-level FIFO buffer where it is temporarily stored to the second-level FIFO buffer through a first-level polling arbiter; the second-level FIFO buffer is a common storage area shared by multiple modules. Next, the on-chip information is transmitted from the second-level FIFO buffer to the on-chip logic analyzer through a second-level polling arbiter. Finally, the on-chip information is processed by the on-chip logic analyzer to achieve multiplexing processing of the on-chip information generated by multiple modules.
[0021] In the technical solution of this application, the multiplexing of on-chip information not only solves problems such as resource contention, increased chip area, increased system complexity, and performance bottlenecks, but can also be easily expanded by adding more modules and corresponding FIFO buffers without significantly increasing the design complexity. The polling arbiter can dynamically adjust the order of information transmission according to different requirements and priorities, improving the flexibility of the system. The design of the FIFO buffer and the polling arbiter can ensure the integrity and order of information, reducing data loss and errors. In the technical solution of this application, through the multi-level buffers and the polling arbiter in the system-on-chip for multiplexing on-chip information, the debugging and tracing efficiency of the system-on-chip is effectively improved, while reducing system complexity and cost. Description of the Drawings
[0022] Figure 1 is a schematic flowchart of a method for multiplexing on-chip information according to an embodiment of this application;
[0023] Figure 2 is a schematic diagram of the principle of a system-on-chip according to an embodiment of this application;
[0024] Figure 3 is a schematic structural diagram of a system-on-chip according to an embodiment of this application;
[0025] Figure 4 is a schematic structural diagram of a device for a method for multiplexing on-chip information according to an embodiment of this application;
[0026] Figure 5 is a schematic structural diagram of an electronic device according to an embodiment of this application;
[0027] Figure 6 is a schematic structural diagram of a media device according to an embodiment of this application. Detailed Embodiments
[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0029] A system-on-chip is an integrated circuit that integrates multiple functional modules, and these modules work together to achieve a complete system function. The system-on-chip integrates tasks that are usually completed by multiple independent modules onto the same circuit, thereby improving performance, reducing power consumption and cost, and reducing the overall size of the system.
[0030] In the related art, it is necessary to manage and transmit debug information and trace information generated by each module in the system-on-chip. Debug information and trace information usually need to be transmitted at the same time, so they will compete for the bandwidth resources of the same bus. This competition may cause transmission delays for some information, affecting the real-time performance of debugging and tracing. Since the transmission requirements for debug and trace information may vary greatly at different times, the bus bandwidth is not fully utilized during some time periods, resulting in resource waste.
[0031] In addition, the prior art usually uses two independent buses to transmit debug information and trace information respectively. Although this method avoids the competition between information to a certain extent, it increases the physical area of the chip, especially the wiring and connection parts of the bus. Although adding independent buses improves the independence and reliability of information transmission, it also increases the manufacturing cost and debugging complexity of the chip. Especially in high-density integrated circuits, the increase in area is particularly significant. And because debug and trace information are transmitted on different buses, it may cause delays in information merging and processing. Especially when the amount of information is large, this kind of delay may affect the overall performance of the system. When a multi-bus system processes information requests, the response time may increase due to bus arbitration and switching, affecting the real-time performance of the system.
[0032] Therefore, there is an urgent need to design a technical solution to solve at least one of the above technical problems.
[0033] To solve at least one of the technical problems in the related art, the embodiments of the present application provide a multiplexing method for on-chip information and related devices.
[0034] In the technical solution provided by this application, for the resource competition problem in the related art, the debugging information and tracing information generated by each module are first stored in their respective first-level FIFO buffers. In this way, each module has an independent buffer, avoiding direct competition for the bus bandwidth. Through the first-level polling arbiter and the second-level polling arbiter, the information transmission of each module can be managed and scheduled orderly. This arbitration mechanism ensures that the information of each module is transmitted at an appropriate time, avoiding delays and bottlenecks caused by resource competition.
[0035] Secondly, for the problem of increased chip area, multiple modules share a second-level FIFO buffer as a common storage area, rather than setting up independent buses and buffers for each module separately. This greatly reduces the chip area occupied, especially for high-density integrated circuits. Through the FIFO buffer and the polling arbiter, the interconnection structure between modules can be simplified, reducing the wiring complexity and further saving the chip area.
[0036] Thirdly, each module has an independent first-level FIFO buffer. This modular design makes the control logic of the system simpler and more intuitive. During debugging and testing, the FIFO buffer of each module can be verified separately, rather than the entire complex bus system. Through the shared polling arbiter, the information transmission of different modules can be managed and scheduled uniformly, reducing the complexity brought by multiple independent buses. This unified scheduling mechanism also facilitates the design and maintenance of the system.
[0037] Finally, for the problem of performance bottlenecks, the FIFO buffer can temporarily store information, relieve the instantaneous pressure on the bus bandwidth, and reduce the transmission delay. The polling arbiter ensures that the information is transmitted at an appropriate time, avoiding long waits. The shared second-level FIFO buffer and the polling arbiter can quickly respond to information requests, reducing the time for bus switching and arbitration. This makes the system more efficient and real-time when processing information requests. The on-chip logic analyzer can centrally process the on-chip information transmitted from the second-level FIFO buffer, optimize the data processing flow, and improve the overall system performance.
[0038] In the technical solution of this application, the multiplexing of on-chip information not only solves problems such as resource competition, increased chip area, increased system complexity, and performance bottlenecks, but can also be easily expanded by adding more modules and corresponding FIFO buffers without significantly increasing the design complexity. The polling arbiter can dynamically adjust the order of information transmission according to different requirements and priorities, improving the flexibility of the system. The design of the FIFO buffer and the polling arbiter can ensure the integrity and order of information, reducing data loss and errors. In the technical solution of this application, the multiplexing method of on-chip information effectively improves the debugging and tracing efficiency of the on-chip system while reducing the system complexity and cost by reasonably managing and scheduling information transmission.
[0039] The multiplexing scheme for on-chip information provided by the embodiments of the present application can also be executed by an electronic device, which can be a server, a server cluster, or a cloud server. The electronic device can also be a terminal device such as a mobile phone, a computer, a tablet computer, a wearable device, or a dedicated device (such as a dedicated terminal device with a multiplexing method system for on-chip information, etc.). The above-mentioned chips introduced in the above embodiments can also be installed in these electronic devices. Alternatively, these electronic devices can also install a service program for executing the multiplexing scheme for on-chip information.
[0040] Figure 1 It is a schematic flowchart of a multiplexing method for on-chip information provided by the embodiments of the present application, as Figure 1 shown, the method includes the following steps:
[0041] 101. Obtain the on-chip information generated by each module in the system-on-chip and store it in a first-in first-out (FIFO) buffer that matches each module;
[0042] 102. According to the shared scheduling rule, transfer the on-chip information from their respective currently cached first-level FIFO buffers to the second-level FIFO buffer through a first-level polling arbiter;
[0043] 103. Transfer the on-chip information from the second-level FIFO buffer to the on-chip logic analyzer through a second-level polling arbiter;
[0044] 104. Process the on-chip information through the on-chip logic analyzer to implement multiplexing processing of the on-chip information generated by multiple modules.
[0045] In the embodiments of the present application, the on-chip information at least includes: debug information (Debug) and / or trace information (Trace). Debug information (Debug) mainly focuses on the current specific status information and is used to help engineers diagnose and solve problems in the system during development and maintenance. Register content: Displays the current value of the register, helping engineers understand the internal state of the CPU or other modules. Signal status: Reflects the status of each signal line in the system, including input / output signals, interrupt signals, etc., helping engineers analyze the timing and status of the signals. Memory content: Displays the data in the current memory, including variables, data structures, etc., helping engineers understand the usage of the system memory. Error code: Records the errors or exceptions that occur in the system, helping engineers quickly locate problems. Current instruction address: Displays the address of the currently executed instruction, helping engineers track the execution flow of the program. Module status: Records the current status of the module, such as running status, paused status, reset status, etc.
[0046] Trace information mainly captures sequential event information, which is used to record the detailed execution process and state changes of the system over a period of time. For example, instruction pipeline: records the execution order of CPU instructions, including stages such as instruction fetching, decoding, execution, and write-back for each instruction. Data cache operations: records cache hits and misses of data caches, including operations such as cache reads, writes, and replacements. System calls: records the sequence of system calls, including the functions called, parameters, return values, etc. Task scheduling: records the scheduling of tasks, including events such as task creation, destruction, and switching. Interrupt handling: records the triggering and handling process of interrupts, including interrupt types, handling times, etc. Data flow: records the transmission path and operations of data in the system, including events such as data reads, writes, and transmissions.
[0047] In the related art, since multiple on-chip modules may generate debug and trace information simultaneously, how to efficiently manage and schedule this information to avoid transmission delays and resource waste is a challenge.
[0048] In the embodiments of the present application, each module is equipped with a FIFO (First In First Out) buffer for temporarily storing its trace information. Here, a first-level FIFO buffer is separately configured for each module to temporarily store the trace and debug information generated by the module. This can ensure that there is a temporary storage area for information before transmission, avoiding direct competition for the bus bandwidth.
[0049] As an optional embodiment, the first-level FIFO buffers matched with each module are FIFO buffers independently set on the side of each module. Among them, the first first-level FIFO buffers matched with each module are respectively used to store the trace information of each module, and the first first-level FIFO buffers matched with each module are connected to the second-level FIFO buffer. In the embodiments of the present application, the second-level FIFO buffer is a common storage area shared by multiple modules.
[0050] Further optionally, a second first-level FIFO buffer is also provided in the system-on-chip, and the second first-level FIFO buffer is used to store the debug information generated by each module. Further, the second-level FIFO buffer and the second first-level FIFO buffer in the system-on-chip are connected to the on-chip logic analyzer through the second-level polling arbiter.
[0051] In practical applications, assuming that the system-on-chip includes n modules, and each module generates different debug information, a first-level FIFO buffer is separately configured for each module. For the specific structure of the system-on-chip, please refer to Figure 2 As shown, assuming that a first-level FIFO buffer is separately configured for each module Figure 2The level-1 caches in it, namely level-1 cache 1, level-1 cache 2, ……, level-1 cache n. Optionally, a second level-1 FIFO buffer can also be set up to receive the debugging information of n modules. Here, the second level-1 FIFO buffer is denoted as level-1 cache n.
[0052] Exemplarily, the circuit structure of the system-on-chip can refer to Figure 3 the circuit schematic diagram shown.
[0053] In the embodiments of the present application, the polling arbiter is used to extract trace information from each FIFO according to a predetermined scheduling policy. Further optionally, the Weighted Round Robin (WRR) policy is adopted. WRR is a scheduling algorithm that can be used in the present application for the fair allocation of shared resources, such as on-chip network traffic management and CPU time slice allocation. Compared with the simple Round Robin scheduling, WRR introduces the concept of weight, enabling different tasks or data streams to obtain different resource allocation ratios according to their importance or requirements.
[0054] Specifically, in the embodiments of the present application, each piece of on-chip information generated by a module is assigned a weight, which reflects the share of resources it can obtain in one cycle. The larger the weight, the more resources are allocated. The scheduler allocates resources to each task in a cyclic order, but the allocated quantity is different according to their respective weights. If the weight of one task is twice that of another task, then the resources it obtains in the same round will also be twice that of the latter.
[0055] In each polling scheduling cycle, the available resources are accumulated to each task according to the weight until it reaches the corresponding weight value, and then resources are continued to be allocated to the next task. If a task cannot be completely processed within its allocated time slice, it will continue to be processed in the next cycle. This can ensure that all tasks can ultimately obtain the processing opportunity. The selected trace information will be directly written into a common FIFO for further processing. The secondary FIFO buffer serves as a common storage area shared by multiple modules and is used to temporarily store the information transmitted from the primary FIFO buffer.
[0056] In the system-on-chip, there are a debug bus and a trace bus. The debug bus and the trace bus will not be enabled simultaneously to avoid race conditions. When the trace information (Trace) is selected, the trace bus is enabled; when the debug information (Debug) is selected, the debug bus is enabled. The user can choose to focus on the debug information or the trace information according to the need. When the debug information is selected, the priority of the debug bus is always higher than that of the trace bus; when the trace information is selected, the priority of the trace bus is always higher than that of the debug bus.
[0057] In the last stage, through an additional polling arbiter, debug or trace information is selected to enter the OCLA. An additional polling arbiter is used to select the information transmitted on the debug bus or the trace bus to enter the OCLA. The fixed-priority scheduling policy ensures that the debug bus and the trace bus are not enabled simultaneously at the same time. Specifically, when the user selects the trace bus, the priority of the trace bus is always higher than that of the debug bus. Conversely, when the user selects the debug bus, the priority of the debug bus is always higher than that of the trace bus.
[0058] In 101, the module generates debug or trace information and temporarily stores it in its respective FIFO buffer. Each module generates debug information or trace information during its operation and writes this information into its corresponding first-level FIFO buffer.
[0059] In 102, the polling arbiter alternately extracts trace information from each FIFO according to the set priority. The first-level polling arbiter extracts information from the first-level FIFO buffers of each module in turn according to the weighted round-robin scheduling policy and writes the extracted information into the second-level FIFO buffer.
[0060] In 103, the extracted information is multiplexed to select the debug bus or the trace bus for transmission. By using the union data structure, different data types are allowed to share the same memory location. This means that debug information and trace information can share the same transmission path. According to the user's selection, the corresponding debug bus or trace bus is enabled. When the trace bus is enabled, the second-level polling arbiter transfers the information from the second-level FIFO buffer to the trace bus. When the debug bus is enabled, the information is transferred to the debug bus.
[0061] In practical applications, a union is a data structure that allows different data types to share the same memory location. Although it can only store one value at a time, it can flexibly switch between different data types. During the multiplexing process, the union data structure can be used for the information transmission of the debug bus and the trace bus. By sharing the same memory space, the two buses can use the same path, thus saving chip area and simplifying the design.
[0062] Finally, in 104, the On-Chip Logic Analyzer (OCLA) is used to monitor and store the information transmitted from the debug bus or the trace bus for subsequent analysis and debugging. The selected information is monitored and stored by the OCLA. Debug or trace information is selected to enter the OCLA through an additional polling arbiter. Here, the OCLA monitors and stores this information, providing a platform for centralized processing and analysis. The user can view and analyze the debug information or trace information through the OCLA for system debugging and optimization.
[0063] Combining the above steps, each module has an independent buffer, which avoids direct competition for the bus bandwidth. It ensures the orderly transmission of information, prevents the information of some modules from being ignored for a long time, and thus improves the overall response speed of the system. Multiple modules share a secondary FIFO buffer, reducing the additional chip area occupied. Through the multi-level FIFO buffer and polling arbiter, the interconnection structure between modules is simplified, and the wiring complexity is reduced. Each module has an independent primary FIFO buffer, making the system design more modular and the control logic more concise. The primary and secondary polling arbiters provide a unified scheduling mechanism, simplifying the system design and maintenance. The primary and secondary FIFO buffers can temporarily store information, relieve the instantaneous pressure on the bus bandwidth, and reduce the transmission delay. The polling arbiter can quickly respond to information requests, reduce the time for bus switching and arbitration, and improve the real-time performance of the system. OCLA centrally processes the information transferred from the secondary FIFO buffer, optimizing the data processing flow and improving the performance of the overall system.
[0064] Through the above multi-level buffering and arbitration mechanism, the on-chip information multiplexing method not only effectively manages the debug and trace information, but also ensures the efficient transmission and processing of information, solving problems such as resource competition, increased chip area, increased system complexity, and performance bottlenecks.
[0065] As an alternative embodiment, in 102, according to the shared scheduling rule, the on-chip information can be respectively transferred from the respective currently cached primary FIFO buffers of each module to the secondary FIFO buffer through the primary polling arbiter, which can be implemented as the following steps:
[0066] 1021. Obtain the task information of each module; the task information at least includes: the task scheduling information of each module, and the data processing task information currently executed by each module;
[0067] 1022. Based on the task information, configure the corresponding task weights for each module in each polling cycle;
[0068] 1023. According to the task weights, read the corresponding on-chip information from the primary FIFO buffers of each module through the primary polling arbiter and store it in the secondary FIFO buffer.
[0069] For example, by introducing a weight configuration and a scheduling mechanism based on task information in step 102, the performance and efficiency of the on-chip information multiplexing system can be significantly improved. Suppose there are four modules (A, B, C, D) in an on-chip system, and each module generates debug and trace information. The data processing task executed by module A is very important and needs to be monitored frequently, while the task of module B is relatively less important, and the tasks of modules C and D are of medium importance.
[0070] The task scheduling information of module A shows that its task frequency is high, and the current data processing task is very critical. The task scheduling information of module B shows that its task frequency is low, and the current task is relatively less important. The task scheduling information of modules C and D shows that their task frequencies are medium, and the current tasks are also of medium importance. Based on the above task information, the weight of module A is configured as 4, the weight of module B is configured as 1, and the weights of modules C and D are configured as 2. The information generated by module A will be preferentially extracted by the first-level polling arbiter to ensure that its important information can be transmitted to the second-level FIFO buffer in a timely manner. The information of module B will be extracted when there is no high-priority information to be transmitted due to its low weight, thus reducing resource waste. The information of modules C and D will be extracted at an appropriate time according to their medium weights to ensure fair utilization of resources.
[0071] In an on-chip system for real-time processing tasks, the debug information generated by module A needs to be processed within a very short time to ensure the real-time performance of the system. The information generated by modules B and C can tolerate a certain delay. The task scheduling information and data processing task information of module A show that it has extremely high requirements for real-time performance. The task scheduling information and data processing task information of modules B and C show that they have lower requirements for real-time performance. A higher weight is configured for module A to ensure the priority transmission of its information. Lower weights are configured for modules B and C to reduce resource competition.
[0072] In this way, the debug information of module A will be preferentially extracted and transmitted, reducing the delay caused by resource competition and ensuring the real-time performance of the system. Although there will be some delay in the information of modules B and C, it will not significantly affect the overall performance of the system.
[0073] In an on-chip system for complex calculations, the amount of data processing tasks generated by module A is large, while the amount of tasks generated by module B is small. Module A generates a large amount of debug and trace information and needs to be transmitted frequently. Module B generates a small amount of debug and trace information with a low transmission requirement. A higher weight is configured for module A to support its large information transmission requirement. A lower weight is configured for module B to reduce unnecessary resource occupation. The large amount of information of module A can be transmitted to the second-level FIFO buffer in a timely manner, optimizing the data processing flow. The small amount of information of module B will not occupy too many resources, improving the overall resource utilization rate of the system.
[0074] In a system - on - chip for handling mission - critical processing, the tasks processed by Module A have an important impact on the system reliability, while the tasks of Module B have lower requirements for reliability. The task scheduling information and data - processing task information of Module A show that its tasks have an important impact on the system reliability. The task scheduling information and data - processing task information of Module B show that its requirements for reliability are lower. A higher weight is configured for Module A to ensure that its information can be processed and transmitted preferentially. A lower weight is configured for Module B to reduce the over - attention to its information. The critical information generated by Module A can be processed and stored by OCLA in a timely manner, improving the system reliability. Although the information of Module B is transmitted slower, it does not affect the overall reliability of the system.
[0075] In a system - on - chip for handling multi - module parallel processing, the task information of each module is different. Through weight configuration and a scheduling mechanism based on task information, the system design can be simplified. The task information generated by Module A needs to be monitored and processed frequently. The task information generated by Module B needs to be concerned only under specific circumstances. The task information generated by Modules C and D needs to be checked regularly. A higher weight is configured for Module A to ensure the high - frequency transmission of its information. A lower weight is configured for Module B to transmit its information only under specific circumstances. Medium weights are configured for Modules C and D to transmit their information regularly. Through weight configuration, the system design and maintenance are simplified, and complex control logic is reduced. The system can dynamically adjust resource allocation according to the importance and requirements of tasks, improving the flexibility and scalability of the system.
[0076] By introducing a task weight configuration and weighted round - robin scheduling mechanism based on task information in Step 102, the resource utilization efficiency of the on - chip information multiplexing system can be significantly improved, transmission delay can be reduced, data - processing flow can be optimized, system reliability can be enhanced, and system design can be simplified. This mechanism ensures the orderly and efficient transmission of information, making the system - on - chip more flexible and reliable when processing debugging and tracing information generated by multiple modules.
[0077] As an optional embodiment, in 1023, according to the task weights, reading the corresponding on - chip information from the first - level FIFO buffers of each module by a first - level round - robin arbiter and storing it in the second - level FIFO buffer can be achieved as follows:
[0078] First, according to the task weights in each polling cycle, determine the resource configuration ratio configured for each module in each polling cycle; the resource configuration ratio is the ratio between the available resource amount corresponding to each module in each polling cycle and the total available resource amount. Then, according to the resource configuration ratio in each polling cycle, the first - level round - robin arbiter reads the debugging information to be extracted from each module from the first - level FIFO buffer to the second - level FIFO buffer.
[0079] Exemplarily, in step 1023, by determining the resource allocation ratio in each polling cycle according to the task weights, and reading the debug information from the first-level FIFO buffers of each module according to these ratios and storing it in the second-level FIFO buffer, more efficient and orderly resource management can be achieved. Suppose there are three modules (A, B, C) in a system, and each module generates debug information and trace information. Module A is used to process high-frequency tasks, generates a large amount of debug information, has a high requirement for real-time performance, and has a weight of 4. Module B is used to process low-frequency tasks, generates a small amount of debug information, has a low requirement for real-time performance, and has a weight of 1. Module C is used to process medium-frequency tasks, generates a medium amount of debug information, has a medium requirement for real-time performance, and has a weight of 2.
[0080] Further assume that the total available resources in each polling cycle are 100 units.
[0081] Based on this, the first-level polling arbiter reads the debug information from the first-level FIFO buffers of each module according to the above resource allocation ratios and stores it in the second-level FIFO buffer. Module A generates a large amount of debug information and has a high requirement for real-time performance. By allocating a higher resource ratio (57%), it ensures that its information can be preferentially extracted and transmitted, reducing the delay caused by resource competition. At the same time, Module B generates less debug information and has a low requirement for real-time performance. By allocating a lower resource ratio (14%), it avoids waste of resources. The resource allocation is more reasonable, improving the resource utilization efficiency and ensuring the timely processing of high-frequency tasks. The debug information generated by Module A needs to be processed within a very short time to ensure the real-time performance of the system. With a higher resource ratio (57%), the information of Module A can be preferentially extracted and transmitted in each polling cycle, reducing the transmission delay. The information transmission delay of critical tasks is significantly reduced, improving the real-time response ability of the system. Module C generates a medium amount of debug information and has a medium requirement for real-time performance. By allocating a medium resource ratio (29%), it ensures that its information can be extracted and transmitted at an appropriate time, meeting its needs without overusing resources. The data processing flow is more optimized, achieving dynamic balance of resources and improving the overall performance of the system. The tasks processed by Module A have an important impact on the system reliability. With a higher resource ratio (57%), it ensures that the critical debug information generated by it can be processed and transmitted in time, reducing the risk of system failures caused by information delay. The system reliability is significantly improved, and the information of critical tasks can be monitored and processed in time.
[0082] Through weight configuration and resource allocation ratio, the system-on-chip can dynamically adjust resource allocation, avoiding complex fixed resource allocation schemes. For example, Module B generates less debugging information and has lower real-time requirements, so a lower resource allocation ratio (14%) is assigned, simplifying the system control logic. System design and maintenance are more straightforward, reducing the complexity of the control logic, improving the scalability and flexibility of the system.
[0083] By determining the resource allocation ratio in each polling cycle based on the task weights in step 1023 and reading and transmitting the debugging information according to these ratios through a first-level polling arbiter, reasonable allocation and efficient management of resources can be achieved. This mechanism not only optimizes the information transmission and processing flow, improves resource utilization efficiency and real-time response capabilities, but also simplifies system design, enhances the overall performance and reliability of the system, making the system-on-chip more efficient and flexible in processing debugging information generated by multiple modules.
[0084] Based on the system-on-chip structure described in the above embodiments, in 103, by transmitting the on-chip information from the secondary FIFO buffer to the on-chip logic analyzer through a secondary polling arbiter, it can be achieved as follows:
[0085] 1031, determine the target on-chip information to be selected in the current polling cycle through the secondary polling arbiter;
[0086] 1032, based on the information type to which the target on-chip information belongs, activate the matching debug bus or trace bus;
[0087] 1033, transmit the target on-chip information from the secondary FIFO buffer to the on-chip logic analyzer through the debug bus or trace bus.
[0088] In the embodiments of the present application, the bus types at least include: a debug bus and a trace bus. In practical applications, the debug bus and the trace bus cannot be enabled simultaneously, otherwise a race condition will occur, resulting in transmission conflicts of on-chip information. In a system-on-chip, the debug bus and the trace bus are two different communication buses used to transmit different debugging and tracing information. The design purpose of these two buses is to optimize the transmission and processing of on-chip information, avoid competition conditions and transmission conflicts, and improve the debugging and tracing capabilities of the system.
[0089] The debug bus is mainly used to transmit data related to system debugging, such as register status, memory content, hardware breakpoints, etc. The data transmitted by the debug bus is usually low-frequency, high-priority information that needs to be accessed and processed in a timely manner during the debugging process. It helps developers and test engineers locate and solve system problems and ensures the normal operation of the system. It can adopt serial or parallel transmission methods, supporting point-to-point or broadcast transmission. Usually, the bandwidth requirement is low, but low latency and high reliability need to be ensured.
[0090] The trace bus is mainly used to transmit data related to system operating status and event tracing, such as instruction execution sequences, data streams, event logs, etc. The data transmitted by the trace bus is usually high-frequency and low-priority information, which is used to record the operating trajectory and events of the system. It helps developers and test engineers analyze the operating status of the system, optimize performance and functions. It adopts a high-speed serial transmission method, supporting high bandwidth and large data transmission. The bandwidth requirement is relatively high, which is used to transmit a large amount of trace information generated at high speed.
[0091] In the above system-on-chip structure, the specific steps of transmitting on-chip information from the secondary FIFO buffer to the on-chip logic analyzer through the secondary polling arbiter are as follows:
[0092] The secondary polling arbiter reads the on-chip information to be transmitted from the secondary FIFO buffer. According to the predetermined polling rules and priorities, it determines the target on-chip information to be transmitted in the current polling cycle. For example, the target information can be selected according to the weight ratio of each module or the urgency of the information. For instance, assume that the secondary FIFO buffer stores the debug information and trace information of modules A, B, and C. If the weight ratio (module A: 57%, module B: 14%, module C: 29%) is used for selection, and if module A has more information and a higher weight, its information will be preferentially selected.
[0093] The secondary polling arbiter identifies the type of the target on-chip information and determines whether it is debug information or trace information. According to the information type, it activates the corresponding debug bus or trace bus. Debug information uses the debug bus, and trace information uses the trace bus. For example, if the target on-chip information is the debug information generated by module A, the secondary polling arbiter will activate the debug bus. If the target on-chip information is the trace information generated by module C, the secondary polling arbiter will activate the trace bus.
[0094] The secondary polling arbiter transmits the target on-chip information from the secondary FIFO buffer to the on-chip logic analyzer through the activated debug bus or trace bus. In the same polling cycle, the debug bus and the trace bus cannot be enabled simultaneously to avoid race conditions and transmission conflicts.
[0095] For example, assume that the target on-chip information selected in the current polling cycle is the debug information of module A. The secondary polling arbiter activates the debug bus and transmits this debug information to the on-chip logic analyzer. In the next polling cycle, assume that the target on-chip information selected is the trace information of module C. The secondary polling arbiter activates the trace bus and transmits this trace information to the on-chip logic analyzer.
[0096] In this way, through the dynamic selection mechanism of the secondary polling arbiter, the reasonable allocation of resources can be ensured. The information of high-frequency tasks is preferentially transmitted, while the information of low-frequency tasks is transmitted on demand, avoiding resource waste. By preferentially selecting high-priority information types (such as debug information), the transmission delay can be significantly reduced, ensuring the real-time processing of critical information. The clear division of labor and mutually exclusive enabling mechanism of the debug bus and the trace bus avoid competition and conflicts in information transmission, making the data processing flow more optimized and orderly. By ensuring the mutually exclusive enabling of the debug bus and the trace bus, the risk of transmission conflicts is reduced, and the reliability and stability of the system are improved. The dynamic selection mechanism of the secondary polling arbiter and the clear division of labor of the bus simplify system design and maintenance, reduce the complexity of control logic, and improve the flexibility and scalability of the system.
[0097] By introducing the dynamic selection mechanism of the secondary polling arbiter and the mutually exclusive enabling mechanism of the debug bus and the trace bus in step 103, the on-chip information can be effectively managed and transmitted, improving the resource utilization efficiency of the system, reducing the transmission delay, optimizing the data processing flow, enhancing the system reliability, and simplifying the system design. This mechanism ensures the orderly and efficient transmission of information, making the on-chip system more flexible and reliable when processing debug and trace information generated by multiple modules.
[0098] Further optionally, in 1031, based on the task information of each module, the target on-chip information to be executed is selected through the secondary polling arbiter. Alternatively, in 1031, the secondary polling arbiter can also receive a user instruction and determine the target on-chip information to be executed based on the user instruction.
[0099] For example, the user can actively select the information that needs to be preferentially transmitted. For example, in a critical debug phase, the user can instruct to preferentially transmit the debug information of module A to ensure the efficient progress of the debug process. The user's control ability over the system is enhanced, and the information transmission strategy can be flexibly adjusted according to actual needs. In different debug scenarios, the user can dynamically adjust the type of information to be preferentially transmitted through instructions. For example, in a performance optimization scenario, the user can instruct to preferentially transmit the trace information of module C to help analyze the system performance. The system can better adapt to various debug and optimization scenarios, improving the flexibility and applicability of the system.
[0100] For example, the user can specify the modules that need to be focused on. For example, during the debugging process, the user can instruct to preferentially transmit the debugging information of Module A to help quickly locate problems. The debugging efficiency is significantly improved, and the user can obtain the required information more quickly, accelerating the problem-solving speed. The user can instruct not to transmit the information of Module B temporarily to avoid unnecessary resource occupation. Unnecessary information transmission is reduced, further improving the resource utilization efficiency and the real-time performance of the system. By dynamically adjusting the information transmission strategy through user instructions, the system maintainer can flexibly adjust according to the actual situation and quickly respond to different debugging requirements. System maintenance is more convenient, the response speed is faster, and the maintainability of the system is improved.
[0101] In the embodiments of the present application, the multiplexing of on-chip information not only solves problems such as resource contention, increased chip area, increased system complexity, and performance bottlenecks, but can also be easily expanded by adding more modules and corresponding FIFO buffers without significantly increasing the design complexity. The polling arbiter can dynamically adjust the order of information transmission according to different requirements and priorities, improving the flexibility of the system. The designs of the FIFO buffer and the polling arbiter can ensure the integrity and order of information, reducing data loss and errors. The technical solution of the present application, the multiplexing method of on-chip information effectively improves the debugging and tracing efficiency of the on-chip system by reasonably managing and scheduling information transmission, while reducing system complexity and cost.
[0102] In another embodiment of the present application, a multiplexing device for on-chip information is further provided, and the device is applied to an on-chip system including multiple modules. Refer to Figure 4 as described, the device includes the following units:
[0103] A first-in first-out (FIFO) buffer, configured to acquire and store on-chip information generated by each module in the on-chip system; the on-chip information includes at least: debugging information and / or tracing information;
[0104] A first-level polling arbiter, configured to transmit the on-chip information from their respective currently cached first-level FIFO buffers to a second-level FIFO buffer according to a shared scheduling rule; the second-level FIFO buffer is a common storage area shared by multiple modules;
[0105] A second-level polling arbiter, configured to transmit the on-chip information from the second-level FIFO buffer to an on-chip logic analyzer according to a static scheduling rule;
[0106] The on-chip logic analyzer is configured to process the on-chip information to implement multiplexing processing of on-chip information generated by multiple modules.
[0107] Further optionally, the first-level FIFO buffer is configured to transfer the on-chip information from their respective currently cached first-level FIFO buffers to the second-level FIFO buffer respectively through a first-level polling arbiter according to a shared scheduling rule, and is configured as follows:
[0108] Obtain the task information of each module; the task information includes at least: the task scheduling information of each module and the data processing task information currently executed by each module;
[0109] Configure corresponding task weights for each module in each polling cycle based on the task information;
[0110] According to the task weights, read the corresponding on-chip information from the first-level FIFO buffers of each module and store it in the second-level FIFO buffer.
[0111] Further optionally, the first-level FIFO buffer is configured to read the corresponding on-chip information from the first-level FIFO buffers of each module through a first-level polling arbiter according to the task weights and store it in the second-level FIFO buffer, and is configured as follows:
[0112] Determine the resource allocation ratio configured for each module in each polling cycle according to the task weights in each polling cycle; the resource allocation ratio is the ratio between the available resource amount corresponding to each module and the total available resource amount in each polling cycle;
[0113] Read the debug information to be extracted from each module from the first-level FIFO buffer to the second-level FIFO buffer according to the resource allocation ratio in each polling cycle.
[0114] Further optionally, the first-level FIFO buffer matched with each module is a FIFO buffer independently set on the side of each module;
[0115] Among them, the first first-level FIFO buffers matched with each module are respectively used to store the trace information of each module, and the first first-level FIFO buffers matched with each module are connected to the second-level FIFO buffer.
[0116] Further optionally, a second first-level FIFO buffer is also provided in the system-on-chip, and the second first-level FIFO buffer is used to store the debug information generated by each module.
[0117] Further optionally, the second-level FIFO buffer and the second first-level FIFO buffer in the system-on-chip are connected to the on-chip logic analyzer through the second-level polling arbiter.
[0118] Further optionally, a secondary polling arbiter is configured to transfer the on-chip information from the secondary FIFO buffer to an on-chip logic analyzer, specifically configured as follows:
[0119] Determine the target on-chip information to be selected in the current polling cycle;
[0120] Based on the information type to which the target on-chip information belongs, activate a matching debug bus or trace bus; wherein the bus types at least include: a debug bus and a trace bus;
[0121] Transfer the target on-chip information from the secondary FIFO buffer to the on-chip logic analyzer through the debug bus or the trace bus.
[0122] Further optionally, a secondary polling arbiter is configured to determine the target on-chip information to be selected in the current polling cycle, specifically configured as follows:
[0123] Based on the task information of each module, select the target on-chip information to be executed through the secondary polling arbiter; or receive a user instruction through the secondary polling arbiter and determine the target on-chip information to be executed based on the user instruction.
[0124] The system can implement various steps in the above method embodiments, which will not be elaborated here for the time being.
[0125] In the embodiments of the present application, a multiplexing device for on-chip information is adopted,....
[0126] Please refer to Figure 5 , Figure 5 which is a schematic diagram of an embodiment of an electronic device provided in the embodiments of the present application. As Figure 5 shown, the embodiments of the present application provide an electronic device 500, including a memory 510, a processor 520, and a computer program 511 stored in the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 511, the following steps are implemented: obtain the on-chip information generated by each module in the system-on-chip and store it in a first-in first-out (FIFO) buffer that matches each module; the on-chip information at least includes: debug information and / or trace information; according to a shared scheduling rule, transfer the on-chip information from their respective currently cached first-level FIFO buffers to a second-level FIFO buffer through a first-level polling arbiter; the second-level FIFO buffer is a common storage area shared by multiple modules; transfer the on-chip information from the second-level FIFO buffer to an on-chip logic analyzer through a second-level polling arbiter; process the on-chip information through the on-chip logic analyzer to implement multiplexing processing of the on-chip information generated by multiple modules.
[0127] Please refer to Figure 6 ,Figure 6 Schematic diagram of an embodiment of a computer-readable storage medium provided by an embodiment of the present application. As Figure 6 shown, this embodiment provides a computer-readable storage medium 600, on which a computer program 611 is stored. When the computer program 611 is executed by a processor, the following steps are implemented: obtaining on-chip information generated by each module in the system-on-chip and storing it in a first-in, first-out (FIFO) buffer that matches each module; the on-chip information at least includes: debugging information and / or tracing information; according to a shared scheduling rule, transmitting the on-chip information from their respective currently cached first-level FIFO buffers to a second-level FIFO buffer through a first-level polling arbiter; the second-level FIFO buffer is a common storage area shared by multiple modules; transmitting the on-chip information from the second-level FIFO buffer to an on-chip logic analyzer through a second-level polling arbiter; processing the on-chip information through the on-chip logic analyzer to implement multiplexing processing of the on-chip information generated by multiple modules.
[0128] It should be noted that in the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0129] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.
[0130] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0131] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
Claims
1. A method for multiplexing on-chip information, characterized in that: The method is applied to a system on chip including a plurality of modules, and the method comprises: The on-chip information generated by each module in the on-chip system is obtained and stored in a first-level first-out FIFO buffer matching each module; the on-chip information at least includes: debugging information and / or tracing information; the first-level FIFO buffer matching each module is a FIFO buffer independently set on each module side; the first-level FIFO buffer matching each module is used to store the tracing information of each module respectively; the on-chip system is also provided with a second-level FIFO buffer, and the second-level FIFO buffer is used to store the debugging information generated by each module; According to the shared scheduling rule, the on-chip information is respectively transferred from the currently temporarily stored first-level FIFO buffer to the second-level FIFO buffer through the first-level polling arbitrator; the second-level FIFO buffer is a common storage area shared by multiple modules; the first first-level FIFO buffer matched with each module is connected to the second-level FIFO buffer; the second-level FIFO buffer and the second first-level FIFO buffer in the on-chip system are connected to the on-chip logic analyzer through the second-level polling arbitrator; The method of transmitting the on-chip information from the respective currently temporarily stored first-level FIFO buffers to the second-level FIFO buffers through the first-level polling arbitrator according to the shared scheduling rule includes: Acquire task information of each module; the task information at least includes: task scheduling information of each module and data processing task information currently executed by each module; configure a corresponding task weight for each module in each polling cycle based on the task information; read corresponding on-chip information from the primary FIFO buffer of each module through a primary polling arbitrator according to the task weight, and store it in a secondary FIFO buffer; The on-chip information is transmitted from the secondary FIFO buffer to the on-chip logic analyzer through the secondary polling arbiter; wherein the on-chip information is transmitted from the secondary FIFO buffer to the on-chip logic analyzer through the secondary polling arbiter, comprising: Determine the target on-chip information to be selected in the current polling cycle through the secondary polling arbiter; start a matching debug bus or trace bus based on the information type of the target on-chip information; wherein the bus type includes at least: a debug bus and a trace bus; transmit the target on-chip information from the secondary FIFO buffer to the on-chip logic analyzer through the debug bus or the trace bus; The on-chip information is processed by the on-chip logic analyzer to implement multiplexing processing of on-chip information generated by multiple modules.
2. The on-chip information multiplexing method according to claim 1, characterized in that: The step of reading corresponding on-chip information from the primary FIFO buffer of each module through a primary polling arbitrator according to the task weight and storing the information in the secondary FIFO buffer comprises: According to the task weight in each polling cycle, determine the resource allocation ratio configured for each module in each polling cycle; the resource allocation ratio is the ratio between the available resource amount corresponding to each module and the total available resource amount in each polling cycle; According to the resource allocation ratio in each polling cycle, the debugging information to be extracted in each module is read from the first-level FIFO buffer to the second-level FIFO buffer through the first-level polling arbitrator.
3. The on-chip information multiplexing method according to claim 1, characterized in that: The step of determining the target on-chip information to be selected in the current polling cycle by the secondary polling arbiter includes: Based on the task information of each module, the target on-chip information to be executed is selected through the secondary polling arbitrator; or, The user instruction is received through the secondary polling arbiter, and the target on-chip information to be executed is determined based on the user instruction.
4. An on-chip information multiplexing device, characterized in that: The device is applied to a system on chip including a plurality of modules, and the device comprises the following units, wherein: A first-level first-out FIFO buffer is configured to obtain and store on-chip information generated by each module in the system-on-chip; the on-chip information at least includes: debugging information and / or tracing information; the first-level FIFO buffer matched with each module is a FIFO buffer independently set on the side of each module; the first-level FIFO buffer matched with each module is used to store the tracing information of each module respectively; the system-on-chip is also provided with a second-level FIFO buffer, and the second-level FIFO buffer is used to store the debugging information generated by each module; The first-level polling arbiter is configured to transfer the on-chip information from the first-level FIFO buffer currently temporarily stored to the second-level FIFO buffer according to the shared scheduling rule; the second-level FIFO buffer is a common storage area shared by multiple modules; the first first-level FIFO buffer matched with each module is connected to the second-level FIFO buffer; the second-level FIFO buffer and the second first-level FIFO buffer in the on-chip system are connected to the on-chip logic analyzer through the second-level polling arbiter; The first-level polling arbiter, according to the shared scheduling rule, transmits the on-chip information from the first-level FIFO buffer currently temporarily stored to the second-level FIFO buffer through the first-level polling arbiter, and is specifically configured as follows: Acquire task information of each module; the task information at least includes: task scheduling information of each module and data processing task information currently executed by each module; configure a corresponding task weight for each module in each polling cycle based on the task information; read corresponding on-chip information from the primary FIFO buffer of each module through a primary polling arbitrator according to the task weight, and store it in a secondary FIFO buffer; A secondary polling arbiter is configured to transfer the on-chip information from the secondary FIFO buffer to the on-chip logic analyzer according to a static scheduling rule; wherein the secondary polling arbiter, through the secondary polling arbiter, transfers the on-chip information from the secondary FIFO buffer to the on-chip logic analyzer, and is specifically configured to: determine the target on-chip information to be selected in the current polling cycle through the secondary polling arbiter; based on the information type to which the target on-chip information belongs, start a matching debug bus or trace bus; wherein the bus type includes at least: a debug bus and a trace bus; and transfer the target on-chip information from the secondary FIFO buffer to the on-chip logic analyzer through the debug bus or the trace bus; The on-chip logic analyzer is configured to process the on-chip information to implement multiplexing processing of the on-chip information generated by multiple modules.
5. An electronic device, characterized in that: include: Memory for storing computer software programs; A processor is used to read and execute the computer software program, thereby implementing the on-chip information multiplexing method as described in any one of claims 1-3.
6. A computer-readable storage medium, characterized in that: The storage medium stores a computer software program, and when the computer software program is executed by a processor, the on-chip information multiplexing method according to any one of claims 1 to 3 is implemented.
7. A chip, characterized in that: The chip is loaded with a computer software program and / or a hardware unit, and the computer software program and / or the hardware unit is used to implement the on-chip information multiplexing method as described in any one of claims 1-3.
Citation Information
Patent Citations
Multi-channel arbitration circuit of on-line simulation debugger and scheduling method thereof
CN109062661A