Computer-implemented method for determining a configuration for performing direct memory accesses in a computing unit

The proposed procedure optimizes DMA access configurations in computing units to prevent overloading and ensure timely completion of DMA operations, addressing the challenges of managing multiple DMA accesses in existing technologies.

DE102023210898A1Pending Publication Date: 2025-05-08ROBERT BOSCH GMBH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
DE102023210898
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing memory direct access units in computing units, such as system-on-a-chip (SOC) or microcontrollers, face challenges in efficiently managing a large number of DMA accesses without overloading the DMA units or communication systems, which can lead to delays and failure to meet time specifications.

Method used

A computer-implemented procedure is proposed to determine an optimized configuration for DMA access in computing units. This involves an optimization process that considers the specified access properties of individual DMA accesses and the available DMA units, using methods such as genetic algorithms to find the best DMA configuration that avoids overloading and meets temporal specifications.

Benefits of technology

The optimized DMA configuration ensures that DMA accesses are carried out efficiently, preventing overloading of DMA units and communication systems. This results in timely completion of DMA operations, adhering to specified deadlines and avoiding delays, thereby enhancing the reliability and performance of computing units, especially in safety-critical applications like vehicle control systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for determining a configuration for performing direct memory accesses in a computing unit (100), wherein a plurality of direct memory accesses are to be performed according to predetermined access characteristics, and wherein a plurality of direct memory access units (110, 111, 112, 113, 114) are configured to perform direct memory accesses, the method comprising the following steps: performing an optimization procedure and determining an access configuration as a result of the optimization procedure, wherein the determined access configuration defines which direct memory accesses of the plurality of direct memory access units are performed by which direct memory access units (110, 111, 112, 113, 114) of the plurality of direct memory access units.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to a computer-implemented method for determining a configuration for performing direct memory accesses in a computing unit, in particular a system-on-a-chip or a microcontroller, in particular of a vehicle, as well as a configuration computing unit and a computer program for carrying out the same, as well as a computing unit, in particular a microcontroller, which is configured according to such a method. Background of the invention

[0002] In certain computing units, e.g., in a system-on-a-chip (SoC) or a microcontroller in vehicles, direct memory access (DMA) can often be performed for simple copying operations to relieve the load on the computing unit's processors. A corresponding direct memory access unit for performing such access can read from one memory area (e.g., from RAM, non-volatile memory (NVM), a peripheral unit, etc.) and write to another memory area. Using such direct memory access means that the processors' program flow does not need to be interrupted. The copying process itself can take place in the background without intervention from the processors.

[0003] Direct memory access units typically have a specific number of DMA channels, which can have independent configurations regarding read addresses, write addresses, data width, and the number of read / write accesses. Each DMA channel can be configured for a specific application, possibly project-specifically, e.g., for reading an analog / digital conversion value and writing it to a CPU-level RAM memory for further processing of the raw value and translating it into a physical quantity. Thus, individual direct memory access units can each perform a multitude of read / write accesses with a specific number of DMA channels in the respective processing unit.

[0004] Direct memory accesses on a DMA channel can be initiated and executed, for example, by a hardware event or a software event. However, the more such direct memory accesses are to be performed, the greater the risk of overloading the respective direct memory access unit itself or the bus system over which the respective data is transferred. Disclosure of the invention

[0005] According to the invention, a computer-implemented method for determining a configuration for performing direct memory accesses in a computing unit, as well as a configuration computing unit and a computer program for implementing the method, as well as a computing unit, in particular a microcontroller configured according to such a method, are proposed, having the features of the independent patent claims. Advantageous embodiments are the subject of the dependent claims and the following description.

[0006] The computing unit can be designed, in particular, as a system-on-a-chip (SoC) or as a microcontroller. The computing unit can, for example, be provided to perform functions (particularly safety-critical ones) that are executed for the operation and / or control of a device comprising the computing unit. The computing unit can, in particular, be used in a vehicle and, there, execute, for example, engine control or driver assistance functions.

[0007] A plurality of direct memory accesses, hereinafter also referred to as "DMA accesses," are to be performed according to predefined access properties or constraints. For example, during these DMA accesses, a memory area and / or a peripheral unit of the computing unit is to be accessed, expediently without intervention by a processor unit of the computing unit. Each direct memory access can comprise a read and a write access. The predefined access properties describe, for example, temporal properties or specifications according to which the respective access is to be performed or which are to be adhered to during the respective access.

[0008] A plurality of direct memory access units or direct memory access modules, hereinafter also referred to as "DMA units," are each configured to perform direct memory accesses. These DMA units can be integrated into the computing unit as internal units or can be connected to the computing unit as external units.

[0009] Within the framework of the proposed method, an optimization method or optimization process is carried out, in particular depending on the specified access characteristics of the individual DMA accesses and the individual available direct memory access units. The result of this optimization method is an optimized access configuration, hereinafter also referred to as the "DMA configuration," wherein this access configuration defines which direct memory accesses of the plurality of direct memory accesses are to be performed by which direct memory access units of the plurality of direct memory access units. In particular, the specific access configuration defines which DMA accesses are to be performed via which DMA channel of the individual DMA units.

[0010] The proposed method can be implemented, for example, during a configuration phase of the respective computing unit. The specific access configuration can then be implemented in the respective computing unit. After commissioning, the individual DMA accesses can be performed during regular operation of the computing unit according to the implemented access configuration.

[0011] With the help of the optimization method, an optimized DMA utilization profile can be expediently determined as the access configuration, which describes how the individual DMA accesses to be carried out can be carried out by the individual available DMA units, so that there can be as little as possible an overload of the individual DMA units or of the vehicle's communication systems and so that time specifications for the individual DMA units can be met.

[0012] The invention utilizes the measure of mathematically describing or abstracting the DMA accesses to be performed by the computing unit and performing mathematical, stochastic, or numerical optimization to find an optimal access configuration. During the optimization process, for example, various possible access configurations can be determined and analyzed, particularly with regard to whether individual DMA units or individual communication systems are overloaded, until the best possible access configuration is found.

[0013] Particularly useful in the course of the optimization process, optimal DMA configurations can be determined based on the temporal characteristics of required DMA channels and the available DMA units, which provide an optimized DMA load for all DMA units. In this way, unfavorable DMA configurations in which a DMA unit could be overloaded can be conveniently identified, and one or more DMA configurations can be generated that can lead to an optimized DMA load for each DMA unit.

[0014] Optimization methods or optimization are generally understood as analytical or numerical calculation methods to find optimized, especially minimized or maximized parameters of a complex system. For this purpose, an optimization problem can be formulated, where a solution space Ω, i.e. a set of possible solutions or variables x→, and a target function f are specified. To solve this optimization problem, a set of values ​​of the variables or solutions x→∈Ω searched so that f(x→) a given criterion is met, for example, maximum or minimum. Furthermore, boundary or secondary conditions can also be specified, whereby feasible solutions x→ must fulfill these specified boundary conditions.

[0015] Particularly usefully, the optimization process can include a so-called genetic or evolutionary algorithm. In the course of such an optimization process, which is based on biological evolution, a first generation of solution candidates can be generated, for example, randomly, in an initialization step. Each of these individual solution candidates can be assigned a value of the objective function, known as the fitness function, depending on their quality. Depending on these assigned values, individual solution candidates can be selected as the starting point for a subsequent generation. In the course of a generation loop, a new generation of solution candidates can then be generated in each iteration step, for example by combining or modifying the solution candidates selected as the starting point of the previous generation.These newly generated solution candidates can also each be assigned a value of the fitness function, and based on these values, starting points for the next iteration step can be selected. The loop can be executed until a predefined termination criterion is met. One or more of the solution candidates from the last generation can then be selected as the solution to the optimization problem.

[0016] According to one embodiment, the access properties include a period within which individual or multiple direct memory accesses are to be performed. For example, a respective period can be initiated or started by a predefined hardware and / or software event. For example, the respective direct memory accesses can thus be performed periodically with a fixed period. However, it is also conceivable, in particular, for individual or multiple direct memory accesses to be performed non-periodically.

[0017] Alternatively or additionally, according to one embodiment, the access characteristics comprise a number of individual direct memory accesses to be performed per period. For example, several individual read and / or write accesses can be performed within the period.

[0018] Alternatively or additionally, according to one embodiment, the access properties comprise a duration or transfer duration of a respective direct memory access, in particular an actual or to duration of the respective access.

[0019] Alternatively or additionally, according to one embodiment, the access properties include a maximum or maximum permitted duration or deadline for a respective direct memory access. This maximum duration should not be exceeded, for example, for security reasons, e.g., to meet a real-time criterion.

[0020] Alternatively or additionally, according to one embodiment, the access properties comprise a chain of direct memory accesses, wherein each successfully executed direct memory access triggers another direct memory access. For example, reading a value from a peripheral unit and writing this read value to a RAM of the computing unit as a first direct memory access can trigger a second direct memory access, e.g., reading a timestamp and storing it in the RAM.

[0021] According to one embodiment, a utilization or a utilization profile of the individual direct memory access units is determined and the optimization method is carried out depending on this determined utilization of the individual direct memory access units.

[0022] In particular, to determine the utilization of a respective DMA unit, it is possible to take into account which DMA accesses the respective DMA unit is to perform. For example, the total duration of these DMA accesses to be performed can be determined. In particular, the utilization can be determined depending on the direct memory accesses to be performed in a respective period and depending on the durations and the maximum durations of these direct memory accesses to be performed in the period. It is particularly useful to check whether the direct memory accesses to be performed in a respective period can be performed in this period. In order to be able to precisely determine the utilization of the DMA units, the duration of each direct memory access can be determined, e.g.through actual measurements using the respective computing unit and / or through a particularly cycle-accurate simulation of the respective computing unit.

[0023] For example, during the optimization process, it can be investigated how the utilization of the individual DMA units would change for different access configurations in order to find an optimized DMA configuration with which the best possible utilization of the individual DMA units can be achieved. For example, a percentage utilization or a standardized utilization can be determined as the utilization of a respective DMA unit, so that the utilizations of the individual DMA units can be compared with each other in a representative manner. For example, each DMA channel can be represented with a respective standardized utilization based on the access characteristics of the respective DMA accesses to be performed. This standardized utilization indicates how heavily each DMA channel would load the respective DMA unit in which the channel is configured. Even if different DMA channels have different access characteristics, e.g.different periods, deadlines, transfer durations, etc., they can be normalized for easier analysis and prediction of the DMA configuration.

[0024] According to one embodiment, the optimization method comprises a genetic algorithm. During this process, a first step is performed, wherein a first, original plurality or a first generation of possible access configurations is determined, e.g., by a random permutation. Individual possible access configurations are selected from this first plurality of possible access configurations depending on the respective utilization of the individual direct memory access unit. For example, those possible access configurations are selected according to which the respective utilization of the individual direct memory access units is below a predetermined threshold or which do not lead to an overload of the individual direct memory access units.

[0025] A second step is then carried out iteratively. In this second step, a second plurality or a new generation of possible access configurations is determined as a starting point, depending on the possible access configurations selected in the previous step. For example, these new possible access configurations can be determined by combining individual possible access configurations selected in the previous step and / or by modifying individual possible access configurations selected in the previous step. From the second plurality of possible access configurations, individual possible access configurations are selected as a starting point for the next step or the next generation, depending on the respective workload.Here, too, those possible access configurations are selected according to which the respective utilization of the individual DMA units is below the specified threshold. For example, the fitness function can define a deviation of the utilization of the individual direct memory access unit from an average utilization of all direct memory access units.

[0026] The iteration is repeated until a specified number of second steps is completed and / or until a specified criterion or termination criterion is met. The optimized access configuration is then determined as the result of the optimization procedure, depending on the possible access configurations selected in the previous step.

[0027] According to one embodiment, the plurality of direct memory access units is configured to access a plurality of communication systems. These individual communication systems can each be configured, for example, as a bus or fieldbus system, e.g., as CAN, Ethernet / IP, ProfiNet, Sercos 2, Sercos III, EtherCAT, FlexRay, LIN, MOST, etc. For example, one or more of the communication systems can each be configured as a peripheral bus to connect peripheral units to the computing unit.

[0028] According to one embodiment, a communication system configuration or bus configuration is determined depending on the specific access configuration and, in particular, depending on the access characteristics of the individual DMA accesses. This specific communication system configuration defines via which of the plurality of communication systems the individual direct memory accesses are performed by the respective direct memory access unit according to the specific access configuration.

[0029] DMA accesses can be delayed not only by DMA utilization or overload of the respective DMA units, but also by the load on the individual communication systems. For example, if multiple DMA units access the same communication system, DMA accesses may not be completed on time, even with low DMA utilization of the individual DMA units. This can lead to individual DMA accesses being started with a delay and not being able to complete within the specified maximum duration.

[0030] Such delays can be effectively prevented by determining the communication system configuration. For example, a prediction of the utilization of the individual communication systems can be made to change the allocation of the individual DMA channels to peripheral resources so that delays cannot occur.

[0031] Particularly usefully, the specific access configuration can prevent delays in DMA accesses due to overloads of DMA units, and the specific communication system configuration can prevent delays in DMA accesses due to accesses by multiple DMA units to the same communication system or due to overloads of individual communication systems.

[0032] According to one embodiment, an effective utilization of each direct memory access unit is determined for each communication system, wherein this effective utilization characterizes or takes into account direct memory accesses of the respective direct memory access unit via the respective communication system as well as direct memory accesses of the other, remaining direct memory access units via the respective communication system. The communication system configuration is determined depending on the determined effective utilization of the individual direct memory access units.This effective utilization of a respective DMA unit includes, in particular, the respective unit's own utilization, which corresponds in particular to an ideal transmission time without the influence of other DMA units, and additionally, in particular, an additional load corresponding to the delay of the transmission processes of the respective DMA unit that other DMA units could cause by simultaneously transmitting data via the same communication system.

[0033] According to one embodiment, a dedicated utilization or inherent load ("DMA transfer load") of each direct memory access unit is determined, which characterizes the direct memory accesses of the respective direct memory access unit via a respective communication system of the plurality of communication systems. This dedicated utilization characterizes, in particular, the total duration of all accesses to be performed by the respective DMA unit.

[0034] Furthermore, a communication system load ("bus load") of each communication system is determined, which characterizes the direct memory accesses of the individual direct memory access units via the respective communication system. In particular, this communication system load can be determined depending on the individual direct memory access units' own load.

[0035] Furthermore, an additional load ("Add-On DMA Load") of each direct memory access unit is determined, which characterizes direct memory accesses of the other, remaining direct memory access units via a respective communication system. For a respective DMA unit that performs DMA accesses via a respective communication system, the additional load describes in particular the duration of DMA accesses performed by the other DMA units via this respective communication system. This additional load can, in particular, lead to a delay in the respective DMA accesses of the respective DMA unit. This additional load can be determined in each case depending on the individual load of the individual direct memory access units and depending on the communication system load of the individual communication systems.

[0036] The effective utilization of each direct memory access unit is then determined based on its own utilization and the additional utilization of the individual direct memory access units. This makes it particularly useful to forecast DMA bus profiles with respect to the inherent load and the effective load due to the delay caused by a bottleneck on the communication systems based on the temporal access characteristics of the individual DMA accesses, the DMA units assigned to them, and communication system or peripheral resources. This not only makes it possible to detect and avoid incorrect DMA configurations, which could overload the DMA units themselves, but also to investigate whether there is a risk of effective overloading of a DMA module.

[0037] According to one embodiment, the communication system configuration is determined such that the effective utilization of each direct memory access unit is below a predetermined threshold. This threshold is expediently selected such that DMA accesses of a respective DMA unit via a respective communication system cannot be delayed by accesses from other DMA units via this respective communication system.

[0038] According to one embodiment, the plurality of direct memory access units is configured according to the determined access configuration and, furthermore, in particular, according to the determined communication system configuration. The optimized access configuration and, furthermore, expediently, the determined communication system configuration can thus be implemented in the computing unit or, respectively, in the vehicle. During subsequent operation of the computing unit or, respectively, the vehicle, the plurality of DMA accesses is expediently performed by the corresponding DMA units according to the determined access configuration, furthermore, in particular, via the corresponding communication systems according to the determined communication system configuration.The access configuration and, in particular, the communication system configuration can be determined in particular during a configuration phase of the computing unit, during which the computing unit manufactured by a computing unit manufacturer is to be configured for its intended use in a particular vehicle. For example, the computing unit can be provided by the respective manufacturer with fixed wiring between the communication or bus systems, the DMA units, respective triggers, etc. These wirings can, in particular, be multiplexer connections, whereby, for example, a trigger can be connected to several DMA units, whereby, for example, a DMA unit can read and write via several buses, etc.During the configuration phase, the access configuration and the communication system configuration can be determined in particular in such a way that the DMA accesses to be carried out in subsequent, regular operation are distributed among the individual DMA units and bus systems depending on the multiplexer connections specified by the manufacturer in such a way that the individual DMA units and individual bus systems cannot be overloaded.

[0039] A configuration computing unit according to the invention is configured, in particular in terms of programming, to carry out a method according to the invention. Such a configuration computing unit can be, for example, a computer, PC, or similar device.

[0040] The implementation of a method according to the invention in the form of a computer program or computer program product with program code for carrying out all method steps is also advantageous, since this entails particularly low costs, in particular if an executing control unit is also used for other tasks and is therefore already present. Finally, a machine-readable storage medium is provided with a computer program stored thereon, as described above. Suitable storage media or data carriers for providing the computer program are, in particular, magnetic, optical and electrical memories, such as hard disks, flash memories, EEPROMs, DVDs, and others. Downloading a program via computer networks (Internet, intranet, etc.) is also possible. Such a download can be wired or cable-based or wireless (e.g. via a WLAN network, a 3G, 4G, 5G or 6G connection, etc.).

[0041] The invention also relates to a computing unit, in particular a system-on-a-chip (SoC) or a microcontroller, which is configured according to a method according to the invention.

[0042] Further advantages and embodiments of the invention will become apparent from the description and the accompanying drawings.

[0043] The invention is illustrated schematically in the drawing using exemplary embodiments and is described below with reference to the drawing. Short description of the drawings Fig. 1 schematically shows a computing unit which may form the basis of an embodiment of the method according to the invention. Fig. 2 shows schematically an embodiment of the method according to the invention as a block diagram. Embodiment(s) of the invention

[0044] In Fig. 1, a computing unit is schematically shown and designated 100. The computing unit 100 can be embodied, for example, as a microcontroller. The computing unit 100 can be provided, for example, in a vehicle for performing safety-critical vehicle functions, e.g., in the course of engine control or driver assistance functions.

[0045] The microcontroller 100 comprises a processor unit 150, a RAM memory unit 160 (SYSRAM) and a timer 170 (GTM module), which are connected to a higher-level communication system 180. For example, the Fig. Element 160 shown in Figure 1 may represent a memory block or a memory segment of the RAM memory. Element 170 may, for example, represent a memory block or a memory segment of the GTM module.

[0046] The microcontroller 100 further comprises a plurality of peripheral units, wherein Fig. 1 Memory units of these individual peripheral units are schematically illustrated as a common memory block 130. These peripheral units can each be provided, in particular, as modules within the microcontroller 100, e.g., to control the outside world or units outside the microcontroller 100. For example, one or more of the peripheral units can be configured as analog-to-digital converters (ADCs), via which sensors of the vehicle can be connected to the microcontroller 100.

[0047] The individual peripheral units can each be connected via different communication systems 120, e.g. via a peripheral bus system. Fig. For example, Figure 1 shows three peripheral bus systems 121, 122, and 123. In the following, the peripheral bus system 121 is also referred to as PERBUS0, the peripheral bus system 122 as PERBUS1, and the peripheral bus system 123 as PERBUS2.

[0048] Microcontroller 100 has a plurality of direct memory access units (DMA units) 110, each configured to perform direct memory accesses (DMA accesses). For example, the individual DMA units 110 can access peripheral units 130, RAM 160, and timer 170 using such DMA accesses without interrupting the program flow of processor unit 150 for these accesses.

[0049] The individual peripheral units or peripheral modules can each have dedicated memory regions, for example. Fig. The vertical lines shown in Figure 1 between the shared memory block 130 of the peripheral units and the peripheral bus systems 120 are intended to represent such dedicated memory regions or memory slots. For example, a delta-sigma ADC module can have a specific memory region, with one of the vertical lines then representing this slot for the corresponding delta-sigma module. The DMA modules 110 can also have dedicated memories as part of the peripheral memory 130 and can access this via specific peripheral buses 120. For example, a specific DMA transfer can be provided from a specific DMA memory region to another memory region, e.g., the system RAM 160.

[0050] In Fig. For example, Figure 1 shows four peripheral DMA units 111, 112, 113, and 114. In the following, DMA unit 111 is also referred to as DMAO, DMA unit 112 is also referred to as DMA1, DMA unit 113 is also referred to as DMA2, and DMA unit 114 is also referred to as DMA3.

[0051] It is understood that the microcontroller 100 may include further elements, such as further RAM memories, e.g., system RAM blocks, cluster local RAM, etc., further storage units, e.g., non-volatile memories (NVM), further peripheral units, further communication systems, etc.

[0052] During operation of microcontroller 100, a plurality of direct memory accesses are to be performed, each according to predetermined access characteristics. These individual DMA accesses are to be performed by the individual DMA units 110 in such a way that no overloading of individual DMA units 110 occurs, which could lead to delays in the execution of individual accesses. Furthermore, the DMA accesses are to be performed in such a way that no overloading of the individual peripheral bus systems 120 occurs, which could also lead to delays in the execution of individual accesses.

[0053] For this purpose, according to one embodiment of the method according to the invention, a configuration for performing the direct memory accesses is determined, as described below with reference to Fig. 2 is explained.

[0054] In Fig.2, an embodiment of a method according to the invention is shown schematically as a block diagram.

[0055] During a configuration phase 210 of the microcontroller 100, it is to be determined which DMA accesses are to be performed by which DMA unit 110 and, furthermore, via which bus systems 120 these DMA accesses are to be performed. For example, the microcontroller 100 can be provided by a manufacturer with fixed multiplexer wiring between the communication systems 120 and the DMA units 110. During configuration phase 210, the microcontroller 100 is to be configured based on these predetermined multiplexer connections in such a way that no overloading of the individual DMA units 110 and the individual bus systems 120 can occur during subsequent regular operation.

[0056] For this purpose, access properties of the individual DMA accesses to be performed are determined in a step 211.

[0057] In a step 212, an optimization method is carried out depending on the access properties of the individual DMA accesses and depending on the available DMA units 110. For example, this optimization method can comprise a genetic or evolutionary algorithm.

[0058] In a step 213, an optimized access configuration or DMA configuration is determined as a result of the optimization method, wherein the determined access configuration defines which DMA accesses are performed by which DMA units 110.

[0059] In a step 214, a communication system configuration is determined depending on the specific access configuration and depending on the access properties, wherein this communication system configuration defines via which peripheral bus system 120 the individual DMA accesses are carried out by the respective DMA unit 110 according to the specific access configuration.

[0060] In a step 215, the microcontroller 100 is configured according to the determined access configuration and the determined communication system configuration. The microcontroller 100 can be used, for example, in a vehicle.

[0061] The vehicle and microcontroller 100 are then put into operation. During regular operation 220 of microcontroller 100, the individual DMA accesses are performed according to the implemented access and communication system configuration.

[0062] The following explains, using an example, how the access properties of the individual DMA accesses to be performed can be determined in step 211.

[0063] For example, the access properties include a period within which one or more direct memory accesses are to be performed. For example, the period can be triggered or started by a hardware or software event. The respective direct memory accesses can thus be performed periodically with a fixed period. However, one or more direct memory accesses can also be performed non-periodically. The access properties further include a number of individual direct memory accesses to be performed per period, a duration of each direct memory access, and a maximum duration or deadline for each direct memory access, within or by which the respective direct memory access is to be terminated.Furthermore, the access properties can define a chain of direct memory accesses, whereby each successful direct memory access triggers another direct memory access.

[0064] The following explains, using an example, how the optimization procedure can be carried out in steps 212 and 213 and how the access configuration can be determined.

[0065] In the following example, the following five DMA accesses are to be performed four times. 1. DMA access: Reading from an ADC 130 and writing to RAM 160. ADC values ​​are to be read from an ADC module and written to RAM (1st DMA channel ADC→RAM). The characteristics of this DMA access are: access period = 3µs, number of accesses = 1, transfer duration = 200ns, and deadline = 3µs. Possible DMA units that can be configured for this access include DMAO 111, DMA1 112, or DMA3 114. 2. DMA access: Read timer and write to RAM A timestamp is to be read from timer 170, and this value is to be written to RAM (2nd DMA channel, timer → RAM). The characteristics of this DMA access are: access period = 3µs, number of accesses = 1, transfer duration = 300ns, deadline = 3µs. This DMA access is to be performed in a chain with the first DMA access, i.e., triggered by the first DMA access. In a mathematical model, these two DMA accesses would load the respective DMA module at 17%. 3. DMA access: Read DMA register and write to RAM DMA register values ​​are to be read and written to RAM (3rd DMA channel DMA→RAM). The characteristics of this DMA access are: access period = 40µs, number of accesses = 1, transfer duration = 500ns, and deadline = 3µs. For example, the possible DMA modules that can be configured for this access are the DMAO 111, DMA2 113, or DMA3 114 units. 4. DMA access: Reading from RAM and copying to SYSRAM The timestamp value should be read from RAM and written to SYSRAM (4th DMA channel RAM→SYSRAM). The characteristics of this DMA access are: access period = 40µs, number of accesses = 1, transfer duration = 100ns, and delay = 40µs. This DMA access should be chained to the third access, meaning it should be triggered by the third DMA access immediately after this access is completed. The DMA unit configured for the third DMA access must also be configured for this fourth DMA access. 5. DMA access: Reading from RAM and copying to SYSRAM The ADC values ​​should be read from RAM and written to SYSRAM (5th DMA channel RAM→SYSRAM). The characteristics of this DMA access are: access period = 40µs, number of accesses = 13, transfer duration = 300ns, and delay = 40µs. This DMA access should also be performed in a chain with the third DMA access, i.e., it should be triggered by the third DMA access immediately after the completion of this access. The DMA unit configured for the third DMA access must also be configured for this fifth DMA access.

[0066] In a mathematical model, the third, fourth and fifth DMA accesses would load the respective DMA unit at 26%.

[0067] For example, each DMA access can be represented with a normalized DMA load value based on its characteristic parameters. This value indicates how much each DMA access would load the DMA in which it is configured. Even though different DMA accesses have different characteristic parameters, such as different periods, deadlines, transfer durations, etc., they can be normalized to facilitate analysis and prediction of the DMA configuration. This normalization also makes it possible to generate optimized DMA configurations.

[0068] To find the optimal DMA configurations, in general, each required DMA channel can be defined by its characteristic parameters.

[0069] An example of an optimization algorithm is a genetic algorithm. For example, each individual in the genetic algorithm represents a permutation of possible access configurations. For example, a DMA configuration would be one in which most DMA channels (e.g., 16 of 20) are configured in the DMAO module and only four in the DMA1 module. This DMA configuration is assumed to overload the DMAO module. This configuration would therefore be identified as an unsuitable individual during optimization. This individual would not survive; it would be suspended and not be part of the next generation. Only the fittest individuals of a generation survive, and only from these individuals are new individuals—i.e., new possible access configurations—generated, e.g., by crossing and mutating to generate the next generation of possible DMA configurations.

[0070] During the genetic algorithm, a first generation of possible DMA configurations is generated by randomly permuting the DMAs for each DMA channel. DMA configurations that overload the DMA bus will not survive. For example, a threshold of 80% can be specified, above which a DMA module is overloaded. This threshold can be defined, for example, depending on the microcontroller used and as needed.

[0071] In an iteratively executed second step, a new DMA configuration is formed from two selected possible DMA configurations by crossover. This means that part of the DMA configuration is taken from a first parent and another part from a second parent. The size of the parts can be fixed or changed randomly. Here, too, only those DMA configurations will survive that do not overload any DMA module. Subsequently, a so-called mutated DMA configuration is generated in which, upon random selection of one or more DMA channels, the DMA module is randomly changed. In each step of generating a new member, only those DMA configurations that do not overload any DMA module survive. The fittest DMA configurations are selected to be part of the next generation, i.e., the next iteration. The fitness function defines the degree to which the load of the DMA modules deviates from the mean load.

[0072] After a large number of generations, the algorithm will find one or more optimal DMA configurations. The number of DMA configurations can be flexibly specified. One or more of these suggested DMA configurations can then be used as the optimized access configuration for the microcontroller's actual DMA configuration.

[0073] The following example explains how the communication system configuration can be determined in step 214.

[0074] Not only the DMA load itself, but also the load on the PERBUS buses can delay DMA accesses. For example, DMA accesses may not be completed on time even with a low DMA load if other DMA modules are using the same PERBUS. The DMA transfers may then be started with a delay and therefore may not be completed within the specified deadline. According to the embodiment of the present invention, such overloading is therefore predicted.

[0075] Based on the temporal characteristics of the DMA accesses and the DMA modules assigned to them according to the access configuration determined in step 213, and further based on the peripheral resources, DMA bus profiles can be predicted with respect to the intrinsic load as well as the effective load due to the delay on PERBUS buses caused by the bottleneck. This not only allows for the detection of incorrect DMA configurations in which the DMA modules could be overloaded, but also allows for the investigation of whether an effective overload of a DMA module is to be expected.

[0076] For this purpose, an own load or DMA self-load is taken into account, which characterizes a best possible normalized transmission time without collision on a PERBUS.

[0077] Furthermore, an effective utilization or effective DMA load is taken into account, which characterizes a normalized transfer time, which leads to a longer total DMA transfer time if the buses are used by other units at the same time in the worst case.

[0078] In this way, it can be taken into account that the DMA transfers of a DMA module can also be influenced by other units, e.g., other DMA units, the CPU, etc., if this data is transferred over the same PERBUS bus. For the sake of simplicity, only other DMA units will be considered in this example, and the time delay of DMA transfers by other units, e.g., when the CPU uses the same bus, will be neglected. This neglect can be justified on the basis of system measurements, according to which the time required by all CPU cores to read or write data to the peripheral units is approximately 2% of the total time for all units to read and write to the peripheral units. In particular, writing to or reading from the same RAM or the system RAM has no influence on the DMA transfers. Parallel reading of timestamps from the timer orIn particular, sharing a GTM module with multiple DMA modules is possible without delaying DMA transfer times. Multiple DMA modules are not allowed to write to the GTM module, so this could, in principle, affect DMA transfer times. For this example, assume that the GTM module version does not allow different DMA modules to write to the memory segments reserved for GTM.

[0079] Determining the communication system configuration also covers the aforementioned neglected influences, such as other units and DMA modules accessing the peripheral bridges. However, the following example specifically considers only additional DMA units and time delays due to bottlenecks in the PERBUS system.

[0080] In particular, the following parameters are determined: 1. individual utilization of each DMA unit

[0081] Each DMA transfer operation has a DMA transfer load I dt . The dead load on the respective DMA module L self,DMAi corresponds to the sum of all DMA transfer operations assigned on the DMA module I dt . 2. Communication system utilization of each communication system

[0082] Every DMA access that reads or writes over the peripheral bus has a PERBUS read or PERBUS write load. Similar to the definition of the DMA load, one can also define the read and write load. To calculate the read load, I read the reading time is taken into account. Accordingly, the write load I write the transfer time of the writing phase is used.

[0083] In particular, only read / write operations via the peripheral bus are considered below. Reading or writing in other memory segments is neglected. Therefore, depending on the DMA transfer process, the sum of the read and transfer load is either equal to or less than the DMA transfer load. An example where the DMA transfer load is greater than the sum of the read and write load is when only one phase, e.g., the read phase, is performed via the PERBUS and the other phase, e.g., the write phase, is performed via a Network-on-Chip (NOC). The transfer process depends, for example, on the complex driver software for the respective peripheral module. If the transfer is designed such that both read and write operations are performed via the PERBUS, then the sum of the read and write transfer loads via the PERBUS is equal to the DMA transfer load. In general, the following applies in particular: I read + I write ≤ I dt

[0084] The communication system load or PERBUS load for a respective PERBUS bus 'i' (PERBUSi) is the sum of the read and write loads of all DMA transfers that transfer data over the corresponding PERBUS: LPERBUSi=∑j=DMAeijlij 3. additional utilization of each direct memory access unit

[0085] The additional utilization or add-on DMA load of a respective DMA unit 'k' (DMAk) corresponds to the sum of the read and write loads of all other DMA transfers that use the same PERBUS as the respective DMA unit 'k': Laddon,DMAk=∑i=PERBUS∑j=DMAj≠keikeijlij It is ik = 0 if the DMA module 'k' does not transmit via the PERBUS 'i', and e ik = 1 if the DMA module 'k' transmits via the PERBUS 'i'.

[0086] The full load that DMAk places on PERBUSi is the sum of all read / write transfer phases from DMAk through PERBUSi: lik=∑r / won PERBUSi and DMAklr / w 4. Effective utilization of each direct memory access unit

[0087] The effective load of the DMAk module is composed of its own load, which corresponds to the ideal transfer time without the influence of other DMA modules, and the additional load corresponding to the delay of the DMAi module's transfer operations that other DMA modules could cause by simultaneously transferring data over the same PERBUSES: Leff,DMAk=Lself,DMAk+Laddon,DMAk

[0088] The following example considers that the following six DMA accesses are to be performed four times: 1. DMA access: reading from ADC and writing to RAM

[0089] ADC values ​​are to be read from the ADC module and written to RAM (1st DMA channel ADC→RAM). The characteristics of this DMA access are: access period = 2µs, number of accesses = 1, transfer duration = 200ns, and delay = 2µs. For example, according to the specific access configuration, DMA unit 112 (DMA1) is to perform these accesses. The load of this transfer is, for example, 10%.

[0090] The reading process occurs via the PERBUS bus. Each peripheral resource is mapped to a specific PERBUS. For example, the respective ADC peripheral unit should be assigned to PERBUS1. For simplicity, we assume that the read and write load is 1 / 2 of the total transfer load, i.e., we assume that half of the transfer time is used for reading and half for writing.

[0091] The write operation of this access is ignored, as the impact of reading or writing on the memory is negligible. In a mathematical model, this access would place a 10% load on the DMA1 module and a 5% load on the PERBUS1 module.

[0092] Assume that there are three more DMA accesses corresponding to this: ADC values ​​are read from the ADC peripheral memory card and stored in RAM. 2. DMA access: reading from ADC and writing to RAM

[0093] Allocation to DMA1 and PERBUS1. For example, the transfer load on DMA1 is 10% and the read load on PERBUS1 is 5%. 3. DMA access: reading from ADC and writing to RAM

[0094] Allocation to DMA1 and PERBUS1. For example, the transfer load on DMA1 is 10% and the read load on PERBUS1 is 5%. 4. DMA access: reading from ADC and writing to RAM

[0095] Allocation to DMA1 and PERBUS1. For example, the transfer load on DMA1 is 10% and the read load on PERBUS1 is 5%.

[0096] The sum of these four loads allocated to DMA1 on PERBUS1 is, for example, 20%. 5. DMA access: Read SPI value and copy to RAM

[0097] SPI values ​​are to be read and written to RAM (5th DMA channel SPI→RAM). The characteristics of this DMA access are: access period = 16µs, number of accesses = 40, transfer duration = 150ns, and deadline = 16µs. The utilization of this access is therefore approximately 36%. For example, according to the access configuration, this transfer should be performed by the DMA2 unit. For example, PERBUS1 should be used. The load of this transfer on PERBUS1 is then approximately 18%. 6. DMA access: Reading from RAM and copying to SPI

[0098] An SPI value stored in RAM is now to be stored in the peripheral memory for the user (6th DMA channel RAM→SPI). The characteristics of this DMA access are: access period = 16µs, number of accesses = 40, transfer duration = 150ns, and delay = 16µs. The utilization of this transfer is therefore approximately 38%. For example, according to the access configuration, this transfer should be performed by the DMA2 unit. For example, PERBUS1 should be used. The load of this transfer on PERBUS1 is approximately 18%. The sum of these two loads that DMA2 imposes on PERBUS1 is 36%.

[0099] In this example, it is predicted that transmission over PERBUS1 will affect the effective load of DMA1 and DMA2, even if there is no overload on the PERBUS1 bus. For the example above, an overload is predicted in DMA2.

[0100] DMA1: Dead load = 40% Effective load = 40% + 2*18% = 76%. For transfers configured on the DMA1 bus, no transfer delays are expected that could jeopardize the deadline.

[0101] DMA2: Dead load = 72% Effective utilization = 72% + 4*5% = 92% For transfers configured on DMA2, it is expected that there may be certain situations where a DMA transfer of the SPI values ​​cannot be completed within the deadline. Transfers initiated on DMA2 may be delayed by transfers initiated by DMA1 that began shortly before DMA2. The bottleneck here is the use of the same PERBUS1 bus, even though this peripheral bus's utilization is approximately 56%. While higher-priority transfers can be configured to interrupt lower-priority DMA transfers on the same DMA module, two transfers from two different DMA modules on the same PERBUS cannot be configured to interrupt each other.

[0102] As a solution to this situation, the communication system configuration can be determined such that the 5th and 6th DMA access should be performed via a PERBUS other than PERBUS1, e.g. via PERBUS2.

[0103] By designing the invention, DMA accesses can thus be executed in the microcontroller in such a way that delays can occur due to overloads of DMA units and due to accesses by multiple DMA units to the same communication system.

Claims

[1] Computer-implemented method for determining a configuration for performing direct memory accesses in a computing unit (100), wherein a plurality of direct memory accesses are to be performed in each case according to predetermined access properties and wherein a plurality of direct memory access units (110, 111, 112, 113, 114) are configured to perform direct memory accesses, the method comprising the following steps: Carrying out (212) an optimization method and determining (213) an access configuration as a result of the optimization method, wherein the determined access configuration defines which direct memory accesses of the plurality of direct memory accesses are carried out by which direct memory access unit of the plurality of direct memory access units (110, 111, 112, 113, 114). [2] The method of claim 1, wherein the access properties comprise one or more of the following properties: a period within which one or more direct memory accesses are performed; a number of individual direct memory accesses performed per period; a duration of a respective direct memory access; a maximum duration of each direct memory access; a chain of direct memory accesses, where each direct memory access triggers another direct memory access. [3] The method of claim 1 or 2, further comprising: Determining a utilization of the plurality of direct memory access units (110, 111, 112, 113, 114) and Carrying out the optimization method (212) depending on the determined utilization of the plurality of direct memory access units (110, 111, 112, 113, 114). [4] The method of claim 3, wherein performing the optimization process comprises: Carrying out a first step, wherein a first plurality of possible access configurations is determined and wherein individual possible access configurations are selected from this first plurality of possible access configurations depending on the respective utilization of the individual direct memory access unit; repeatedly performing a second step, wherein in the second step a second plurality of possible access configurations is determined depending on the possible access configurations selected in the respective previous step, in particular by combining individual ones of the possible access configurations selected in the respective previous step with one another and / or by changing individual ones of the possible access configurations selected in the respective previous step, and wherein individual possible access configurations are selected from this second plurality of possible access configurations depending on the respective utilization of the individual direct memory access unit; Determining, after a predetermined number of second steps have been carried out and / or after a predetermined criterion has been met, the access configuration as a result of the optimization process depending on the possible access configurations selected in the previous step. [5] A method according to any one of the preceding claims, wherein the plurality of direct memory access units (110, 111, 112, 113, 114) are arranged to access a plurality of communication systems (120, 121, 122, 123), the method further comprising: Determining (214) a communication system configuration depending on the determined access configuration, wherein the determined communication system configuration defines via which communication system of the plurality of communication systems (120, 121, 122, 123) the individual direct memory accesses are carried out according to the determined access configuration by the respective direct memory access unit (110, 111, 112, 113, 114). [6] The method of claim 5, further comprising: Determining an effective utilization of each direct memory access unit of the plurality of direct memory access units (110, 111, 112, 113, 114) for each communication system of the plurality of communication systems (120, 121, 122, 123), wherein this effective utilization characterizes direct memory accesses of the respective direct memory access unit via the respective communication system as well as direct memory accesses of the respective other direct memory access units via the respective communication system; and Determining the communication system configuration depending on the determined effective utilization of the individual direct memory access units (110, 111, 112, 113, 114) for the individual communication systems (120, 121, 122, 123). [7] The method of claim 5 or 6, further comprising: Determining a specific utilization of each direct memory access unit (110, 111, 112, 113, 114), which in each case characterizes direct memory accesses of the respective direct memory access unit via a respective communication system of the plurality of communication systems (120, 121, 122, 123); Determining a communication system utilization of each communication system of the plurality of communication systems (120, 121, 122, 123), which in each case characterizes direct memory accesses of the individual direct memory access units via the respective communication system, in particular depending on the individual utilization of the individual direct memory access units; Determining an additional utilization of each direct memory access unit (110, 111, 112, 113, 114), which in each case characterizes direct memory accesses of the respective other direct memory access units via a respective communication system of the plurality of communication systems, in particular depending on the individual utilization of the individual direct memory access units and depending on the communication system utilization of the individual communication systems; and Determining the effective utilization of each direct memory access unit (110, 111, 112, 113, 114) depending on its own utilization and on the additional utilization of the individual direct memory access units. [8] Method according to one of claims 5 to 7, wherein the communication system configuration is determined such that the effective utilization of each direct memory access unit (110, 111, 112, 113, 114) is below a predetermined threshold value. [9] A method according to any one of the preceding claims, further comprising: Configuring (215) the plurality of direct memory access units (110, 111, 112, 113, 114) according to the determined access configuration. [10] Computing unit (100), in particular system-on-a-chip or microcontroller, comprising a plurality of direct memory access units (110, 111, 112, 113, 114) configured according to a method according to claim 9. [11] Configuration computing unit which is configured to carry out all method steps of a method according to one of claims 1 to 9. [12] Computer program which causes a computing unit to carry out all the method steps of a method according to one of claims 1 to 9 when executed on the computing unit. [13] A machine-readable storage medium having stored thereon a computer program according to claim 12.

Citation Information

Patent Citations

  • Optimized multiport NVMe controller for multipath input / output applications

    US10817446B1

  • Neural processing unit (NPU) direct memory access (NDMA) memory bandwidth optimization

    US20200104691A1

  • Efficient buffering technique for transferring data

    US20220147280A1

  • Software management of direct memory access commands

    US20230195664A1

  • US000010817446B1