Slave core memory extension optimization method under master-slave core architecture
By adjusting the slave kernel memory allocation layout and optimizing the buffer channel, the problem of insufficient memory in the master-slave kernel architecture is solved, dynamic expansion of slave kernel memory and efficient data transmission are achieved, and the performance and model execution efficiency of heterogeneous multi-core systems are improved.
Patent Information
- Application Number
- CN202510788479.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-13
AI Technical Summary
Under the master-slave core architecture, the existing memory expansion scheme cannot meet the memory requirements of the deep learning network model, resulting in bottlenecks in master-slave core communication and the problem of less memory allocated by slave cores, limiting the operation performance and transmission speed of slave cores.
By adjusting the layout of the kernel memory allocation, allocating and recycling memory on demand, combining the FreeRTOS memory allocation algorithm and OpenAMP architecture, the buffer channel size and number are optimized, and dynamic expansion and cyclic allocation of kernel memory are achieved to meet the requirements of DNN inference models of different sizes.
It improves the memory utilization rate and data transmission efficiency of the kernel, shortens the transmission time, improves the operation performance of the kernel and the execution efficiency of the model, and maintains the detection effect of the model.
Smart Images

Figure CN120295804A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to an optimization method for slave core memory expansion under a master-slave core architecture. Background Art
[0002] OpenAMP (Open Asymmetric Multi-Processing) is an open-source software framework for managing resources in heterogeneous multi-core systems, especially suitable for embedded systems containing multi-core processors with the ARM architecture. This framework provides a mechanism that allows for efficient data and control message passing between independent operating systems or bare-metal applications. The OpenAMP framework allows different types of processor cores to work together on the same chip, enabling the system to dynamically allocate different cores according to the computing requirements and energy efficiency requirements of tasks.
[0003] In an AMP system, the master-slave cores communicate through shared memory. The master core is responsible for memory management. There are two buffers in each communication direction, namely USED and AVAIL. The buffers can be divided into blocks according to the message format in RPMsg and linked to form a ring. When the master core needs to communicate with the slave core, it can be divided into four steps: 1. The master core first obtains a block of memory from USED; 2. Fill the message according to the message protocol; 3. Link the memory to AVAIL; 4. Trigger an interrupt to notify the slave core that there is a message to process. Conversely, when the slave core needs to communicate with the master core, it is similar.
[0004] The message transmission between the master and slave cores is completed through a ring buffer established by Virtio on the shared memory. Among them, the buffer size is determined by the number of buffer channels and the channel size defined in the system configuration (the product of the two). Affected by traditional message passing, the channel size is usually set very small, such as 512B, resulting in a relatively small buffer size, usually 512x512B.
[0005] With the popularization of Deep Neural Networks (DNNs), more and more business scenarios require deploying DNN model inference to Internet of Things (IoT) devices. Even for lightweight neural network model inference, it requires an input size of nearly 0.5 MB and a memory space of over 100 MB. Common lightweight neural network models, such as Resnet-18, require 144.5 MB of running memory, and Mobilenet_v1 requires 133.1 MB of running memory. Affected by the current buffer size and the fixed memory allocation size of the slave core, the master-slave core transmission performance and the slave core running efficiency of common heterogeneous multi-core IoT devices are limited. Most existing memory expansion solutions are based on enclaves in a trusted execution environment or the secure world, and there are fewer solutions for expanding the memory of the slave core in a master-slave core structure. With the development of neural networks, if a DNN model is to be run on the slave core, it is also necessary to expand the memory of the slave core as needed. The following problems exist: (1) There is a bottleneck in master-slave core communication. Under the OpenAMP architecture, master-slave core messages are completed through a circular buffer established on shared memory by Virtio. Usually, the channel size is small and the number of channels is large. Such a channel allocation strategy can no longer meet the requirements of multi-core DNN collaborative inference, with poor performance and slow speed, resulting in the transmission speed becoming the performance bottleneck of high-real-time systems.
[0006] (2) The allocable memory of the slave core in the master-slave core is less. In the existing master-slave core architecture, the slave core generally runs RTOS as a real-time core, which leads to less allocable memory for the slave core. When it is necessary to run a DNN model on the slave core, the allocable memory will limit the performance of the model and may even not meet the dynamic memory allocation requirements for model operation. Therefore, how to improve the utilization rate of the slave core memory in a resource-constrained master-slave core architecture is particularly important. Summary of the Invention
[0007] The present invention aims to solve the above problems. To this end, the present invention provides a method for optimizing the expansion of the slave core memory under a master-slave core architecture, which adjusts the memory allocation layout of the slave core under the master-slave core architecture, adjusts the size of the slave core memory as needed, and performs cyclic allocation according to the different memory spaces occupied during model operation, so that it can run inference models of different sizes.
[0008] The present invention provides a method for optimizing the expansion of the slave core memory under a master-slave core architecture, and the technical solution adopted is as follows: including: Obtain the model that needs to be run on the slave core, and calculate the total memory required during model operation; Allocate slave core memory from the memory space according to the total memory and the size of the allocable memory of the slave core; Transmit the data of the model to the slave core memory; The slave core performs model inference calculation.
[0009] Further, when the total memory is less than or equal to the allocable memory of the slave core, allocate slave core memory from the memory space; When the total memory is greater than the allocable memory of the slave core, calculate the memory required for each round of model operation; according to the size of the memory for each round and the allocable memory of the slave core, allocate and recycle the slave core memory during each round of model inference calculation.
[0010] Further, when the memory for each round is less than or equal to the allocable memory of the slave core, allocate a memory space not less than the memory for each round as the slave core memory; after each round of model inference calculation is completed, release the slave core memory back to the memory space, and when the next round of model inference calculation starts, still perform memory allocation in the above-mentioned memory space; When the memory for each round is greater than the allocable memory of the slave core, calculate the memory required for each layer of the model during operation, allocate a memory space not less than the memory for each layer as the slave core memory; after each layer of model inference calculation is completed, save the slave core memory occupied by the output data and release the other slave core memory.
[0011] Further, the calculation formula for the memory for each layer is: where, is the memory for each layer, is the read memory size, is the write memory size, is the input data size, is the number of input data, is the output data size, is the number of output data, is the convolution kernel size.
[0012] Further, the allocation and recycling of the slave core memory are implemented based on the memory allocation algorithm of FreeRTOS.
[0013] Further, when transmitting the data of the model to the slave core memory, according to the size of the data and the pipeline size, allocate the number of buffer channels, and then transmit the data.
[0014] Further, pre-allocate buffer channels and buffer memory; When the data size is less than or equal to the pipeline size, directly transmit the data; When the data size is greater than the pipeline size, recycle the pre-allocated buffer channels and buffer memory, and calculate the multiple of the data size and the buffer memory size; allocate the number of buffer channels according to the multiple and transmit the data.
[0015] Further, if the multiple is greater than or equal to 1, set the buffer channel number to 1, set the channel size to the size of the buffer memory, and transfer the data sequentially. For the data transferred in the last transfer, set the buffer channel number to 2, and the channel sizes to the size of the data transferred in the last transfer and the remaining memory size of the buffer respectively. If the multiple is less than 1, set the buffer channel number to 2, and the channel sizes to the size of the data transferred in the last transfer and the remaining memory size of the buffer respectively.
[0016] Further, based on the OpenAMP architecture, data transfer between the main core and the slave core is performed through virtio and RPMsg.
[0017] One or more of the above technical solutions in the embodiments of the present invention have at least one of the following technical effects: In the main-slave core architecture, the present invention adjusts the memory allocation layout of the slave core, adjusts the memory size of the slave core as needed, and performs cyclic allocation according to the different memory spaces occupied during model operation, enabling it to run DNN inference models of different sizes. For DNN models, if the model parameters are numerous and the model is large, then during operation, through the main-slave core memory recycling and allocation method, its dynamic memory allocation requirements during model operation are met; if the model is small, it can be directly run without reallocation.
[0018] In the main-slave core framework based on the OpenAMP architecture, the present invention adjusts the size of the buffer in the shared memory as needed. By reallocating the buffer channels at the software level, different channel sizes and channel numbers are assigned to data transfer tasks in different scenarios, so as to improve the data transfer bottleneck and speed without exceeding the buffer memory size. The present invention increases the upper limit of large data transfer and reduces the memory space consumed by short data transfer; at the same time, the channel number is adjusted accordingly, saving the memory size occupied by the buffer.
[0019] The additional aspects and advantages of the present invention will be partly given in the following description, partly will become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 It is a flowchart of the method provided by the present invention.
[0022] Figure 2 is the flowchart of dynamic allocation from the nuclear memory provided by the present invention. Specific embodiments
[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the scope of protection of the present invention. The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.
[0024] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., mean that the specific features, structures, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0025] The following combines Figure 1 and Figure 2 to further elaborate on the present invention and describe an optimization method for expanding the memory of the slave core in a master-slave core architecture of the present invention: In this embodiment, as Figure 1 shown, an optimization method for expanding the memory of the slave core in a master-slave core architecture is provided, including the following steps: S1: Obtain the model to be run on the slave core and calculate the total memory required during the model operation.
[0026] The models running on the slave core are generally lightweight DNN models. Models with extremely large model parameter scales are not within the scope of consideration of the present invention because these models are more suitable for running on the master core. For models where the memory required for each layer of the model exceeds the size of the slave core memory, methods such as model quantization can be used to reduce their memory requirements.
[0027] S2: Allocate slave core memory from the memory space according to the total memory and the size of the memory that can be allocated to the slave core.
[0028] For a deep learning model, even when running a lightweight DNN model on the coprocessor, its runtime memory is affected by factors such as the size of model parameters, input data size, batch size, model structure and number of layers, and data type, and still requires a certain amount of space. Therefore, it is necessary to reallocate and layout the coprocessor memory space. Based on the OpenAMP architecture, in this embodiment, it is first assumed that a shared memory space of 409MB is allocated for the coprocessor in the device tree file configuration, with the starting address being 0xb0100000. After deducting the memory occupied by the running of OpenAMP-related programs, the allocable memory for the coprocessor is 255MB. When performing DNN inference calculations, the coprocessor dynamically allocates memory for the program through the pvPortmalloc() function to ensure that the memory size meets the running requirements of the DNN network, enabling the coprocessor to successfully execute the lightweight DNN network.
[0029] When allocating coprocessor memory from the memory space, a dynamic allocation strategy is adopted, and the process is as Figure 2 shown. The specific process is as follows: Compare the size relationship between the total memory and the allocable memory of the coprocessor.
[0030] 1. When the total memory is less than or equal to the allocable memory of the coprocessor, directly allocate coprocessor memory from the memory space. In this case, when the model is running, the model dynamically allocates coprocessor memory: allocate memory for the model parameters, input data, temporary variables, etc. of each training round. Until the model inference calculation ends, the memory used during the model runtime will not exceed the allocable memory of the coprocessor, and there is no need to consider memory fragmentation and memory release and recycling issues.
[0031] 2. When the total memory is greater than the allocable memory of the coprocessor, calculate the memory required for each round of model runtime; according to the size of the round memory and the allocable memory of the coprocessor, allocate and recycle the coprocessor memory during each round of model inference calculation. In the master-slave coprocessor mode, the memory space required for the model parameters, input data of each round, and temporary variables running on the coprocessor is relatively fixed. Therefore, when the total memory required during model runtime is greater than the allocable memory of the coprocessor, this embodiment can recycle the coprocessor memory according to the size of the round memory required for each training round.
[0032] (1) When the round memory is less than or equal to the allocable memory of the coprocessor, allocate coprocessor memory by round, and allocate a memory space not less than the round memory as the coprocessor memory. After the completion of this round of model inference calculation (model training), this embodiment uses the vPortFree() function to release the allocated coprocessor memory back to the memory space. When the next round of model inference calculation starts, memory allocation is still performed in the above-mentioned memory space. Therefore, in this case, the allocation and release of coprocessor memory are performed in the same memory space for each round of model inference calculation.
[0033] In this embodiment, the size of the slave core memory is slightly larger than that of the round memory. The value of the round memory can be added with an adjustment value to obtain the value of the slave core memory. The adjustment value is set according to the actual situation. Specifically, the adjustment value can be a fixed value, such as 64B or 512B.
[0034] (2)When the round memory is greater than the allocable memory of the slave core, the slave core memory is allocated layer by layer. When the model performs inference calculation for each layer, memory allocation and recycling are carried out once. The layer memory required for each layer of the model during operation is calculated, and a memory space not less than the layer memory is allocated as the slave core memory. In this case, it is necessary to recycle and allocate the slave core memory in a loop to ensure that there is no memory shortage for each training batch.
[0035] In this embodiment, memory is allocated for each layer of the model: for each layer of the model, before the inference calculation starts, the required memory space capacity such as the input data size, output data size, and convolution kernel size of this layer is determined according to the following formula. The calculation formula for the layer memory is: Among them, is the layer memory, is the read memory size, is the write memory size, is the input data size, is the number of input data, is the output data size, is the number of output data, is the convolution kernel size.
[0036] After the inference calculation of each layer of the model is completed, the slave core memory occupied by the output data (i.e., the data required for the next layer of inference calculation) is saved, and the other slave core memory is released. In this way, sufficient memory space is ensured for each round of model training to ensure the smooth execution of each layer of the model and the transmission of input information to the next layer.
[0037] For the case where the layer memory is greater than or equal to the allocable memory of the slave core, it is considered that the model has a very large scale of model parameters and is not a model that can be run in this embodiment.
[0038] The allocation and recycling of the slave core memory are implemented based on the memory allocation algorithm of FreeRTOS. In this embodiment, under the Feiteng OpenAMP framework, the slave core runs the FreeRTOS operating system. FreeRTOS is an open-source real-time operating system (RTOS), mainly for embedded systems. When a program runs on FreeRTOS, memory management and allocation are crucial. Memory management involves effectively organizing the program's data and code in memory so that the operating system and hardware can execute the program efficiently. The memory layout usually includes the following main parts: code segment, data segment, heap, stack, and free storage area. FreeRTOS provides a dedicated dynamic memory allocation algorithm, using the pvPortMalloc() and pvPortFree() functions instead of the malloc() and free() functions in the standard C library, and implements the memory allocation algorithm by subdividing the array into smaller blocks. The array is defined by static declaration and marked by configTOTAL_HEAP_SIZE. Before actually allocating memory from the array, the application also appears to consume a large amount of RAM. FreeRTOS allocates memory through a fitting algorithm, supports memory allocation and release, and combines adjacent free memory blocks (merges) into a larger block, thus minimizing the risk of memory fragmentation.
[0039] S3: Transmit the data of the model to the slave core memory. Specifically, according to the data size, pipeline size, allocated buffer channel number, and buffer memory, transmit the data.
[0040] This embodiment is based on the OpenAMP architecture and uses virtio and RPMsg for data transfer between the master core and the slave core. During the data transfer process in this embodiment, the buffer settings are optimized. The purpose of the optimization is to make the size of the memory occupied by the buffer meet the data transfer requirements to ensure the transfer efficiency; the premise of the optimization is to avoid memory overflow caused by an overly large buffer memory area, thus overwriting the memory areas where the operating system or other components are located. According to experiments, for the same-sized data, the one-time transfer is faster than multiple transfers. This embodiment implements a large-scale data transfer scheme based on the minimum number of transfers, reducing the number of transfers and increasing the transfer speed. Of course, when the data size is close to the pre-set pipeline size, this embodiment uses the original data transfer scheme.
[0041] The specific optimization method is as follows: Pre-allocated buffer channels and buffer memory. The pre-set data transfer pipeline size of OpenAMP in this embodiment is 512B, the number of buffer channels is 512, and the size of the buffer memory is 256KB.
[0042] When the data size is less than or equal to the pipeline size, there is no need to re-allocate buffer channels, etc., and the data is directly transmitted.
[0043] When the data size is greater than the pipeline size, that is, when the data scale is large, the pre-allocated buffer channels and buffer memory are recycled, and the multiple of the data size and the buffer memory size is calculated; then, according to the multiple relationship, the number of buffer channels is allocated, and the data is transmitted. Specifically: If the multiple is greater than or equal to 1, that is, the data size ≥ buffer memory size, the number of buffer channels is set to 1, the channel size is set to the buffer memory size, and the data is transmitted in turn. For the data transmitted last time, the number of buffer channels is set to 2, and the channel sizes are the size of the data transmitted last time and the remaining memory size of the buffer respectively. When the data transmission size or the size of the data transmitted last time is equal to the buffer memory size, the remaining memory size of the buffer is 0, and the actual number of buffer channels at this time is 1.
[0044] If the multiple is less than 1, that is, the data size < buffer memory size, the number of buffer channels is set to 2, and the channel sizes are the size of the data transmitted last time and the remaining memory size of the buffer respectively.
[0045] When the data transmission size or the size of the data transmitted last time is less than the buffer memory size, the number of buffer channels is set to 2 to meet the minimum transmission times, and to ensure that when other tasks are processed simultaneously, the transmission times of other tasks are as small as possible.
[0046] S4: The slave core performs model inference calculation.
[0047] In this embodiment, the size of the buffer in the shared memory is adjusted as needed on the master-slave core framework based on the OpenAMP architecture. By adjusting the number of channels and the channel size, the performance requirements under different transmission requirements are solved. By allocating the channel size as needed, the upper limit of large data transmission is increased, and the memory space consumed by short data transmission is reduced. At the same time, the number of channels is adjusted accordingly, saving the memory size occupied by the buffer. This embodiment adjusts the slave core memory allocation layout under the master-slave core architecture and adjusts the slave core memory size as needed. For example, for a DNN model, if the model parameters are numerous and the model is large, during operation, through the master-slave core memory recycling and allocation method, its dynamic memory allocation requirements during model operation are met; if the model is small, it can be directly run without re-allocation.
[0048] In this embodiment, under the Feiteng OpenAMP framework, the heterogeneous multi-core development board is equipped with 2 FTC664 cores at 1.8 GHz and 2 FTC310 cores at 1.5 GHz. Among them, the FTC664 core numbered 3 serves as a slave core, and the 3 cores numbered 0-2 serve as master cores. The OpenAMP framework enables developers to efficiently exchange data and cooperate on tasks among these heterogeneous processors through standardized APIs and communication mechanisms. Its core components include: RemoteProc: Used for the management of remote processors, responsible for loading, starting, stopping, and monitoring the status of remote processors; RPMsg (Remote Processor Messaging): Provides a messaging mechanism to support communication between processors; Virtio: Used to standardize the interface of virtual devices, simplifying device virtualization and sharing.
[0049] Among these three core components, Virtio is a virtual device standard for device virtualization between virtual machines and host machines. RPMsg is a key component in the OpenAMP framework for implementing message passing in heterogeneous multi-processor systems. It is built on top of Virtio and RemoteProc and supports two-way communication between processors through a lightweight messaging protocol. OpenAMP, RPMsg, and Virtio together build a powerful framework to support communication and collaboration in heterogeneous multi-core processor systems. OpenAMP simplifies the management of heterogeneous processors, RPMsg provides an efficient messaging mechanism, and Virtio standardizes the device interface, promoting the efficient management and sharing of virtual devices.
[0050] In this embodiment, the end-to-end inference latency and segmented inference latency are tested in the serial operation of the multi-modal network (Baseline) in a normal environment and the master-slave core cooperation mode. In addition, the results of directly performing multi-core parallel acceleration on the multi-modal network without using the master-slave core mode are also tested, as shown in Table 1. The end-to-end latency for directly serially running the multi-modal network to complete one inference calculation is approximately 15 s, the end-to-end latency for completing one inference calculation in the master-slave core mode is approximately 4.5 s, and the end-to-end latency for direct parallel acceleration is approximately 3 s. This result proves that the master-slave core mode fully exploits the parallel computing performance of the heterogeneous multi-core CPU, achieving an acceleration ratio of 3.4 times, which is close to the effect of direct parallel acceleration.
[0051] According to the test results of the segmented inference delay, after the master-slave core transmission optimization and the lightweight FCN model design, the input-output delay between the master and slave cores is only 1 ms, which is almost negligible compared to the end-to-end delay. In terms of the execution efficiency of the slave core ResNet model, all 4 cores are directly used in parallel acceleration for execution, obtaining the shortest execution delay. The master and slave cores use 3 cores for execution, and the delay is higher than that of direct parallel acceleration. However, in the FCN task, the deeply optimized master-slave core mode shows a leading execution efficiency, proving the significant advantage of the master-slave core operation mode in running high-real-time tasks.
[0052] Table 1
[0053] In this embodiment, the model accuracy, recall rate, precision rate, and F1 score indicators of the model running in the normal environment and the master-slave core cooperation mode are further tested. The results are shown in Table 2. Due to the memory resource limitation of the slave core, the weights of the FCN model processed by the slave core in this embodiment are quantized. The multi-modal model running in the master-slave core mode still shows good detection effects, with an accuracy of 80.45%, a recall rate of 87.25%, a precision rate of 85.55%, and an F1 score of 80.11%. Compared with directly running the multi-modal myopia detection model, there is no significant loss of accuracy.
[0054] Table 2
[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An optimization method for slave core memory expansion under a master-slave core architecture, characterized in that Including: Obtain the model to be run on the slave core, and calculate the total memory required during the model running; Allocate slave core memory from the memory space according to the total memory and the size of the memory allocable by the slave core; Transfer the data of the model to the slave core memory; The slave core performs model inference calculation.
2. The method for optimizing the slave core memory expansion under the master-slave core architecture according to claim 1, wherein When the total memory is less than or equal to the memory allocable by the slave core, allocate slave core memory from the memory space; When the total memory is greater than the memory allocable by the slave core, calculate the round memory required during each round of model running; According to the round memory and the size of the memory allocable by the slave core, during each round of model inference calculation, allocate and recycle the slave core memory.
3. The method for optimizing the memory expansion of the slave core under the master-slave core architecture according to claim 2, wherein When the round memory is less than or equal to the memory allocable by the slave core, allocate a memory space not less than the round memory as the slave core memory; after each round of model inference calculation is completed, release the slave core memory back to the memory space, and when the next round of model inference calculation starts, still perform memory allocation in the above memory space; When the round memory is greater than the memory allocable by the slave core, calculate the layer memory required during each layer of the model running, and allocate a memory space not less than the layer memory as the slave core memory; After each layer of the model inference calculation is completed, save the slave core memory occupied by the output data, and release the other slave core memory.
4. The method for optimizing the slave core memory expansion under the master-slave core architecture according to claim 3, wherein The calculation formula of the layer memory is: Among them, is the layer memory, is the read memory size, is the write memory size, is the input data size, is the number of input data, is the output data size, is the number of output data, is the convolution kernel size.
5. The method for optimizing the memory expansion of the slave core under the master-slave core architecture according to claim 2, characterized in that The allocation and recycling of the slave core memory are implemented based on the memory allocation algorithm of FreeRTOS.
6. The method for optimizing the memory expansion of the slave core under the master-slave core architecture according to claim 1, wherein When transferring the data of the model to the slave core memory, allocate the number of buffer channels according to the size of the data and the pipe size, and then transfer the data.
7. A method for optimizing the expansion of slave core memory under a master-slave core architecture according to claim 6, characterized in that Pre-allocated buffer channels and buffer memory; When the data size is less than or equal to the pipe size, directly transfer the data; When the data size is greater than the pipe size, recycle the pre-allocated buffer channels and buffer memory, and calculate the multiple of the data size and the buffer memory size; allocate the number of buffer channels according to the multiple, and transfer the data.
8. The slave core memory expansion optimization method under the master-slave core architecture according to claim 7, wherein If the multiple is greater than or equal to 1, set the number of buffer channels to 1, set the channel size to the size of the buffer memory, transfer the data in sequence, and for the data transferred last time, set the number of buffer channels to 2, and the channel sizes are the size of the data transferred last time and the remaining memory size of the buffer respectively; If the multiple is less than 1, set the number of buffer channels to 2, and the channel sizes are the size of the data transferred last time and the remaining memory size of the buffer respectively.
9. The slave core memory expansion optimization method under the master-slave core architecture according to claim 1, wherein Based on the OpenAMP architecture, data transfer between the master core and the slave core is performed through virtio and RPMsg.
Citation Information
Patent Citations
Implementation method of dual-core shared network port, intelligent terminal and storage medium
CN111400214A
Multilayer neural network calculation method and device based on heterogeneous multi-core processor
CN115526302A
Multi-core system and dynamic module loading method thereof, medium and processor chip
CN117234607A
Memory management method for multi-core system, multi-core system, equipment and medium
CN119201482A
Scalable ARM-based heterogeneous multi-core system on a single chip to accelerate memory utilization and performance
DE202024102523U1