Business processing method and device, equipment and medium
By allocating multiple memory areas in the system on chip and adopting specific business processing methods, the problem of low software and hardware access efficiency in traditional SoC systems is solved, and high-performance software and hardware interaction and data processing are achieved.
Patent Information
- Application Number
- CN202311536855.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
In traditional SoC systems, a single SRAM is difficult to meet the high-performance access of software and hardware at the same time, and the software and hardware interaction access is inefficient, especially when data synchronization operations are frequent.
A business processing method is adopted by allocating multiple memory areas in the on-chip system, including the first memory located on the central processing unit side, the second memory between the central processing unit and the total work engine, and the third memory on the overall work engine side. This method acquires the target service through the central processor, converts it into a hardware-identifiable data structure, and stores it in a cache. The total work engine processes the service based on the third memory, and stores the result in the second memory, and finally stores the result in the first memory by the central processor.
It realizes high-performance access to software and hardware, improves the efficiency of software and hardware interaction, and improves overall performance by reducing the length of access paths and avoiding frequent flush operations.
Smart Images

Figure CN120020723A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and particularly to a service processing method, apparatus, device, and medium. Background Art
[0002] Currently, in a traditional SoC (System on Chip) system, each module of the system allocates memory resources from a single SRAM (Static Random-Access Memory). Therefore, business data with different characteristics such as system image code, pure software access data, software-hardware interaction access data, and pure hardware access data are stored in the on-chip SRAM. Specifically, different business data has different processing methods: 1. For pure software access data, by enabling the Cache to improve the CPU's access performance to the SRAM, the higher the Cache (high-speed cache memory) hit rate, the higher the overall performance of the software access data; 2. For software-hardware interaction access data, before the software hands over the data to the hardware for access, a flush operation (flushing the cache, synchronizing the data in the cache to the memory) is performed to ensure that the data in the Cache is written to the memory, so as to ensure that the hardware reads the correct data. After the hardware finishes writing, the software performs an invalidate operation (emptying the cache), and then reads the memory data to ensure that the software correctly reads the data written to the memory by the hardware; 3. For pure hardware access data, after the software hands it over to the hardware, the hardware accesses the memory through the DMA (Direct Memory Access) method.
[0003] However, when accessing from the CPU and the hardware IP (Intellectual Property core), the SRAM may be closer to the CPU (Central Processing Unit) side or closer to the hardware IP side. Since the total access path length of the CPU and the hardware IP is longer, the access rate is lower. Therefore, a single SRAM device is difficult to simultaneously meet the high-performance access requirements of software and hardware. In addition, for software-hardware interaction access data, the software performs a flush before the hardware reads and an invalidate after the hardware writes. When the software-hardware interaction is relatively frequent, the flush operation synchronizes the data from the cache to the memory, which generally takes a long time, especially when synchronizing a large block of data at once, the latency is even higher, thus having a greater impact on the overall performance.
[0004] In summary, how to achieve high-performance access for software and hardware and improve the efficiency of software-hardware interaction is an urgent problem to be solved currently. Summary of the Invention
[0005] In view of this, an object of the present invention is to provide a service processing method, apparatus, device and medium, which can achieve high-performance access to software and hardware and improve the efficiency of software and hardware interaction. The specific solution is as follows:
[0006] In a first aspect, the present application discloses a service processing method applied to a system on a chip. The on-chip static random access memory of the system on a chip includes a first memory on the side where the central processing unit is located, a second memory at an intermediate position between the central processing unit and the total working engine, and a third memory on the side where the total working engine is located. The method includes:
[0007] Through the central processing unit, obtain a target service, determine target resources required to process the target service based on the memory resources of the first memory and according to the target service, and store the target resources in the first memory;
[0008] Through the central processing unit, extract the target resources from the first memory, convert the target resources into converted resources in a hardware-recognizable data structure, and store the converted resources in a cache;
[0009] Through the total working engine, obtain the converted resources from the cache, process the target service based on the converted resources according to the memory resources of the third memory, and store the processing result in the second memory;
[0010] Through the central processing unit, store the processing result from the second memory to the first memory.
[0011] Optionally, the step of "through the total working engine, obtain the converted resources from the cache" includes:
[0012] Through the total working engine, directly obtain the converted resources from the cache according to the accelerator coherence port.
[0013] Optionally, the on-chip static random access memory of the system on a chip includes a first memory connected to the on-chip interconnect subsystem of the central processing unit, a second memory connected to the on-chip interconnect subsystem for software and hardware interaction, and a third memory connected to the on-chip interconnect subsystem of the hardware; wherein, the on-chip interconnect subsystem of the central processing unit is connected to the central processing unit, and the on-chip interconnect subsystem of the hardware is connected to the total working engine;
[0014] Correspondingly, the step of "through the central processing unit, obtain a target service, determine target resources required to process the target service based on the memory resources of the first memory and according to the target service, and store the target resources in the first memory" includes:
[0015] Through the central processing unit, obtain the target service, determine the target resources required to process the target service based on the memory resources of the first memory, and store the target resources in the first memory through the on-chip interconnect subsystem of the central processing unit;
[0016] Correspondingly, the process of the central processing unit extracting the target resources from the first memory, converting the target resources into a converted resource in a hardware-recognizable data structure, and storing the converted resource in the cache includes:
[0017] Through the central processing unit, extract the target resources from the first memory through the on-chip interconnect subsystem of the central processing unit, convert the target resources into a converted resource in a hardware-recognizable data structure, and store the converted resource in the cache;
[0018] Correspondingly, the process of the total work engine obtaining the converted resource from the cache, processing the target service based on the converted resource according to the memory resources of the third memory, and storing the processing result in the second memory includes:
[0019] Through the total work engine, obtain the converted resource from the cache, process the target service based on the converted resource according to the memory resources of the third memory, and store the processing result in the second memory through the on-chip hardware interconnect subsystem and the on-chip software-hardware interaction interconnect subsystem;
[0020] Correspondingly, the process of the central processing unit storing the processing result from the second memory to the first memory includes:
[0021] Through the central processing unit, store the processing result from the second memory to the first memory through the on-chip software-hardware interaction interconnect subsystem and the on-chip interconnect subsystem of the central processing unit.
[0022] Optionally, the target resources include an engine list of several work engines stored in the total work engine in the running order for processing the target service;
[0023] Correspondingly, the process of the total work engine obtaining the converted resource from the cache, processing the target service based on the converted resource according to the memory resources of the third memory includes:
[0024] Through any target engine among the several work engines, obtain the converted resource from the cache;
[0025] Through the target engine, allocate the memory resources of the third memory required by the several working engines based on the converted resources, and allocate tasks to the several working engines according to the running order, so that the several working engines process the target service according to the converted resources.
[0026] Optionally, the registers corresponding to the target engine only include a control register and a status register.
[0027] Optionally, the target resources further include input data and input data size, or, the input data, input data size, and output data size;
[0028] Correspondingly, the service processing method further includes:
[0029] When the input data and the input data size in the target resources are empty, the step of storing the target resources in the first memory and subsequent steps are prohibited from being triggered.
[0030] Optionally, the converted resources obtained by converting the target resources into a hardware-recognizable data structure include:
[0031] The converted resources obtained by converting the target resources into a control block data structure.
[0032] In a second aspect, the present application discloses a service processing device, which is applied to a system on a chip. The on-chip static random access memory of the system on a chip includes a first memory on the side where the central processing unit is located, a second memory at an intermediate position between the central processing unit and the total working engines, and a third memory on the side where the total working engines are located. The device includes:
[0033] A first resource storage module, configured to obtain a target service through the central processing unit, determine target resources required to process the target service based on the memory resources of the first memory and according to the target service, and store the target resources in the first memory;
[0034] A second resource storage module, configured to extract the target resources from the first memory through the central processing unit, convert the target resources into a converted resource of a hardware-recognizable data structure, and store the converted resource in a cache;
[0035] A service processing module, configured to obtain the converted resource from the cache through the total working engines, process the target service based on the memory resources of the third memory according to the converted resource, and store the processing result in the second memory;
[0036] A processing result transfer module, configured to store the processing result from the second memory to the first memory through the central processing unit.
[0037] In a third aspect, the present application discloses an electronic device, including:
[0038] A memory for storing a computer program;
[0039] A processor for executing the computer program to implement the foregoing disclosed service processing method.
[0040] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the foregoing disclosed service processing method is implemented.
[0041] The beneficial effects of the present application are as follows: Through the central processing unit, the target service is obtained, the target resources required for processing the target service are determined based on the memory resources of the first memory and according to the target service, and the target resources are stored in the first memory; through the central processing unit, the target resources are extracted from the first memory, the target resources are converted into a converted resource in a data structure recognizable by hardware, and the converted resource is stored in the cache; through the total working engine, the converted resource is obtained from the cache, the target service is processed based on the memory resources of the third memory according to the converted resource, and the processing result is stored in the second memory; through the central processing unit, the processing result is stored from the second memory to the first memory. It can be seen that the central processing unit of the present application processes services based on the first memory resources, the total working engine processes services based on the third memory resources, the first memory is located on the side where the central processing unit is located, and the third memory is located on the side where the total working engine is located. Therefore, the access path is short and the rate is high, which can simultaneously meet the high-performance access of software and hardware; in addition, during the process of software and hardware interacting to access data, the data is directly stored in the cache, and the total working engine retrieves the data from the cache instead of from the second memory, without the flush operation. Therefore, the transmission speed is increased, and the efficiency of software and hardware interaction is further improved. Description of the Drawings
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0043] Figure 1 It is a flowchart of a service processing method disclosed in the present application;
[0044] Figure 2 It is a flowchart of a specific service processing method disclosed in the present application;
[0045] Figure 3 It is a flowchart of a specific service processing method disclosed in this application;
[0046] Figure 4 It is a schematic diagram of a system-on-chip disclosed in this application;
[0047] Figure 5 It is a schematic diagram of the structure of a service processing device disclosed in this application;
[0048] Figure 6 It is a structural diagram of an electronic device disclosed in this application. Specific embodiments
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0050] When a single SRAM is accessed from the CPU and the hardware IP (Intellectual Property core), the SRAM may be close to the CPU (Central Processing Unit) side or close to the hardware IP side. Since the total path length of the access by the CPU and the hardware IP is longer, the access rate is lower. Therefore, it is difficult for a single SRAM device to simultaneously meet the high-performance access requirements of software and hardware. In addition, for the data accessed by software and hardware interaction, the software executes flush before the hardware reads and executes invalidate after the hardware writes. When the software-hardware interaction is relatively frequent, the flush operation synchronizes the data from the cache to the memory, and this operation generally takes a long time, especially when synchronizing a large block of data at one time, the latency is higher, thus having a greater impact on the overall performance.
[0051] Therefore, the embodiments of this application propose a service processing solution that can achieve high-performance access for software and hardware and improve the efficiency of software-hardware interaction.
[0052] The embodiments of this application disclose a service processing method. Refer to Figure 1 As shown, it is applied to a system-on-chip. The on-chip static random access memory of the system-on-chip includes a first memory on the side where the central processor is located, a second memory at the middle position between the central processor and the total work engine, and a third memory on the side where the total work engine is located. The method includes:
[0053] Step S11: Through the central processing unit, obtain a target service, determine target resources required for processing the target service based on the memory resources of the first memory and according to the target service, and store the target resources in the first memory.
[0054] In this embodiment, memory access is divided into three categories: 1) pure software access; 2) software-hardware interaction access; 3) pure hardware access; memory devices are divided into three categories: 1) close to the CPU side; 2) in the middle position; 3) close to the DMA side (which can also be considered as close to the total working engine side, and the working engine exchanges data with the memory through direct memory access); the devices close to the CPU side are mainly used by the pure software memory access module; the devices close to the DMA side are mainly used by the pure hardware memory access module; the memory devices in the middle position are used by the software-hardware interaction management module. The side close to the DMA side is the side where the total working engine is located.
[0055] In this embodiment, the system-on-chip is a system-level chip, which is a product and an integrated circuit with a dedicated target, including a complete system and all the content of the embedded software. As long as the static random access memory remains powered on, the data stored in it can be constantly maintained.
[0056] In this embodiment, the first memory is the memory located on the side where the central processing unit is located, close to the central processing unit. Therefore, when the central processing unit accesses the first memory, the path is shorter. Therefore, the process of determining the target resources required for processing the target service based on the memory resources of the first memory and according to the target service is faster. In addition, the process of storing the target resources in the first memory is also faster.
[0057] In this embodiment, the target resources include an engine list of several working engines stored in the total working engine in the running order for processing the target service; among them, the engine list indicates both the working engines that can participate in processing the target service and the running order of the working engines, and the ID (Identity document, identity identification number) of the working engine is stored in the engine list.
[0058] In this embodiment, the target resources further include input data and the size of the input data, or, the input data, the size of the input data, and the size of the output data; among them, if the engine itself stipulates the format of the output data, the target resources need to include the size of the output data and other output data formats. If the engine itself does not stipulate, the target resources do not need to include the size of the output data and other output data formats.
[0059] Step S12: Through the central processing unit, extract the target resources from the first memory, convert the target resources into converted resources in a data structure recognizable by hardware, and store the converted resources in the cache.
[0060] In this embodiment, the software and hardware interaction defines a unified data structure format (the control block in this application), and uses shared memory for interaction. The converted resource obtained by converting the target resource into a data structure recognizable by the hardware includes: the converted resource obtained by converting the target resource into a control block data structure.
[0061] Step S13: Through the general working engine, obtain the converted resource from the cache, process the target service according to the converted resource based on the memory resource of the third memory, and store the processing result in the second memory.
[0062] In this embodiment, the third memory is located on the side where the general working engine is located and is close to the general working engine. Therefore, when the general working engine accesses the third memory, the path is short. Therefore, the step of processing the target service according to the converted resource based on the memory resource of the third memory and storing the processing result in the second memory is relatively fast.
[0063] It should be noted that the several working engines process the target service according to the converted resource and store the processing result in the second memory, rather than in the third memory, because putting it in the third memory and then having the central controller fetch the processing result is not as fast as directly putting it in the second memory. However, if necessary, it can also be put in the third memory and then fetched by the central processor.
[0064] In this embodiment, the converted resource can be directly obtained from the cache (Cache), without the need to use the flush operation to synchronize the data from the cache to the memory (this operation generally takes a long time), and then read from the memory, saving time, improving efficiency, and enhancing the software-hardware interaction performance.
[0065] In this embodiment, the step of obtaining the converted resource from the cache through the general working engine includes: obtaining the converted resource directly from the cache through the general working engine according to the accelerator coherency port. The accelerator coherency port, that is, ACP (Accelerator Coherency Port), is the slave interface of the DSU (Diode Supply Unit) module, and its interface protocol specification is a subset of the ACE-lite protocol. The ACP slave interface allows the external master interface connected to this interface to access the cachable memory space through the main memory interface (master interface) of the DSU.
[0066] In this embodiment, the memory for software and hardware interaction is managed by ACP, which avoids frequent cache flushing operations and improves the software-hardware interaction performance at the same time.
[0067] Step S14: Store the processing result from the second memory to the first memory through the central processing unit.
[0068] In summary, the software running of this application can access data with the best performance. In the system image, the code segment, data segment, bss segment, etc. are all stored in the first memory. Among all the memory devices accessed by the central processing unit, the first memory is the fastest, so the software running can achieve the best performance. The hardware running can access data with the best performance. Multiple hardware IPs jointly complete a service process, and the data they access is stored in the third memory. Among all the memory devices accessed by the hardware, the third memory is the fastest, so the hardware running can achieve the best performance.
[0069] The beneficial effects of this application are as follows: Through the central processing unit, this application obtains the target service, determines the target resources required to process the target service based on the memory resources of the first memory, and stores the target resources in the first memory; through the central processing unit, extracts the target resources from the first memory, converts the target resources into a converted resource in a data structure recognizable by the hardware, and stores the converted resource in the cache; through the total work engine, obtains the converted resource from the cache, processes the target service based on the converted resource according to the memory resources of the third memory, and stores the processing result in the second memory; stores the processing result from the second memory to the first memory through the central processing unit. It can be seen that the central processing unit of this application processes services based on the first memory resources, and the total work engine processes services based on the third memory resources. Moreover, the first memory is on the side where the central processing unit is located, and the third memory is on the side where the total work engine is located. Therefore, the access path is short and the rate is high, which can simultaneously meet the high-performance access of software and hardware; in addition, during the process of software and hardware interacting to access data, the data is directly stored in the cache, and the total work engine retrieves data from the cache instead of from the second memory, without the flush operation, so the transmission speed is increased, and the efficiency of software and hardware interaction is further improved.
[0070] This embodiment of the application discloses a specific service processing method. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution. Applied to a system-on-chip, the on-chip static random access memory of the system-on-chip includes a first memory on the side where the central processing unit is located, a second memory at the middle position between the central processing unit and the total work engine, and a third memory on the side where the total work engine is located. See Figure 2 shown, specifically including:
[0071] Step S21: Through the central processing unit, obtain a target service, determine target resources required for processing the target service based on the memory resources of the first memory, and store the target resources in the first memory.
[0072] Step S22: Through the central processing unit, extract the target resources from the first memory, convert the target resources into converted resources in a hardware-recognizable data structure, and store the converted resources in the cache.
[0073] In this embodiment, the converting the target resources into converted resources in a hardware-recognizable data structure includes: converting the target resources into converted resources in a control block data structure.
[0074] Step S23: Through any one of the several working engines, obtain the converted resources from the cache.
[0075] In this embodiment, only one of the several working engines can interact with software through a direct memory access device, that is, among a group of hardware IPs, only one IP is reserved for interaction with software. The intellectual property module is a pre-designed circuit function module in an ASIC (Application Specific Integrated Circuit) or an FPGA (Field-Programmable Gate Array).
[0076] Step S24: Through the target engine, allocate the memory resources of the third memory required by the several working engines based on the converted resources and allocate tasks to the several working engines based on the running order, so that the several working engines process the target service according to the converted resources and store the processing results in the second memory.
[0077] In this embodiment, the registers corresponding to the target engine only include control registers and status registers, so as to maximize the reduction of register configuration and reading, reduce time consumption, and improve speed and overall performance.
[0078] In this embodiment, the target resources further include input data and input data size, or, the input data, input data size, and output data size; correspondingly, the service processing method further includes: when the input data and the input data size in the target resources are empty, prohibit triggering the step of storing the target resources in the first memory and subsequent steps.
[0079] It should be noted that when a group of hardware IP engines jointly achieve a business goal, this group of IP engines is divided into two categories: a scheduling engine and a working engine (the scheduling engine is also the so-called target engine). The scheduling engine is responsible for interacting with the software, disassembling and distributing the business goal to each working engine, managing the resources required for the engine work, and coordinating each engine to jointly achieve the business goal.
[0080] It should be noted that for the scheduling engine, the registers it provides should be as few as possible, and only the necessary control registers and status registers need to be retained. For example, the control registers provide a start register, a stop register, and a control block address to respectively control the start, stop, and physical memory address of the control block of the scheduling engine. The status registers provide a completion status and an error status, respectively indicating whether the business goal is completed, whether there is an error, and the error type.
[0081] It should be noted that the control block is used to tell the scheduling engine which working engines will participate in the operation and their operation order. The control block contains the number of participating working engines, a list of working engine IDs stored in sequence, the input data size, the input data, the output data size, and the output data. The input data size and the output data size can be empty, indicating that there is no data input or output for this business goal.
[0082] Step S25: Store the processing result from the second memory to the first memory through the central processing unit.
[0083] The beneficial effects of this application are as follows: Through the central processing unit, this application obtains the target service, determines the target resources required to process the target service based on the memory resources of the first memory and according to the target service, and stores the target resources in the first memory; through the central processing unit, extracts the target resources from the first memory, converts the target resources into converted resources in a hardware-recognizable data structure, and stores the converted resources in the cache; through any one of the several working engines, the target engine, obtains the converted resources from the cache; through the target engine, allocates the memory resources of the third memory required by the several working engines based on the converted resources and allocates tasks to the several working engines based on the running order, so that the several working engines process the target service according to the converted resources and store the processing results in the second memory; through the central processing unit, stores the processing results from the second memory in the first memory. It can be seen that the central processing unit of this application processes services based on the first memory resources, the total working engine processes services based on the third memory resources, and the first memory is located on the side where the central processing unit is located, and the third memory is located on the side where the total working engine is located. Therefore, the access path is short and the rate is high, which can simultaneously meet the high-performance access of software and hardware; in addition, during the process of software and hardware interacting to access data, the data is directly stored in the cache, and the total working engine fetches data from the cache instead of from the second memory, without the flush operation, so the transmission speed is increased, and the efficiency of software and hardware interaction is further improved; in addition, this application has only one target engine interacting with the software and controlling the operation of the working engine, and the registers of the target engine only include control registers and status registers, which reduces the access of other engines to the cache and the second memory, and also reduces the register access, improving the interaction efficiency.
[0084] The embodiment of this application discloses a specific service processing method. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution. It is applied to a system on a chip. The on-chip static random access memory of the system on a chip includes a first memory connected to the on-chip interconnection subsystem of the central processing unit, a second memory connected to the on-chip interconnection subsystem of software and hardware interaction, and a third memory connected to the on-chip interconnection subsystem of hardware; wherein, the on-chip interconnection subsystem of the central processing unit is connected to the central processing unit, and the on-chip interconnection subsystem of hardware is connected to the total working engine. See Figure 3 as shown, and specifically includes:
[0085] Step S31: Through the central processing unit, obtain the target service, determine the target resources required to process the target service based on the memory resources of the first memory and according to the target service, and store the target resources in the first memory through the on-chip interconnection subsystem of the central processing unit.
[0086] Step S32: Through the central processing unit, extract the target resource from the first memory through the on-chip interconnection subsystem of the central processing unit, convert the target resource into a converted resource in a hardware-recognizable data structure, and store the converted resource in the cache.
[0087] Step S33: Through the general working engine, obtain the converted resource from the cache, process the target service according to the converted resource based on the memory resources of the third memory, and store the processing result in the second memory through the on-chip interconnection subsystem of the hardware and the on-chip interconnection subsystem of the software and hardware interaction.
[0088] Step S34: Through the central processing unit, store the processing result from the second memory to the first memory through the on-chip interconnection subsystem of the software and hardware interaction and the on-chip interconnection subsystem of the central processing unit.
[0089] In this embodiment, as shown in Figure 4 is a schematic diagram of a system-on-chip; the on-chip interconnection subsystem of the central processing unit is also called the CPUNoC subsystem, the on-chip interconnection subsystem of software and hardware interaction is also called the DataNoC subsystem, and the on-chip interconnection subsystem of the hardware is also called the DeviceNoC subsystem; according to Figure 4 the structure, the paths for the central processing unit to access different memories are as follows: 1. The central processing unit accesses the first memory: central processing unit - on-chip interconnection subsystem of the central processing unit - first memory; 2. The central processing unit accesses the second memory: central processing unit - on-chip interconnection subsystem of the central processing unit - on-chip interconnection subsystem of software and hardware interaction - second memory; 3. The central processing unit accesses the third memory: central processing unit - on-chip interconnection subsystem of the central processing unit - on-chip interconnection subsystem of the hardware - third memory; the paths for the working engine to access different memories are as follows: 1. The working engine accesses the third memory: working engine - on-chip interconnection subsystem of the hardware - third memory; 2. The working engine accesses the second memory: working engine - on-chip interconnection subsystem of the hardware - on-chip interconnection subsystem of software and hardware interaction - second memory; 3. The working engine accesses the first memory: working engine - on-chip interconnection subsystem of the hardware - on-chip interconnection subsystem of the central processing unit - first memory.
[0090] It should be noted that from the perspective of the central processing unit, the memory access rate ranking is first memory > second memory > third memory; from the perspective of the working engine, the memory access rate ranking is third memory > second memory > first memory.
[0091] It should be noted that the bandwidth and frequency of the on-chip interconnection subsystem of the central processing unit and the on-chip interconnection subsystem of software and hardware interaction are generally higher than those of the on-chip interconnection subsystem of the hardware.
[0092] In this embodiment, the above is divided into three memories, and there are distinctions between software, hardware, and software-hardware interaction. Therefore, the service modules can be divided according to the memory access characteristics into: 1. Pure software access module; 2. Software-hardware interaction management module; 3. Pure hardware access module. And the first memory is allocated to the pure software access module, the second memory is allocated to the software-hardware interaction management module, and the third memory is allocated to the pure hardware access module to implement the steps of the above embodiment.
[0093] It should be noted that the pure software access module manages the first memory, such as code, global variables, run stacks, service data, etc. in the system. All data without direct hardware access is allocated to this type of module for management. At the same time, caching is enabled to maximize the performance of pure software memory access. The software-hardware interaction management module manages the second memory. The second memory has two access methods: 1) Software module writes and hardware module reads; 2) Hardware module writes and software module reads. The former is used for the software module to transfer input data to the hardware module, and the latter is used for the software to obtain the operation result from the hardware module. The pure hardware memory access module manages the third memory. During initialization, the third memory address is configured to the target hardware engine corresponding to the pure hardware memory access module. When it works, it decides how to use the third memory by itself. It can use it by itself or allocate it to other hardware engines for use. But ultimately, the reading and writing of the third memory are both completed by the hardware, and the software does not directly read and write the third memory.
[0094] In a specific embodiment, the above steps can be specifically represented by the following process:
[0095] First, the pure software access module prepares the hardware operation parameters (i.e., target resources) through logical calculation, and the memory required during its operation process is allocated from the first memory.
[0096] Second, the software-hardware interaction management module converts and assembles the calculation result of the previous step into a data structure (control block structure) that can be recognized by the hardware according to the control block structure defined by the software-hardware interaction interface.
[0097] Third, the pure hardware access module obtains and parses the control block content, allocates the memory resources required by the hardware according to the control block information, and coordinates the work of other hardware IPs. The access to the control block is accelerated by ACP and directly read from the cache. The memory is allocated from the third memory. After the execution is completed, the service result is written into the second memory.
[0098] Fourth, the pure software access module reads the service result from the second memory.
[0099] The beneficial effects of this application are as follows: Through the central processing unit, this application acquires a target service, determines target resources required for processing the target service based on the memory resources of the first memory and according to the target service, and stores the target resources in the first memory through the on-chip interconnection subsystem of the central processing unit; through the central processing unit, extracts the target resources from the first memory through the on-chip interconnection subsystem of the central processing unit, converts the target resources into converted resources in a hardware-recognizable data structure, and stores the converted resources in the cache; through the total work engine, obtains the converted resources from the cache, processes the target service based on the converted resources according to the memory resources of the third memory, and stores the processing result in the second memory through the on-chip interconnection subsystem of the hardware and the on-chip interconnection subsystem of the software and hardware interaction; through the central processing unit, stores the processing result from the second memory to the first memory through the on-chip interconnection subsystem of the software and hardware interaction and the on-chip interconnection subsystem of the central processing unit. Thus, it can be seen that the central processing unit of this application processes services based on the first memory resources, the total work engine processes services based on the third memory resources, and the first memory is located on the side where the central processing unit is located, and the third memory is located on the side where the total work engine is located. Therefore, the access path is short and the rate is high, which can simultaneously meet the high-performance access requirements of software and hardware; in addition, during the process of software and hardware interaction to access data, the data is directly stored in the cache, and the total work engine retrieves the data from the cache instead of from the second memory, without the need for a flush operation. Therefore, the transmission speed is increased, and the efficiency of software and hardware interaction is further improved.
[0100] Correspondingly, an embodiment of this application also discloses a service processing device, which is applied to a system on a chip. The on-chip static random access memory of the system on a chip includes a first memory located on the side where the central processing unit is located, a second memory located at the middle position between the central processing unit and the total work engine, and a third memory located on the side where the total work engine is located. Refer to Figure 5 as shown, this device includes:
[0101] A first resource storage module 11, configured to acquire a target service through the central processing unit, determine target resources required for processing the target service based on the memory resources of the first memory and according to the target service, and store the target resources in the first memory;
[0102] A second resource storage module 12, configured to extract the target resources from the first memory through the central processing unit, convert the target resources into converted resources in a hardware-recognizable data structure, and store the converted resources in the cache;
[0103] The service processing module 13 is configured to obtain the converted resource from the cache through the total working engine, process the target service according to the converted resource based on the memory resource of the third memory, and store the processing result in the second memory;
[0104] The processing result transfer and storage module 14 is configured to store the processing result from the second memory to the first memory through the central processing unit.
[0105] Among them, for the more specific working processes of the above-mentioned various modules, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.
[0106] It can be seen that in this application, the central processing unit processes services based on the first memory resource, and the total working engine processes services based on the third memory resource. Moreover, the first memory is located on the side where the central processing unit is located, and the third memory is located on the side where the total working engine is located. Therefore, the access path is short and the rate is high, which can simultaneously meet the high-performance access requirements of both software and hardware. In addition, during the process of software and hardware interacting to access data, the data is directly stored in the cache, and the total working engine fetches data from the cache instead of from the second memory, without the need for a flush operation. Therefore, the transmission speed is increased, and the efficiency of software and hardware interaction is further improved.
[0107] Furthermore, the embodiment of this application also provides an electronic device. Figure 6 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be regarded as any limitation on the scope of use of this application.
[0108] Figure 6 This is a schematic structural diagram of an electronic device 20 provided by the embodiment of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a display screen 23, an input / output interface 24, a communication interface 25, a power supply 26, and a communication bus 27. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the service processing method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0109] In this embodiment, the power supply 26 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 25 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of this application, and specific limitations are not imposed herein; the input / output interface 24 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and specific limitations are not imposed herein.
[0110] In addition, as a carrier for storing resources, the memory 22 can be a read-only memory, a random access memory, a magnetic disk, an optical disk, etc. The resources stored thereon can include a computer program 221, and the storage method can be transient storage or permanent storage. Among them, in addition to the computer program that can be used to complete the business processing method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 221 can further include a computer program that can be used to complete other specific tasks.
[0111] Furthermore, an embodiment of the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the foregoing disclosed business processing method is implemented.
[0112] For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0113] The embodiments in this application are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts between the embodiments, reference can be made to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference can be made to the method part for related parts.
[0114] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0115] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0116] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0117] The above has introduced in detail a service processing method, apparatus, device, and storage medium provided by the present application. Specific examples are used in this text to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A business processing method, characterized in that: Applied to a system on chip, the on-chip static random access memory of the system on chip includes a first memory located on the side where the central processing unit is located, a second memory located in the middle between the central processing unit and the general working engine, and a third memory located on the side where the general working engine is located, including: Acquire a target business through the central processor, determine target resources required for processing the target business based on the memory resources of the first memory and according to the target business, and store the target resources in the first memory; Extracting the target resource from the first memory through the central processor, converting the target resource into a converted resource having a data structure recognizable by hardware, and storing the converted resource into a cache; Obtaining the converted resource from the cache through the general work engine, processing the target business according to the converted resource based on the memory resources of the third memory, and storing the processing result in the second memory; The processing result is stored from the second memory to the first memory through the central processing unit.
2. The service processing method according to claim 1, characterized in that: The obtaining the converted resource from the cache through the general working engine includes: The converted resource is directly obtained from the cache through the general work engine according to the accelerator consistency port.
3. The service processing method according to claim 1, characterized in that: The on-chip static random access memory of the on-chip system includes a first memory connected to the on-chip interconnection subsystem of the central processing unit, a second memory connected to the on-chip interconnection subsystem of the software and hardware interaction, and a third memory connected to the on-chip interconnection subsystem of the hardware; wherein the on-chip interconnection subsystem of the central processing unit is connected to the central processing unit, and the on-chip interconnection subsystem of the hardware is connected to the general working engine; Accordingly, the target business is acquired through the central processor, the target resources required for processing the target business are determined based on the memory resources of the first memory and according to the target business, and the target resources are stored in the first memory, including: Obtaining a target business through the central processor, determining target resources required for processing the target business based on the memory resources of the first memory and according to the target business, and storing the target resources in the first memory through the on-chip interconnect subsystem of the central processor; Accordingly, extracting the target resource from the first memory through the central processor, converting the target resource into a converted resource having a data structure recognizable by hardware, and storing the converted resource in a cache includes: By means of the central processor, the target resource is extracted from the first memory via the central processor on-chip interconnect subsystem, the target resource is converted into a converted resource having a data structure recognizable by hardware, and the converted resource is stored in a cache; Correspondingly, the method of obtaining the converted resource from the cache through the general working engine, processing the target business according to the converted resource based on the memory resource of the third memory, and storing the processing result in the second memory includes: Obtain the converted resources from the cache through the general work engine, process the target business according to the converted resources based on the memory resources of the third memory, and store the processing results in the second memory through the hardware on-chip interconnect subsystem and the software-hardware interactive on-chip interconnect subsystem; Accordingly, storing the processing result from the second memory to the first memory by the central processor includes: The processing result is stored from the second memory to the first memory through the central processing unit, the software-hardware interactive on-chip interconnection subsystem and the central processing unit on-chip interconnection subsystem.
4. The service processing method according to claim 1, characterized in that: The target resource includes an engine list of a plurality of working engines for processing the target business stored in the general working engine in the running order; Correspondingly, obtaining the converted resource from the cache through the general work engine, and processing the target business according to the converted resource based on the memory resource of the third memory includes: Obtaining the converted resource from the cache through any target engine among the plurality of working engines; Through the target engine, the memory resources of the third memory required by the several working engines are allocated based on the converted resources and tasks are allocated to the several working engines based on the running order, so that the several working engines can process the target business according to the converted resources.
5. The service processing method according to claim 4, characterized in that: The registers corresponding to the target engine only include a control register and a status register.
6. The service processing method according to claim 4, characterized in that: The target resource further includes input data and input data size, or the input data, input data size and output data size; Accordingly, the business processing method further includes: When the input data and the input data size in the target resource are empty, triggering the step of storing the target resource in the first memory and subsequent steps is prohibited.
7. The service processing method according to any one of claims 1 to 6, characterized in that: The step of converting the target resource into a converted resource with a hardware-recognizable data structure comprises: The target resource is converted into a converted resource of a control block data structure.
8. A service processing device, characterized in that: Applied to a system on chip, the on-chip static random access memory of the system on chip includes a first memory located on the side where the central processing unit is located, a second memory located in the middle between the central processing unit and the general working engine, and a third memory located on the side where the general working engine is located, including: A first resource storage module, configured to obtain a target service through the central processor, determine a target resource required for processing the target service based on the memory resources of the first memory and according to the target service, and store the target resource in the first memory; A second resource storage module is used to extract the target resource from the first memory through the central processor, convert the target resource into a converted resource with a data structure recognizable by hardware, and store the converted resource in a cache; A business processing module, used to obtain the converted resources from the cache through the general work engine, process the target business according to the converted resources based on the memory resources of the third memory, and store the processing results in the second memory; A processing result transfer module is used to store the processing result from the second memory to the first memory through the central processing unit.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the service processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store a computer program; wherein, when the computer program is executed by a processor, the business processing method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
System and method for multi-level switch firmware combination memory resource allocation
CN120238511A