A multi-stage pipelined multiway data operation and access control system

By optimizing the multi-stage pipelined processor group and bus access control, the problems of complex bus interconnection and low efficiency in multi-core chips in multi-channel multi-stage pipelined parallel processing are solved, achieving high efficiency, low cost access parallelism and dynamic adaptability.

CN115437994BActive Publication Date: 2026-02-24母国标
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211070128.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-02
Publication Date
2026-02-24
Estimated Expiration
2042-09-02

AI Technical Summary

Technical Problem

Existing multi-core chips suffer from complex bus interconnects and architecture design when performing multi-channel, multi-stage pipelined parallel processing. This results in low storage-to-process efficiency, high power consumption and cost, and an inability to adapt to performance degradation caused by changes in algorithms and software.

Method used

It employs a multi-stage pipelined processor group, a basic parameter configuration module, a storage-computation-retrieval parameter configuration module, a storage-computation-retrieval timing control training unit, a storage-computation-retrieval scheduling controller, a CPU unit, a SOC bus, and a parallel data transfer module. Through training and configuration optimization of bus access, it achieves dynamic adaptation to performance changes and improves parallelism and access efficiency.

Benefits of technology

It simplifies system design, reduces the number of buses and memory sets, improves pipelined access efficiency, reduces costs, achieves the highest parallelism and access efficiency, and dynamically adapts to the performance requirements of different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115437994B_ABST
    Figure CN115437994B_ABST
Patent Text Reader

Abstract

The application relates to a multi-stage flow water multi-path data operation and access control system, which comprises a multi-stage flow water processor group, a basic parameter configuration module, a storage-computation-access parameter configuration module, a storage-computation-access timing control training unit, a storage-computation-access scheduling controller, a CPU unit, an SOC bus, a parallel data moving module and a DDR storage unit; under the multi-stage flow water pipeline processing of multi-path input data, the data flow water processing training of storage, computation and access is carried out, the architecture configuration scheme of the multi-stage flow water processor group or processor core and the storage-computation-access timing control method of the multi-stage flow water processing are obtained through the training data operation, and the storage-computation-access processing of the multi-stage flow water is completed according to the architecture configuration scheme and the flow water processing control method obtained through the training; the system can dynamically adapt to different performances and customer demands in various application scenarios, reduces the number of SOC buses and memory units, improves the flow water processing access efficiency, improves the access and computation parallel degree efficiency, simplifies system design and saves cost.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to a multi-stage pipeline multi-path data operation and access control system. BACKGROUND

[0002] In related technologies, the existing multi-core chip data processing technology, in realizing the data calculation and access of multi-path input signals, if multi-path multi-stage pipeline parallel processing is to be achieved, a plurality of sets of SOC buses and memory units need to be used for the corresponding chip, which is complex for the chip internal bus interconnection and architecture design, complex for the multi-set system control, and low for the storage and calculation processing efficiency, and the power consumption, area and cost of the entire system are very high, which is too expensive for ordinary applications and high-end applications.

[0003] At the same time, when the chip is completed and the algorithm and software are changed, the storage, calculation and access time and data volume change due to the change of processor allocation, and the chip control architecture cannot better adapt to the new changes, resulting in reduced performance.

[0004] In addition, in the single-bus system chip or chip sub-bus system of pipelining processing, the storage and reading of data are large, the number of different bus accesses is large, bus access conflicts occur frequently, the efficiency of the bus system is affected, and the performance of the chip is reduced. SUMMARY

[0005] Therefore, the purpose of the present application is to provide a multi-stage pipeline multi-path data operation and access control system to solve at least one problem in the prior art.

[0006] According to a first aspect of an embodiment of the present application, a multi-stage pipeline multi-path data operation and access control system is provided, comprising:

[0007] a multi-stage pipeline processor group, a basic parameter configuration module, a storage, calculation and access parameter configuration module, a storage, calculation and access timing control training unit, a storage, calculation and access scheduling controller, a CPU unit, a SOC bus, a parallel data moving module and a DDR memory unit;

[0008] The multi-stage pipeline processor group acquires multi-path input data to be processed, and performs operation processing on the multi-path input data to obtain operation processing result data.

[0009] The basic parameter configuration module is configured to configure the basic parameters of the storage, calculation and access parameter configuration module, the storage, calculation and access timing control training unit, the storage, calculation and access scheduling controller and the parallel data moving module.

[0010] The memory-computation-fetch parameter configuration module is connected with the memory-computation-fetch timing control training unit, the parallel data moving module and the memory-computation-fetch scheduling controller, configured to obtain the control parameter trained by the memory-computation-fetch timing control training unit, and configured the control parameter to the parallel data moving module and the memory-computation-fetch scheduling controller.

[0011] The memory-computation-fetch timing control training unit is connected with the basic parameter configuration module, the memory-computation-fetch scheduling controller, the memory-computation-fetch parameter configuration module and the parallel data moving module, configured to solve the architecture configuration scheme of the multi-stage pipeline processor group or processor core according to the preset algorithm operation, and the memory-computation-fetch timing control method of the multi-stage pipeline processor group or processor core, and output the parameters of the architecture configuration scheme and the timing control method.

[0012] The memory-computation-fetch scheduling controller is connected with the basic parameter configuration module, the multi-stage pipeline processor group, the memory-computation-fetch parameter configuration module, the memory-computation-fetch timing control training unit, the parallel data moving module and the CPU unit, configured to control the memory-computation-fetch processing of the multi-stage pipeline processor group, control the parallel storage and reading operation of the parallel data moving module, and output the scheduling control data to the memory-computation-fetch timing control training unit.

[0013] The CPU unit is configured to respond to and process the SOC bus occupation application, allocate the control right, and output the bus control right response signal to the memory-computation-fetch scheduling controller.

[0014] The SOC bus is configured to signal interconnection of the CPU unit, the parallel data moving module and the memory unit.

[0015] The parallel data moving module is configured to move the operation processing result data to the specified address of the external memory according to the moving control signal and the control parameter, or move the data in the specified address of the external memory to the multi-stage pipeline processor group cache.

[0016] Preferably, the memory-computation-fetch timing control training unit loads the training algorithm operation basic parameter from the basic parameter configuration module, obtains the scheduling control data of the memory-computation-fetch scheduling controller, and obtains the storage control signal and the reading control signal of the parallel data moving module.

[0017] According to the training algorithm operation basic parameter, the scheduling control data, the storage control signal and the reading control signal, and the operation timing value, the storage timing value and the reading timing value of each data of the multi-stage pipeline processor group, the memory-computation-fetch timing control parameter is output by the preset algorithm operation.

[0018] Preferably, the multi-stage pipelined processor group is further configured to send a completion indication signal to the memory-computation-fetch scheduling controller according to the operation processing result, and send a request indication signal according to the indication signal of the memory-computation-fetch scheduling controller, wherein the request indication signal is a data storage request signal and a data reading request signal;

[0019] The multi-stage pipelined processor group is further configured to receive the operation control signal of the memory-computation-fetch timing control unit, and perform data operation processing according to the operation control signal;

[0020] The multi-stage pipelined processor group is further configured to receive the request response signal of the parallel data moving module, and prepare data read-write operation according to the request response signal, wherein the request response signal is a data read response signal and a data write response signal.

[0021] Preferably, the memory-computation-fetch timing control training unit is further configured to perform memory-computation-fetch timing control training on the to-be-processed multi-channel input data, and obtain the architecture configuration scheme and configuration parameters that can make the multi-stage pipelined operation and parallel memory access optimal and the parallel memory access efficiency highest by using a preset algorithm operation according to the training result data, and output the architecture configuration scheme and configuration parameters to the memory-computation-fetch parameter configuration module.

[0022] Preferably, the memory-computation-fetch scheduling controller applies for the SOC bus control right according to the architecture configuration scheme and the parameters of the memory-computation-fetch timing control method issued by the memory-computation-fetch parameter configuration module, the basic parameters configured by the basic parameter configuration module, and controls the memory-computation-fetch processing of the multi-stage pipelined processor group, the parallel memory and reading operation of the parallel data moving module, and the output of the scheduling control data to the memory-computation-fetch timing control training unit according to the response signal of the CPU unit.

[0023] The memory-computation-fetch scheduling controller is further configured to control the multi-stage pipelined processor group to perform memory-computation-fetch processing in the order of the pipeline in the training mode, control the parallel data moving module to perform data storage and reading operation in the order of the pipeline, and output the scheduling control data to the memory-computation-fetch timing control training unit.

[0024] Preferably, the memory-computation-fetch timing control training unit is connected with the parallel data moving module.

[0025] The memory-computation-fetch timing control training unit is configured to obtain the storage and reading control signal of the parallel data moving module and the control signal of the memory-computation-fetch scheduling controller.

[0026] The memory-computation-fetch timing control training unit is configured to perform memory-computation-fetch training on the to-be-processed multi-channel input data according to the storage and reading control signal of the parallel data moving module and the control signal of the memory-computation-fetch scheduling controller, and obtain the training result.

[0027] The training results are used to perform parameter calculations for the architecture configuration scheme and the storage-computation-retrieval timing control method using a preset algorithm.

[0028] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0029] It is understood that the present invention relates to a multi-stage pipelined multi-channel data processing and access control system, comprising: a multi-stage pipelined processor group, a basic parameter configuration module, an access parameter configuration module, an access timing control training unit, an access scheduling controller, a CPU unit, a SOC bus, a parallel data transfer module, and a DDR memory unit; under multi-stage pipelined processing of multi-channel input data, through data pipelined processing training of access, computation, and retrieval, the architecture configuration scheme of the multi-stage pipelined processor group or processor core and the access timing control method of multi-stage pipelined processing are obtained from the training data; the multi-stage pipelined access processing is completed according to the trained architecture configuration scheme and pipelined processing control method; it can dynamically adapt to different performance and customer needs under various application scenarios, reduce the number of SOC buses and memory units, improve pipelined access efficiency, improve access and computation parallelism efficiency, simplify system design and save costs.

[0030] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0031] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0032] Figure 1 This is a schematic block diagram of a multi-stage pipelined multi-channel data computing and access control system according to an exemplary embodiment;

[0033] Figure 2 This is a schematic diagram of a parallel data transfer module structure according to an exemplary embodiment;

[0034] Figure 3 This is a schematic diagram of the storage-computation timing control training unit structure according to an exemplary embodiment;

[0035] Figure 4 This is a schematic diagram of the storage-compute-access scheduling controller structure according to an exemplary embodiment;

[0036] Figure 5 This is a flowchart illustrating the steps of a multi-level pipelined multiplex data processing and access control method according to an exemplary embodiment;

[0037] Figure 6This is a schematic diagram illustrating timing control based on the ratio of computation time to access time - computation bottleneck type, according to an exemplary embodiment.

[0038] Figure 7 This is a schematic diagram of timing control for access bottleneck types, illustrating that computation time is shorter than access time, according to an exemplary embodiment.

[0039] Figure 8 This is a schematic diagram of timing control for data storage time being longer than computation time in video image processing, according to an exemplary embodiment.

[0040] Figure 9 This is a flowchart illustrating a specific implementation method of this patent in video image processing where data storage takes longer than computation, according to an exemplary embodiment. Detailed Implementation

[0041] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.

[0042] First, in the existing technology, the DSP processing of multiple input intermediate frequency signals in wireless communication is used as an example for illustration:

[0043] In an architecture with four input channels and four pipelines to complete digital signal algorithm processing, each stage is completed by a group of multi-core DSPs, and the result data is transmitted to the next pipeline stage for the next stage of algorithm processing.

[0044] Another option is to use a single bus and memory system to process the four channels using a polling pipeline algorithm. The DMA controller waits for bus responses according to priority to complete the data transfer of the multi-stage pipeline processor group. However, when the pipeline DSP processor performs frequent DMA controller data transfers, severe bus conflicts occur, and the output data buffer latency is very high.

[0045] Alternative Option 2: This option employs four sets of buses, memory systems, and four sets of four-stage pipelined multi-core DSPs. The four channels are processed by parallel pipelined algorithms. The DMA controller waits for bus responses according to priority to complete the data transfer between the multi-stage pipelined processors. However, each stage of the pipelined DSP processor will experience severe bus conflicts when performing frequent DMA controller data transfers. Furthermore, the system is highly complex, costly, and consumes a lot of power.

[0046] Option 3: Use four bus and memory systems and four sets of four-stage pipelined multi-core DSPs. Each bus and memory system corresponds to one stage of pipelined multi-core DSP. The results of the four pipelined algorithms are processed and calculated in parallel and stored in each memory. The next stage of the bus and memory system reads the results of the four pipelined processing from each memory and then divides them. When there are too many data paths, the parallel data bit width will be too large, and there will be no suitable memory unit or more additional memory units will be used. In addition, multiple bus and memory systems result in high system complexity, high cost and high power consumption.

[0047] Example 1

[0048] Figure 1 This is a schematic block diagram illustrating the structure of a multi-stage pipelined multi-channel data computation and access control system according to an exemplary embodiment. See also: Figure 1 A multi-stage pipelined multi-channel data processing and access control system is provided, comprising:

[0049] A, B, ..., N multi-stage pipelined processor group 1. Basic parameter configuration module 2. Storage-computation-retrieval parameter configuration module 3. Storage-computation-retrieval timing control training unit 4. Storage-computation-retrieval scheduling controller 5. CPU unit 6. SOC bus 7. Parallel data transfer module 8 and DDR storage unit 9.

[0050] The multi-stage pipeline processor group acquires multiple input data to be processed, performs calculations on the multiple input data, and obtains the calculation result data.

[0051] The basic parameter configuration module is used to configure the basic parameters of the storage-computation-retrieval parameter configuration module, the storage-computation-retrieval timing control training unit, the storage-computation-retrieval scheduling controller, and the parallel data transfer module.

[0052] The storage-computation-retrieval parameter configuration module is connected to the storage-computation-retrieval timing control training unit, the parallel data transfer module, and the storage-computation-retrieval scheduling controller. It is used to obtain the control parameters trained by the storage-computation-retrieval timing control training unit and configure the control parameters to the parallel data transfer module and the storage-computation-retrieval scheduling controller.

[0053] The storage-compute-retrieval timing control training unit is connected to the basic parameter configuration module, the storage-compute-retrieval scheduling controller, the storage-compute-retrieval parameter configuration module, and the parallel data transfer module, respectively. It is used to calculate and solve the architecture configuration scheme of the multi-stage pipelined processor group or processor core and the storage-compute-retrieval timing control method of the multi-stage pipelined processor group or processor core according to a preset algorithm, and output the parameters of the architecture configuration scheme and the timing control method.

[0054] It should be noted that the storage-compute-access timing control training unit trains and calculates the processing time of multiple processor groups in each pipeline stage, as well as the storage and retrieval time of the pipeline result data to DDR memory, based on parameters such as the number of data to be processed from the multiple input signals, the number of pipeline stages, and the number of data paths. It calculates the parameters of the architecture configuration scheme and storage-compute-access timing control method that meet the output rate requirements, maximize the number of parallel storage and retrieval pipeline stages, and optimize the parallelism of operation and access according to the preset algorithm, and outputs them to the storage-compute-access parameter configuration module.

[0055] The storage-compute-retrieval scheduling controller is connected to the basic parameter configuration module, the multi-stage pipeline processor group, the storage-compute-retrieval parameter configuration module, the storage-compute-retrieval timing control training unit, the parallel data transfer module, and the CPU unit, respectively. It is used to control the storage-compute-retrieval processing of the multi-stage pipeline processor group, control the parallel storage and retrieval operations of the parallel data transfer module, and output scheduling control data to the storage-compute-retrieval timing control training unit.

[0056] It should be noted that the storage-compute-access scheduling controller is connected to the CPU unit for requesting and responding to SOC bus control rights; it is connected to the storage-compute-access parameter configuration module for acquiring output total delay timing parameters, arbitration margin timing parameters, parallel access initial channel number arrangement order parameters, and pipeline storage-compute-access timing control parameters for each stage; it is connected to the storage-compute-access timing control training unit for outputting timing control and timing data signals, which are then used by the storage-compute-access timing control training unit to collect data for the calculation of preset algorithms; and it is connected to the parallel data access module and the AN multi-stage pipeline processor group for controlling the parallel storage of multi-stage pipeline multi-channel signal data after computation and the splitting and acquisition of multi-stage pipeline multi-channel signal data read in parallel from the DDR memory unit.

[0057] The CPU unit is used to respond to and process the SOC bus occupancy request, allocate control rights, and output bus control right response signals to the storage-compute-access scheduling controller.

[0058] The SOC bus is used for signal interconnection between the CPU unit, the parallel data transfer module, and the memory unit;

[0059] The parallel data transfer module is used to transfer the computational processing result data to a specified address in the external memory according to the transfer control signal and the control parameters, or to transfer the data in the specified address in the external memory to the cache of the multi-level pipelined processor group.

[0060] It should be noted that the parallel data transfer module transfers the data of the AN multi-stage pipeline processor group 1 simultaneously and in parallel during each read and write operation.

[0061] The DDR storage unit is used to store and retrieve data from the SOC bus.

[0062] In practice, Figure 2 This is a schematic diagram of a parallel data transfer module structure provided in one embodiment of this application. The parallel data transfer module is used to simultaneously transfer and store the result data in the caches of multiple processor groups in each pipeline to external memory according to the amount of computation data in each stage, or to read computation data from external memory and distribute it to multiple processor group caches simultaneously. The parallel data transfer module, based on the SOC bus access right response signal input by the storage-computation-access scheduling controller, feeds back access response signals to each processor group, performs data reading through each write channel 1-n or data storage operation through each read channel 1-n, and sends read and write control signals to the storage-computation-access timing control training unit for counting during operation timing.

[0063] The access configuration parameters are used to save the initial channel number arrangement parameters for parallel access after training is completed. The combined sequence lookup table stores the channel number arrangement parameters and looks up the corresponding channel number arrangement parameters for each parallel storage or read operation for this parallel data merging and splitting.

[0064] In each parallel memory operation, the internal cache address of each level processor group memory operation is mapped to a unified internal address and stored in a unified external memory address. In each parallel read operation, the unified internal address is mapped to the internal cache address of each level processor group. The number of stored and read data is used to count the number of stored and read data. The storage and read timing adjustment control parameters are used to adjust the storage and read data offset addresses of different pipeline stages.

[0065] Parallel write to external memory operation after computation:

[0066] After each processor group completes the computation of one data stream, it sends a computation completion indication signal to the memory-compute-retrieval scheduler and a write request signal to the parallel data transfer module. Once the parallel data transfer module receives the SOC bus response signal from the memory-compute-retrieval scheduler and sends a write request response signal back to each processor group, the internal parallel write control and counting unit initiates a data transfer operation, moving data from each processor group to the data aggregation buffer control unit. The data aggregation buffer control unit, based on the channel number arrangement parameters obtained from the combination sequence lookup table, completes the splicing of the data moved out by each processor group. The spliced ​​data is then moved in parallel to external memory through the bus timing interface unit until all data specified by the storage data quantity parameter is stored, at which point the SOC bus is released. The write address mapping module completes the mapping between the internal unified address and the cache addresses of each processor.

[0067] or,

[0068] After each processor group completes the computation of one data stream, it sends a computation completion indication signal to the memory-compute-access scheduler. The DMA controller within each processor group then sends a write request signal to the parallel data transfer module. Once the parallel data transfer module receives the SOC bus response signal from the memory-compute-access scheduler and simultaneously sends a write response signal back to the DMA controllers within each processor group that initiated the write request signal, each DMA controller simultaneously transfers data from its processor cache to the parallel data transfer module. The data aggregation buffer control unit, based on the channel number arrangement parameters obtained from the combination sequence lookup table, completes the concatenation of the input data from each DMA controller group. The concatenated data is then transferred in parallel to external memory via the bus timing interface unit until all data specified by the data quantity parameters in each DMA controller group has been written, at which point the SOC bus is released. The write address mapping module completes the mapping between the internal unified address and the source cache address of each processor's internal DMA controller.

[0069] After the upstream pipeline completes storage, the downstream pipeline performs parallel read operations on external memory:

[0070] After the previous stage pipelined processor group completes the data storage specified by the read data quantity parameter, it proceeds to read data from the next stage pipelined processor group. Once the parallel data transfer module receives the data read request signal initiated by each processor group from the in-memory access controller and the SOC bus response signal from the in-memory access scheduling controller, and then sends a read request response signal back to each processor, the parallel read control and counting unit initiates a data read operation. Data is transferred from the external memory at a unified address to the data allocation buffer control unit. The data allocation buffer control unit, based on the channel number arrangement order parameter obtained from the combinatorial sequence lookup table, splits the transferred data. The split data is then simultaneously transferred to the caches of each processor group through the pipelined parallel read / write interface unit until all the data specified by the read data quantity parameter has been read, at which point the SOC bus is released. The read address mapping module completes the mapping between the internal unified address and the cache addresses of each processor.

[0071] or,

[0072] After the previous-level pipelined processor group completes the storage of data specified by the read data quantity parameter, it proceeds to the next-level pipelined data read operation. Once the parallel data transfer module receives a data read request signal from the DMA controller within each processor group, controlled by the Memory-Compute-Access Controller (MCA), and receives a SOC bus response signal from the MCA scheduler, and sends a read request response signal back to the DMA controller within each processor group, the parallel read control and counting unit initiates a data read operation. Data is transferred from the external memory at a unified address to the data allocation buffer control unit. The data allocation buffer control unit, based on the channel number arrangement order parameter obtained from the combined sequence lookup table each time, splits the transferred data, and then each DMA controller simultaneously transfers it to the cache of its respective processor group. This process continues until all data specified by the read data quantity parameter in each DMA controller group has been read, at which point the SOC bus is released. The read address mapping module completes the mapping between the internal unified address and the destination cache address of each processor's internal DMA controller.

[0073] Preferably, the storage-computation-retrieval timing control training unit loads the basic parameters for training algorithm operation from the basic parameter configuration module, obtains the scheduling control data of the storage-computation-retrieval scheduler, and obtains the storage control signal and read control signal of the parallel data transfer module.

[0074] Based on the training algorithm's computational base parameters, scheduling control data, storage control signals, and read control signals, as well as the computational timing values, storage timing values, and read timing values ​​of each data stream in the multi-stage pipeline processor group, the preset algorithm calculates and outputs the storage-computation-retrieval timing control parameters.

[0075] Preferably, the multi-stage pipeline processor group is further configured to send a completion indication signal to the storage-processing scheduling controller according to the processing result, and to send a request indication signal according to the indication signal of the storage-processing scheduling controller, wherein the request indication signal is a data storage request signal and a data retrieval request signal;

[0076] The multi-stage pipelined processor group is also used to receive the operation control signal from the storage-operation-access timing controller, and to perform data operation processing according to the operation control signal;

[0077] The multi-stage pipeline processor group is also used to receive request response signals from the parallel data transfer module, and prepare data read and write operations according to the request response signals, wherein the request response signals are data read response signals and data write response signals.

[0078] It should be noted that the multi-stage pipelined processor group issues a completion indication when the calculation is finished and a read completion indication when the reading is finished.

[0079] Preferably, the storage-computation-retrieval timing control training unit is further used to perform storage-computation-retrieval timing control training on the multi-channel input data to be processed, and to use a preset algorithm to calculate the architecture configuration scheme and its configuration parameters that can optimize the parallelism of pipeline operation and storage at each level and maximize the parallel storage efficiency based on the training results data, and output them to the storage-computation-retrieval parameter configuration module.

[0080] It should be noted that in training mode, the storage-compute-retrieval scheduling controller loads the configuration parameters input by the basic parameter configuration module. Based on the current architecture configuration scheme, and according to the valid bus control response signal input by the CPU unit, it controls the first-level unit of the pipeline processor group to perform a read operation on any one channel of data through the parallel data transfer module according to the configured data quantity. After the read is completed, the data operation is started, and after the operation is completed, the data storage is started. After the storage is completed, it controls the next-level unit of the pipeline processor group to repeat the above read, operation, and storage operation process through the parallel data transfer module. During the above process, the time consumed by the read, operation, and storage operations of each level of the pipeline processor group is counted according to the start and completion control signals of each operation and the read and write control signals input by the parallel data transfer module. After the last level of the pipeline processor group completes the above operations, the total operation time count value is collected, saved, and output to the storage-compute-retrieval timing control training unit.

[0081] See Figure 3 During the training phase, the training results data input by the storage and retrieval scheduling controller are saved through the acquisition of training data and control signals. Based on the basic parameters input by the basic parameter configuration module, and according to the required amount of input data, the architecture configuration parameter training and algorithm operation unit trains the data reading, operation and storage operations of each level of pipeline processing. Based on the time constraints of input and output data processing, the architecture configuration scheme that optimizes the parallelism of pipeline operation and retrieval at each level and maximizes the parallel access efficiency is obtained according to the preset algorithm.

[0082] The timing control parameter training and algorithm operation unit loads the training strategy (user-configured delay priority strategy, user-configured Nth level pipeline baseline timing arrangement strategy, default configuration strategy) input by the configuration strategy loading unit according to the configuration strategy loading unit, and uses the preset algorithm to operate all levels of pipeline processing, parallel storage and retrieval of timing control parameters with the best parallelism and the highest parallel access efficiency, and outputs them to the exit condition verification operation unit.

[0083] The exit condition verification calculation unit verifies that the exit condition constraints meet the requirements according to the data processing time constraints from input to output and the order of reading, calculation and storage of each level of pipeline processing. After verification, the output is saved to the configuration and control parameter saving output unit for storing the timing control parameters of the storage and calculation scheduling controller.

[0084] Overall training and computation processing logic (which can be implemented using on-chip hardware circuitry or by using a CPU for training computation):

[0085] The storage-compute-retrieval scheduling controller loads the configuration parameters input from the basic parameter configuration module. Based on the current architecture configuration scheme, and according to the valid bus control response signal input by the CPU unit, it controls the first-level unit of the pipeline processor group to perform a read operation on any one channel of data through the parallel data transfer module, according to the configured data quantity. After the read operation is completed, the data operation is started, and after the operation is completed, the data storage is started. After the storage is completed, it controls the next-level unit of the pipeline processor group to repeat the above read, operation, and storage operation process through the parallel data transfer module. During the above process, the time consumed by the read, operation, and storage operations of each level of the pipeline processor group is counted according to the start and completion control signals of each operation and the read and write control signals input by the parallel data transfer module. After the last level of the pipeline processor group completes the above operations, the total operation time count value is collected, saved, and output to the storage-compute-retrieval timing control training unit.

[0086] The architecture configuration parameter training and algorithm operation unit uses the basic parameters of the basic parameter loading unit, the training data and the clock count corresponding to the storage and retrieval operation time count value input by the control signal acquisition unit as training data, and the processing delay requirement clock count obtained by the data input rate as a constraint condition to calculate the number of operation units required for each stage of pipeline processing. If the number of operation units for each stage is greater than the total number of operation units for this stage set in the basic configuration parameters, an alarm is output and the maximum limit condition is increased. The above operation training is repeated until a new number of operation units for each stage is obtained that meets the total number of operation units for this stage set in the basic configuration parameters.

[0087] Based on the new pipelined processing architecture, the timing control parameter training and algorithm operation unit re-performs the storage-computation-retrieval scheduling control training of the above-mentioned pipelined processing at each level, and collects the clock count corresponding to the new storage-computation-retrieval operation time count value from the storage-computation-retrieval scheduling controller as intermediate training data.

[0088] Based on the training intermediate data, the basic parameters of the basic parameter configuration module, and the preset algorithm operation to obtain the storage-retrieval timing control parameters are as follows:

[0089] The maximum value of the clock cycle count B for each pipeline stage is obtained by comparison and solution. Then, the sum of the access clock count C of the previous stage and the read clock count D of the current stage is calculated as the maximum value of the access clock cycle count A.

[0090] Compare the maximum value A and the maximum value B.

[0091] If the maximum value A >= the maximum value B, the maximum value A is used as the reference control parameter for the maximum clock cycle count. According to the configured algorithm strategy, the control timing arrangement calculation of the parallel reading, parallel storage and the pipeline operation of each level of the pipeline processor group of one data is performed. The clock cycle count timing control parameters of the parallel reading and parallel storage of each level of the pipeline processor group are obtained respectively. The clock cycle count timing control parameters of different pipeline operations of each level of the pipeline processor group are also obtained.

[0092] If the maximum value A < the maximum value B, the maximum value B is used as the reference control parameter for the maximum clock cycle count. According to the configured algorithm strategy, the control timing arrangement calculation of the parallel reading, parallel storage and the pipeline operation of each level of the pipeline processor group for one channel data is performed. The clock cycle count timing control parameters for the parallel reading and parallel storage of each level of the pipeline processor group are obtained respectively. The clock cycle count timing control parameters for different pipeline operations of each level of the pipeline processor group are also obtained.

[0093] According to the configured algorithm strategy, the initial channel number arrangement order parameters for parallel access are found and calculated. Using the access timing control parameters and architecture configuration parameters, after retraining the access scheduling control, the access control signal input from the parallel data transfer module is used. The access scheduling controller collects the new access clock cycle count required for access, performs precise correction and adjustment, and uses it as the timing control parameters to be verified after training. The correctness of timing control is verified according to the output conditions.

[0094] The export conditions include:

[0095] Determine if the number of clock cycles required to complete one pipeline cycle for all channels is less than or equal to the number of clock cycles for data block input; determine if the start time for reading any channel number in the current pipeline is greater than or equal to the end time for storing the corresponding channel number in the previous pipeline; determine if the start time for calculating any channel number in the current pipeline is greater than or equal to the end time for reading the corresponding channel number in the current pipeline; determine if the start time for storing any channel number in the current pipeline is greater than or equal to the end time for calculating the corresponding channel number in the current pipeline.

[0096] Based on the unified clock cycle count value and the clock cycle count value, the exit condition verification is performed sequentially according to the starting channel number. It is determined whether the number of clock cycles required for all channels to complete one pipeline processing cycle is less than or equal to the number of data block input clock cycles. It is determined whether the start time of reading any channel number in the current pipeline is greater than or equal to the end time of storing the corresponding channel number in the previous pipeline. It is determined whether the start time of calculation for any channel number in the current pipeline is greater than or equal to the end time of reading the corresponding channel number in the current pipeline. It is determined whether the start time of storage for any channel number in the current pipeline is greater than or equal to the end time of calculation for the corresponding channel number in the current pipeline. If the exit condition is met, the storage-computation-retrieval timing control parameters (parallel read clock cycle count value and parallel storage clock cycle count value parameters), total delay clock count parameters, arbitration margin clock count parameters, and parallel storage-retrieval initial channel number arrangement order parameters are output to the storage-computation-retrieval parameter configuration module.

[0097] In work mode,

[0098] The first step involves the storage-compute-retrieval scheduling controller loading the configuration parameters input by the basic parameter configuration module. The storage-compute-retrieval parameter configuration module inputs the configuration and scheduling control parameters, communicates with the CPU unit, and initializes the internal architecture of each pipeline level, SOC bus, configuration parameters, etc., as needed.

[0099] In the second step, the storage-processing-retrieval scheduling controller requests SOC bus control from the CPU unit based on the operation completion indication signal of the one-channel data output by the multi-level pipeline processor group and the count value signal of its internal operation control timing. It then controls the pipeline processor group to initiate a write request based on the operation completion indication signal of the multi-level pipeline processor group. Based on the bus control right response valid signal input by the CPU unit, it controls the parallel data transfer module to respond to the parallel write request. According to the direct memory access method, based on the number of stored data corresponding to the one-channel of each pipeline processor group, it simultaneously transfers and stores the operation result data cached in each pipeline processor group to the specified address of the external memory unit in parallel until the data storage specified by the number of stored data parameters is completed. The parallel data transfer module then outputs a parallel storage completion indication signal.

[0100] The third step involves the parallel storage completion indication signal input by the parallel data transfer module, and the storage control timer of the storage-computation-retrieval scheduling controller reaching a count value. The module then requests SOC bus control rights from the CPU unit. Based on the valid bus control response signal input by the CPU unit, it controls the multi-stage pipelined processors to initiate parallel read requests. The parallel data transfer module responds to these read requests and, using direct memory access, moves data from the specified address in the external memory unit according to the number of data reads corresponding to each pipelined processor stage. This allocates one data path corresponding to each pipelined processor stage. Based on the address information of each pipelined processor stage, the data is simultaneously and concurrently stored in the cache of each pipelined processor stage until the data specified by the number of reads parameter is read. Finally, the parallel data transfer module outputs a parallel read completion indication signal.

[0101] The fourth step is to control the multi-level pipeline processor groups to complete the calculation and processing of their respective corresponding 1-channel data according to the reading completion indication signal input by each level of pipeline processor group in the storage, computation and retrieval scheduling controller, and output the calculation completion indication signal to the storage, computation and retrieval scheduling controller according to the reading control clock counter of each level of pipeline processor group in the storage, computation and retrieval scheduling controller.

[0102] Following a sequential polling method based on multiple input data streams, the storage-compute-retrieval scheduling controller repeats steps two through four above. See Table 1:

[0103]

[0104]

[0105] Table 1

[0106] Preferably, the storage-compute-retrieval scheduling controller requests SOC bus control rights based on the architecture configuration scheme and storage-compute-retrieval timing control method parameters issued by the storage-compute-retrieval parameter configuration module and the basic parameters configured by the basic parameter configuration module. Based on the response signal of the CPU unit, it controls the storage-compute-retrieval processing of the multi-stage pipelined processor group, controls the parallel storage and retrieval operations of the parallel data transfer module, and outputs scheduling control data to the storage-compute-retrieval timing control training unit.

[0107] The storage-compute-retrieval scheduling controller is also used in training mode to control the multi-stage pipelined processor group to perform storage-compute-retrieval processing step by step in pipeline order, control the parallel data transfer module to perform data storage and retrieval operations step by step in pipeline order, and output scheduling control data to the storage-compute-retrieval timing control training unit.

[0108] See Figure 4 , Figure 4This is a schematic diagram of the storage-compute-retrieval scheduling controller structure in one embodiment of this application. Specifically, the processing flow completed by the storage-compute-retrieval scheduling controller is as follows:

[0109] The processor group controlling each stage of pipeline processing performs computation, storage, and retrieval operations on each of the multiple input data streams according to the configured polling priority order. The control process is as follows:

[0110] Timing is performed according to the parameters of the configured storage-and-retrieval timing control method. Once one input data path is ready in the cache of each processor group, the input data is processed. After processing one input data path and storing the result data in the processor cache, the current processor group begins processing the next input data path once the next path is ready in the cache. Simultaneously, if the previous result data path in the cache of the current processor group is ready, and the parallel read timing control timer and the processing timing control timer for the previous data path in the current pipeline are complete, the parallel data transfer module is controlled to adjust the data transfer according to the stored data. According to the number of data specified by the quantity parameter, the previous result data in the cache of each level of processor group is moved and stored in parallel to the specified address of external memory. After the parallel storage is completed and the parallel storage control timer is completed, the parallel data moving module is controlled to read the result data of the previous pipeline from the specified address of external memory as the input data of the next level pipeline. According to the number of data specified by the data reading quantity parameter, the data is simultaneously distributed to the cache of each processor group of each next level pipeline according to the specified number of data for each different path, until the specified number of data for each path of each level is read. The above process is repeated.

[0111] The overall processing logic is as follows:

[0112] The loading of basic parameters from the basic parameter configuration module is complete. The loading of configuration and control parameters from the timing control method configured by the storage-computation-retrieval timing control training unit is complete (including system total delay parameters, bus arbitration margin time parameters, storage-computation-retrieval path number starting sequence number parameters, parallel storage timing control parameters, parallel read timing control parameters, and pipeline operation timing control parameters at each stage).

[0113] After switching to the working mode, the working mode control state machine starts. If it receives a first-level input data preparation completion indication signal from the pipeline processor group or if the parallel storage data transfer is completed and the control timing value is restored to the initial value (e.g., zero value), it requests the SOC bus occupancy right from the control communication unit through the CPU, outputs a valid SOC bus occupancy indication signal to the parallel data transfer module, and simultaneously outputs a read indication signal to control each level of the processor group to initiate a read request. At this time, the parallel read control timing starts and waits for the parallel data transfer module to simultaneously distribute and transfer data from external memory to the cache of each processor group. When it receives a read completion indication signal from any level of the processor group, and the timing of the operation sequence control of this level is at the initial value (e.g., zero value), it outputs an operation indication signal to control this level of the processor group to start operation processing, and simultaneously starts the corresponding operation control timing. After the parallel data transfer module inputs a parallel read completion indication signal, and the parallel read control timing value reaches the set value of the parallel read timing control parameter, the state machine restores the parallel read control timing to the initial value (e.g., zero value).

[0114] During the process, if a completion indication signal is received from any level of processor group, the operation control timing of this level will be restored to the initial value (such as zero value). However, the completion indication status of this level will be stored internally until the write request indication signal corresponding to this level is valid after the parallel memory is started, at which point the completion indication status of this level will be cleared.

[0115] When the parallel read control timing is restored to its initial value (e.g., zero), and all the operation completion indication states stored internally in each stage are valid, the CPU requests and communicates with the control to obtain the SOC bus occupancy right, outputs the SOC bus occupancy valid indication signal to the parallel data transfer module, and simultaneously outputs a storage indication signal to control each level of the processor group to initiate a storage request. At this time, the parallel storage control timing is started and waits for the parallel data transfer module to simultaneously transfer data from the cache of each processor group to the external memory in parallel. After the parallel data transfer module inputs the parallel storage completion indication signal, and the parallel storage control timing value reaches the set value of the parallel storage timing control parameter, the state machine restores the parallel storage control timing to its initial value (e.g., zero). Thus, the state machine runs in a loop according to the above method.

[0116] During the training phase, the training mode control state machine sequentially controls the parallel data transfer module to read, process, and store data for each stage of the pipeline from front to back until the last stage obtains the output data for the first time. The training result feedback module then outputs the collected processor operation time, storage time, and read time for each stage, as well as the order of calculation, storage, and read, to the storage-computation-retrieval timing control training unit.

[0117] Preferably, the storage-computation-retrieval timing control training unit is connected to the parallel data transfer module;

[0118] Used to acquire the storage and read control signals of the parallel data transfer module, and to acquire the control signals of the storage-compute-retrieval scheduling controller;

[0119] The storage and retrieval control signals of the parallel data transfer module and the control signals of the storage-computation-retrieval scheduling controller are used to perform storage-computation-retrieval training on the multi-channel input data to be processed, and the training results are obtained.

[0120] The training results are used to perform parameter calculations for the architecture configuration scheme and the storage-computation-retrieval timing control method in the preset algorithm.

[0121] See Figure 5 , Figure 5 This application provides a flowchart of a multi-level pipelined multi-channel input data processing and access control method, including:

[0122] Step S101: Obtain the multi-channel input data to be processed, and perform step-by-step calculation, storage and reading of the multi-channel input data to obtain the calculation result data;

[0123] Step S102: Obtain the training result data of the storage-computing-retrieval timing control training unit for algorithm operation to obtain the architecture configuration scheme parameters;

[0124] Step S103: Using the new architecture configuration scheme with the parameters set by the architecture configuration scheme, repeat step S101;

[0125] Step S104: Obtain the training result data obtained by the storage-computation-retrieval timing control training unit for algorithm operation to obtain timing control parameters for pipelined processing operation, parallel storage and retrieval, and verify that the parameters meet the exit conditions;

[0126] Step S105: Obtain the multi-channel input data to be processed, perform calculation and processing control according to the timing control parameters, start the calculation and processing of the multi-channel input data, and obtain the calculation and processing result data;

[0127] Step S106: Perform parallel storage control according to the timing control parameters, and move the operation processing result data to a specified address in the external memory in parallel;

[0128] Step S107: Perform parallel read control according to the timing control parameters, and read the result data stored in the specified address of the external memory into multiple processor group caches in parallel for the next operation.

[0129] Step S108: Repeat steps S105 to S108 until all data has been processed.

[0130] This application provides a multi-stage pipelined multi-channel data processing and access control method.

[0131] By acquiring multiple input data to be processed, the multiple input data is processed step by step, including operation, storage, and retrieval, to obtain the operation processing result data; the result data obtained from the training of the storage-computation-retrieval timing control training unit is used for algorithm operation to obtain architecture configuration scheme parameters; a new architecture configuration scheme is set using the architecture configuration scheme parameters, and step S101 is repeated; the result data obtained from the training of the storage-computation-retrieval timing control training unit is used for algorithm operation to obtain timing control parameters for pipelined operation, parallel storage, and retrieval, and the parameters are verified to meet the exit conditions; the multiple input data to be processed is acquired, and operation processing control is performed according to the timing control parameters to start the operation processing of the multiple input data to obtain the operation processing result data; parallel storage control is performed according to the timing control parameters to move the operation processing result data in parallel to a specified address in external memory; parallel retrieval control is performed according to the timing control parameters to read the result data stored in the specified address in external memory in parallel into multiple processor group caches for the next operation processing; steps S105 to S108 are repeated until all data is processed. Intelligent training can yield architecture configurations and timing control methods that optimize the parallelism of pipelined operations and access at each stage and maximize parallel access efficiency. This can reduce the number of SOC buses, storage controllers, and their storage media, simplifying design, saving costs, and adapting to new changes and customer needs.

[0132] The multi-stage pipelined multi-channel data computation and access control system provided in this application employs a parallel data transfer module added within the chip to the SOC bus for multiple processing units. Through regularized storage and retrieval control combined with maximum computational parallelism, the storage processing time and retrieval processing time of each stage are simultaneously aligned, avoiding frequent SOC bus occupancy conflicts and maximizing the efficiency of each data access, thereby minimizing the number of SOC bus and DDR memory units required. This solves the problem of frequent SOC bus occupancy conflicts that occur within a single bus system when multiple processor units perform pipelined computation and large-scale data transfer, improving SOC bus access efficiency. Using a single SOC bus and memory system also saves bus and memory resources / costs. For multiple input data streams, instead of using more processors to compute each stream in parallel at each pipeline stage, serial polling of each stream saves processor resources / costs. The control scheme can be modified through the chip's training function to adapt to flexible and varied algorithms.

[0133] In some embodiments, an architecture in which multiple processor groups (or processing units) within a single SOC chip are connected to the system SOC interconnect bus via a parallel data transfer module, jointly performing multi-stage pipelined computing applications.

[0134] Figure 6 This is a timing control diagram of the computation bottleneck type, provided by one embodiment of this application, showing that the computation time is longer than the access time.

[0135] Figure 7 This is a timing control diagram of an embodiment of the present application, showing that the computation time is shorter than the access time, and the access bottleneck type is shorter.

[0136] Figure 8 This is a timing control diagram illustrating the bottleneck type of video signal processing where data storage takes longer than computation time, provided in one embodiment of this application.

[0137] The present invention also provides a specific usage example, please refer to [link / reference]. Figure 9 , Figure 9 This is a flowchart illustrating a specific implementation method of input video signal data according to an embodiment of this application, as follows: Figure 9 As shown,

[0138] Given a defined hierarchical division of the input data stream pipeline processing algorithm, the number and bit width of the output data are determined. The required DRAM bit width for parallel access is the sum of the bit widths of the 6 pipeline stages.

[0139] Training Phase: Based on the configured parameters such as the number of input macroblock operation data, the number of data paths, the number of pipeline stages, the processing latency requirements, the total number of processor cores / processing units at each stage, and the algorithm training strategy, one channel of data is controlled step by step according to the configured amount of data to be processed at each stage. The data is processed step by step from the first stage pipeline input to the last stage pipeline output. The operation time of each pipeline processor group, the storage and retrieval time between the pipeline result data and DRAM are statistically recorded. According to the processing latency requirements and the processing time calculation method, the number of operation processing units required for each stage of the pipeline processor group to meet the 4 input data is calculated, and a new architecture configuration scheme is obtained.

[0140] After initializing the architecture based on the new architecture configuration, the 3rd, 4th, 5th and 6th stage pipelines are processed using 2, 5, 2 and 4 sets of processor cores / processing units for parallel computing. Then, one channel of data is controlled step by step according to the configured amount of data to be processed at each stage. The data is processed step by step from the first stage pipeline input to the last stage pipeline output. The computing time of each stage pipeline processor group and the storage and retrieval time between each stage pipeline result data and DRAM are statistically recorded. In this way, the real processing time data of each stage can be obtained through training.

[0141] Based on the training obtained, real processing time data for each stage is used. According to the configured algorithm training strategy, a preset algorithm is adopted to first calculate and classify the storage and retrieval times of the preceding and following stages of a pipelined process that are greater than the computation time of any other pipelined process. Then, under these conditions, the preset algorithm is used to calculate the storage-computation-retrieval timing control parameters for each stage, so that the calculation results satisfy the optimal parallelism of the 6-stage pipelined operation and storage-retrieval, and the highest parallel storage-retrieval efficiency. At the same time, the total demonstration clock count parameter, the arbitration margin clock count parameter, and the parallel access initial channel number arrangement order parameter are calculated. Finally, the obtained storage-computation-retrieval timing control parameters are judged according to the constraints of the output condition verification operation. If the requirements are met, the total delay clock count parameter, the arbitration margin clock count parameter, the parallel access initial channel number arrangement order parameter, and the storage-computation-retrieval timing control parameters are output and saved to the storage-computation-retrieval parameter configuration module for storage-computation-retrieval timing control execution in the working mode.

[0142] Control execution:

[0143] Once in working mode, after the data of 4-way row macroblocks is cached in DRAM according to the parallel access arrangement, the storage-processing-access scheduling control is started.

[0144] After the memory-compute-access scheduler requests and acquires the SOC bus from the CPU, it controls each processor group to initiate a read request and controls the parallel data transfer module to respond to the request. At the same time, it starts the internal parallel read timer. It controls the parallel data transfer module to read row macroblock data from DRAM in parallel according to the number of macroblock data to be read in parallel, splits it according to the initial channel number arrangement parameters of the parallel access, and moves it to the input ping-pong buffer of the 6-stage pipeline. When the data of any stage of pipeline is read, it outputs a read completion indication signal to the memory-compute-access scheduler. When the memory-compute-access scheduler receives any read completion indication signal and the counter corresponding to the ping-pong operation timer is at its initial zero value, it starts the corresponding ping-pong operation timer of this stage of pipeline and controls this stage of pipeline to start the operation. After the operation is completed, the result data is output to the output ping-pong buffer of the pipeline.

[0145] During the process, the parallel data transfer module outputs a parallel reading completion indication, and after the parallel reading timer of the storage-computation-retrieval scheduling controller reaches the set value of the reading timing control parameter, the parallel data storage operation is started. Furthermore, during the process, once a completion indication signal is received from any level of pipeline processing, the storage-computation-retrieval scheduling controller keeps the previous operation control timing value in the ping-pong counter of this level unchanged, waits for the parallel storage indication signal of the corresponding ping-pong sequence, and then restores this counter to the initial value of zero and starts the parallel data storage operation.

[0146] The data combination lookup table varies depending on the number of pipeline stages and the number of data paths. For example, if the current number of pipeline stages is set to M=6 and the number of paths is set to N=4, and the initial channel number arrangement order for parallel access is 3-1-3-1-3-1, then the channel numbers corresponding to the data initially read and split into 6 pipeline stages are: pipeline stage 1: channel number 3, pipeline stage 2: channel number 1, pipeline stage 3: channel number 3, pipeline stage 4: channel number 1, pipeline stage 5: channel number 3, pipeline stage 6: channel number 1;

[0147] The second parallel read splits the data to the 6-level pipeline with the following channel numbers: pipeline level 1: channel number 4, pipeline level 2: channel number 2, pipeline level 3: channel number 4, pipeline level 4: channel number 2, pipeline level 5: channel number 4, pipeline level 6: channel number 2.

[0148] The data split into six pipelines for the third parallel read corresponds to the following channel numbers: pipeline 1: channel number 1, pipeline 2: channel number 3, pipeline 3: channel number 1, pipeline 4: channel number 3, pipeline 5: channel number 1, pipeline 6: channel number 3.

[0149] The channel numbers corresponding to the data split into 6 pipelines in the fourth parallel read are: pipeline 1: channel number 2, pipeline 2: channel number 4, pipeline 3: channel number 2, pipeline 4: channel number 4, pipeline 5: channel number 2, pipeline 6: channel number 4.

[0150] The above loop is performed according to the order of the counts read.

[0151] After initiating the parallel data storage operation, the storage-compute-access scheduling controller requests and acquires the SOC bus from the CPU, then controls each level of the processor group to initiate a storage request and controls the parallel data transfer module to respond to the request. Simultaneously, it starts the internal parallel read timing. The parallel data transfer module reads row macroblock data from an output ping-pong buffer corresponding to the 6-stage pipeline according to the number of parallel storage macroblocks. Following the initial parallel access channel number arrangement parameter, it retrieves the previous channel number from the data combination lookup table, concatenates the read data from the 6-stage pipeline, and stores it in parallel at a specified address in DRAM. At this point, the concatenation of the 6-stage pipeline data in parallel storage corresponds to the following channel numbers: Pipeline 1: Channel 2, Pipeline 2: Channel 4, Pipeline 3: Channel 2, Pipeline 4: Channel 4, Pipeline 5: Channel 2, Pipeline 6: Channel 4.

[0152] After the parallel storage operation is started, the ping-pong counter corresponding to the pipeline operation is cleared and restored to its initial value of zero according to the channel number in the table above.

[0153] Once the parallel data transfer module inputs a parallel storage completion indication signal and the internal parallel storage timer of the storage-computation-retrieval scheduling controller reaches the set value of the parallel storage timing control parameter, it outputs a parallel storage completion indication signal to the storage-computation-retrieval scheduling controller to start the parallel read operation. During the process, once a completion indication signal is received from any level of pipeline processing, the storage-computation-retrieval scheduling controller keeps the current operation control timer value of a ping-pong counter corresponding to this level unchanged, and waits for the parallel storage indication signal of the corresponding ping-pong sequence before restoring this counter to its initial value of zero.

[0154] Therefore, by repeating the above operations and access control processing, the processing of 6-level pipelined multi-channel continuous input data streams can be achieved.

[0155] It is understood that the same or similar parts in the above embodiments can be referred to each other, and the contents not described in detail in some embodiments can be referred to the same or similar contents in other embodiments.

[0156] It should be noted that in the description of this invention, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this invention, unless otherwise stated, "a plurality of" means at least two.

[0157] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0158] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

Claims

1. A multi-stage pipelined multi-channel data processing and storage control system, characterized in that, include: Multi-stage pipelined processor group, basic parameter configuration module, storage-computation-retrieval parameter configuration module, storage-computation-retrieval timing control training unit, storage-computation-retrieval scheduling controller, CPU unit, SOC bus, parallel data transfer module and DDR memory unit; The multi-stage pipeline processor group acquires multiple input data to be processed, performs calculations on the multiple input data, and obtains the calculation result data. The basic parameter configuration module is used to configure the basic parameters of the storage-computation-retrieval parameter configuration module, the storage-computation-retrieval timing control training unit, the storage-computation-retrieval scheduling controller, and the parallel data transfer module. The storage-computation-retrieval parameter configuration module is connected to the storage-computation-retrieval timing control training unit, the parallel data transfer module, and the storage-computation-retrieval scheduling controller. It is used to obtain the control parameters trained by the storage-computation-retrieval timing control training unit and configure the control parameters to the parallel data transfer module and the storage-computation-retrieval scheduling controller. The storage-compute-retrieval timing control training unit is connected to the basic parameter configuration module, the storage-compute-retrieval scheduling controller, the storage-compute-retrieval parameter configuration module, and the parallel data transfer module, respectively. It is used to calculate and solve the architecture configuration scheme of the multi-stage pipelined processor group or processor core and the storage-compute-retrieval timing control method of the multi-stage pipelined processor group or processor core according to a preset algorithm, and output the parameters of the architecture configuration scheme and the timing control method. The storage-compute-retrieval timing control training unit loads the basic parameters for training algorithm operation from the basic parameter configuration module, obtains the scheduling control data of the storage-compute-retrieval scheduler, and obtains the storage control signal and read control signal of the parallel data transfer module; and outputs the storage-compute-retrieval timing control parameters by a preset algorithm based on the basic parameters for training algorithm operation, scheduling control data, storage control signal and read control signal, as well as the operation timing value, storage timing value and read timing value of each data path of the multi-stage pipeline processor group. The storage-computation-retrieval timing control training unit is also used to train the storage-computation-retrieval timing control on the multi-channel input data to be processed, and to use a preset algorithm to calculate the architecture configuration scheme and its configuration parameters that can optimize the parallelism of pipeline operation and storage at each level and maximize the parallel storage efficiency based on the training results data, and output them to the storage-computation-retrieval parameter configuration module. The storage-compute-retrieval scheduling controller is connected to the basic parameter configuration module, the multi-stage pipeline processor group, the storage-compute-retrieval parameter configuration module, the storage-compute-retrieval timing control training unit, the parallel data transfer module, and the CPU unit, respectively. It is used to control the storage-compute-retrieval processing of the multi-stage pipeline processor group, control the parallel storage and retrieval operations of the parallel data transfer module, and output scheduling control data to the storage-compute-retrieval timing control training unit. The storage-compute-retrieval scheduling controller requests SOC bus control rights based on the architecture configuration scheme and storage-compute-retrieval timing control method parameters issued by the storage-compute-retrieval parameter configuration module and the basic parameters configured by the basic parameter configuration module. Based on the response signal of the CPU unit, it controls the storage-compute-retrieval processing of the multi-stage pipelined processor group, controls the parallel storage and retrieval operations of the parallel data transfer module, and outputs scheduling control data to the storage-compute-retrieval timing control training unit. The storage-compute-retrieval scheduling controller is also used in training mode to control the multi-stage pipeline processor group to perform storage-compute-retrieval processing step by step in pipeline order, control the parallel data transfer module to perform data storage and retrieval operations step by step in pipeline order, and output scheduling control data to the storage-compute-retrieval timing control training unit. The CPU unit is used to respond to and process the SOC bus occupancy request, allocate control rights, and output bus control right response signals to the storage-compute-access scheduling controller. The SOC bus is used for signal interconnection between the CPU unit, the parallel data transfer module, and the DDR memory unit. The parallel data transfer module is used to transfer the computational processing result data to a specified address in the external memory according to the transfer control signal and the control parameters, or to transfer the data in the specified address in the external memory to the cache of the multi-level pipelined processor group.

2. The system according to claim 1, characterized in that, The multi-stage pipeline processor group is also used to send a completion indication signal to the storage-processing scheduler based on the processing result, and to send a request indication signal based on the indication signal of the storage-processing scheduler. The request indication signal is a data storage request signal and a data retrieval request signal. The multi-stage pipelined processor group is also used to receive the operation control signal from the storage-operation-access timing controller, and to perform data operation processing according to the operation control signal; The multi-stage pipeline processor group is also used to receive request response signals from the parallel data transfer module, and prepare data read and write operations according to the request response signals, wherein the request response signals are data read response signals and data write response signals.

3. The system according to claim 1, characterized in that, The storage-computation-retrieval timing control training unit is connected to the parallel data transfer module; Used to acquire the storage and read control signals of the parallel data transfer module, and to acquire the control signals of the storage-compute-retrieval scheduling controller; The storage and retrieval control signals of the parallel data transfer module and the control signals of the storage-computation-retrieval scheduling controller are used to perform storage-computation-retrieval training on the multi-channel input data to be processed, and the training results are obtained. The training results are used to perform parameter calculations for the architecture configuration scheme and the storage-computation-retrieval timing control method using a preset algorithm.

Citation Information

Patent Citations

  • Hierarchical parallel modular sequence image real-time processing device

    CN102306371A

  • Heterogeneous processing system, processor and task processing method for federated learning

    CN111813526A