Processing device and program for processing device

JP2024000852A5Pending Publication Date: 2025-06-23CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022099799
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-06-21
Publication Date
2025-06-23

AI Technical Summary

Technical Problem

Conventional programmable signal processing devices require an interface circuit between DSPs, a control CPU, and its control program, increasing the circuit scale of the entire system.

Method used

A programmable signal processing circuit with a configuration that includes a plurality of execution units, a register file section with series-connected registers, and a control unit that issues shift signals for parallel operation, allowing data exchange without additional interface circuits.

Benefits of technology

This configuration simplifies the structure for data exchange between parallel program execution units, reducing the circuit scale and improving operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a programmable signal processing circuit that is simplified in the data transfer configuration between program executing sections running in parallel.SOLUTION: A programmable signal processing circuit comprises: a plurality of execution sections that can execute programs in parallel and have memories for storing programs to be executed by themselves; a register file section that has a plurality of registers connected in series for respective ones of the execution sections to utilize, so as to receive a shift signal and transfer data which each of the plurality of registers holds to a register positioned downstream by one in series connection; an issuance section that issues a shift signal if respective ones of the execution sections complete the execution of one cycle of a program; and a shift control section that, if each of the executing sections delivers data to another executing section, stores in program memories of the respective execution sections a program including an instruction that sets a register located one upstream of the register that other execution units refer to when inputting data, as a storage destination of the data resulting from one cycle of processing.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a programmable signal processing circuit, and more particularly to a programmable signal processing circuit having a plurality of different program execution units. [Background technology]

[0002] Conventionally, programmable signal processing devices (signal processors) include those that dynamically configure the signal processing circuit itself, such as FPGAs (Field-Programmable Gate Arrays) and reconfigurable circuits, and those that run programs that execute a sequence of instructions sequentially, such as DSPs.

[0003] For example, Patent Document 1 discloses a method of cascading DSPs that execute different programs to configure a multi-stage configuration. Different programs mean that the processing time required for each program differs between DSPs. Therefore, in order to improve throughput, it is necessary to devise a way to prevent a decrease in operating rate due to the influence of the BUSY state of other DSPs that operate in cooperation.

[0004] In Patent Document 2, DSPs that execute different programs are configured to calculate data read from memory via a DAMC (Direct Memory Access Controller) and write the calculation results to memory. Patent Document 2 also discloses a method of IPC (Inter-Process Synchronization) for DSP synchronization, thereby realizing a data flow equivalent to that of DSPs connected in series. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Patent No. 4222808 [Patent Document 2] JP 2011-89913 A Summary of the Invention [Problem to be solved by the invention]

[0006] However, in the conventional technology disclosed in the above-mentioned Patent Document 2, although the operating rate of the DSP is improved, an interface circuit between the DSPs, a control CPU and its control program, etc. are required, which increases the circuit scale of the entire system.

[0007] The present invention has been made in view of the above problems, and aims to provide a technique for simplifying the structure relating to data transfer between program execution units operating in parallel in a programmable signal processing circuit. [Means for solving the problem]

[0008] In order to solve this problem, for example, a programmable signal processing circuit according to the present invention has the following configuration. 1. A programmable signal processing circuit, comprising: A plurality of execution units each having a program memory and operable in parallel by executing a program stored in the program memory; a register file unit including a plurality of serially connected registers, each of which is available to the plurality of execution units, and which transfers data held in each of the plurality of registers to a register located one register downstream in the serial connection in response to receiving a shift signal; a control unit that writes the program into a program memory in the plurality of execution units, and issues the shift signal to the plurality of execution units when each of the plurality of execution units has finished executing one cycle of the program; the control unit stores the program in a program memory in the execution units, the program including an instruction for specifying, when each of the execution units passes data to another execution unit, a register located immediately upstream of a register referenced when the other execution unit inputs data, as a storage destination for data resulting from the processing in one cycle; Each of the plurality of execution sections is characterized in that, after executing the program for one cycle, in response to receiving the shift signal, it executes the program again. Effect of the Invention

[0009] According to the present invention, it is possible to simplify the structure relating to data transfer between program execution units operating in parallel in a programmable signal processing circuit. [Brief description of the drawings]

[0010] [Figure 1] FIG. 2 is a circuit configuration diagram of a programmable signal processor according to an embodiment. [Diagram 2] FIG. 4 is a circuit configuration diagram of a shift register group in a register file unit according to the embodiment. [Diagram 3] FIG. 4 is a circuit configuration diagram of a first processing unit in the embodiment. [Figure 4] 4 is a diagram showing bit assignments of instructions executed by a first processing unit in the embodiment. [Diagram 5] 3 is an equivalent circuit diagram when each execution unit of the programmable signal processor according to the embodiment executes a program. [Figure 6] FIG. 2 is a diagram showing a list of programs executed by each execution unit of the programmable signal processor according to the embodiment. [Figure 7] FIG. 4 is a timing chart for explaining the processing of the programmable signal processor according to the embodiment. [Figure 8] FIG. 11 is a diagram for comparing various methods for data transfer. [Figure 9] FIG. 4 is a diagram showing a memory map of an IO bus in the embodiment. [Figure 10] FIG. 1 is a block diagram of an electronic device according to an embodiment. [Figure 11] 4 is a flowchart showing a processing procedure of a CPU according to the embodiment. [Figure 12] FIG. 11 is a circuit configuration diagram of a programmable signal processor according to a second embodiment. [Figure 13]FIG. 11 is a diagram showing a program list executed by a first processing unit according to the second embodiment. [Figure 14] FIG. 11 is a circuit configuration diagram of a programmable signal processor according to a third embodiment. [Figure 15] FIG. 13 is a block diagram showing the configuration of an electronic device according to a fourth embodiment. [Figure 16] FIG. 13 is a view showing a program list executed by a first processing unit in the fourth embodiment. [Figure 17] FIG. 13 is a diagram showing bit assignment of floating-point data in the fourth embodiment. [Figure 18] FIG. 13 is an equivalent circuit diagram of a signal processor according to a fourth embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.

[0012] [First embodiment] 10 shows a block configuration diagram of an information processing device in the first embodiment. For ease of understanding, an example will be described in which this device is applied to an image capture device such as a digital camera.

[0013] This device has a programmable signal processor 1001, a CPU (Central Processing Unit) 1002, memories 1003a to 1003d, a memory bus 1002, and an IO bus 1006. The CPU 1002 controls the entire device, and also loads and starts programs into the programmable signal processor 1001. The memories 1003a to 1003d are used to store programs executed by the CPU 1002 and the signal processor 1001, as well as image data to be processed and image data after processing, and are composed of RAM, ROM, etc.

[0014] The memory bus 1002 is a multi-layer type memory bus that connects multiple memories (memories 1003a to 1003d) and multiple bus masters (including the illustrated signal processor 1001 and CPU 1005). The IO bus 1006 is an IO bus that connects multiple devices and controllers (such as the signal processor 1001 and CPU 1005).

[0015] Note that Fig. 10 shows only the main parts related to the present invention. For example, when the configuration of Fig. 10 is applied to an imaging device such as a digital camera, an imaging unit for acquiring image data to be processed, a recording unit for recording the processed image data in a non-volatile memory or the like, and a user interface (operation unit, display unit) between a user and this device are connected to the IO bus 1006 as the above-mentioned devices.

[0016] The CPU 1005 controls the entire device by executing programs stored in memories 1003a to 1003d via a memory bus 1002. The CPU 1005 also controls the signal processor 1001 via an IO bus 1006. Specifically, the CPU 1005 reads a plurality of programs to be executed by the signal processor 1001 from the memories 1003a to 1003d, writes the programs to a plurality of memories (reference numerals 102, 150, 111, and 113 described later) built in the signal processor 1001 via the IO bus 1006, and instructs the signal processor 1001 to start up and executes processing.

[0017] When the signal processor 1001 receives a startup instruction from the CPU 1001 via the IO bus 1006, it reads out the image data to be processed that is stored in the memories 1003a to 1003d, performs filtering processing, and writes the image data after filtering processing to the memories 1003a to 1003d.

[0018] Next, a more detailed description will be given of the configuration of the signal processor 1001 in the embodiment. Fig. 1 is a circuit configuration diagram of the signal processor 1001 in the embodiment.

[0019] As shown in the figure, the signal processor 1001 has a read unit 101, an assignment unit 104, a first processing unit 110, a second processing unit 112, a write unit 114, a register file unit 106, and a shift control unit 117. The read unit 101, the assignment unit 104, the first processing unit 110, and the second processing unit 112 each have a built-in memory for holding a program to be executed by the read unit. The illustrated reference numerals 102, 104, 111, and 113 indicate these memories. As described above, the CPU 1005 stores the programs (four) for the signal processor 1001 read from the memories 1003a to 1003d in the memories 103, 104, 111, and 113 via the IO bus 1006. The CPU 1005 also performs a process of setting a value required for the signal processor 1001 in a control register held by the signal processor 1001.

[0020] The register file unit 106 is composed of a general-purpose shift register group 107, a general-purpose register group 108, and an output register 109. The assignment unit 104, the first processing unit 110, the second processing unit 112, and the writing unit 114 can each directly use each register in the register file unit 106.

[0021] The general-purpose shift register group 107 is composed of a plurality of shift registers connected in series. Each shift register can hold 32-bit data. In response to a shift cycle signal from the shift control unit 117, each shift register transfers the data it holds to the next downstream shift register (described in detail later). Note that, although there is no particular limit to the number of shift registers constituting the general-purpose shift register group 107 in the embodiment, the general-purpose shift register group 107 is composed of 16 shift registers in the embodiment. When expressed as operands in a program executed by the signal processor 1001, these serially connected shift registers are represented as R00, R01, R02, ..., R0e, R0f from the most upstream to the downstream (the two characters on the right are represented in hexadecimal).

[0022] The general-purpose register group 108 is made up of a plurality of registers. There is no particular limit to the number of these registers, but in this embodiment, it is 15. When expressed as operands in a program, the individual registers constituting the general-purpose register group 108 are denoted as R10, R11, R12, ..., R1e. Unlike the shift registers in the general-purpose shift register 107, these registers hold data (32 bits) regardless of the shift cycle signal.

[0023] The output register 109 is a register for holding the data (32 bits) after the filter process. The output register 109 is represented as "R1f" in operand notation.

[0024] As described above, the register file unit 106 in this embodiment has a total of 32 registers.

[0025] The reading unit 101 executes a program stored in the internal program memory 102 to read out image data to be processed from memories 1003a to 1003d in pixel units and store the image data in a FIFO (First In First Out) memory 103. In this way, the processing of the reading unit 101 is a relatively simple process of reading out image data from a specified address and storing the image data in the FIFO memory 103, so a program indicating a read address update process is stored in the program memory 102. Note that the reading unit 101 executes the read process when there is free space in the FIFO memory 103. When there is no free space in the FIFO memory 103 and the memory becomes full, the reading unit 101 temporarily stops reading and waits for free space to be generated.

[0026] Assignment unit 104 executes a program stored in program memory 105 to input image data from FIFO memory 1003 and assign the image data to a corresponding register in register file unit 106. That is, program memory 105 stores a program indicating to which register the input image data is to be assigned. When assignment unit 104 inputs data from FIFO memory 103, a free area for one piece of data is generated in FIFO memory 103.

[0027] The first processing unit 110 executes a program stored in the program memory 111. The first processing unit 110 uses a register in the register file unit 106 when performing various arithmetic processing in accordance with the program.

[0028] The second processing unit 110 executes a program stored in the program memory 112. When performing various arithmetic processing according to the program, the second processing unit 110 performs the processing using a register in the register file unit 106. In addition, the second processing unit 112 stores image data resulting from the final filter processing in the output register 109 (shift register "R1f" in the program).

[0029] The writing unit 114 stores the data held in the output register 109 in the register file unit 106 in the write FIFO memory 116. Then, the writing unit 114 sets the data stored in the FIFO memory 116 in the write buffer 115. The writing unit 114 reads the data from the buffer 115 and writes it to the preset address positions of the memories 1003a to 1003d.

[0030] The structure of the signal processor 1001 in the embodiment has been described above. The FIFO memory 103 is interposed between the readout unit 101 and the assignment unit 104 to absorb the difference in processing timing therebetween. The FIFO memory 103 has a capacity capable of storing image data for three pixels.

[0031] In contrast, there is no interface circuit for inputting and outputting data between the assignment unit 104, the first processing unit 110, and the second processing unit 110. Instead, the general-purpose shift register group 107 in the register file unit 106 is responsible for data transfer between the assignment unit 104, the first processing unit 110, and the second processing unit 110 (details will be described later).

[0032] FIG. 2 shows a circuit diagram of the general-purpose shift register group 107 in the register file unit 106.

[0033] Reference numeral 201 denotes a shift register located at the most upstream position in the series connection, and its operand is "R00." Reference numeral 214 denotes a shift register located second from the most upstream position in the series connection, and its operand is "R01." The operands of the subsequent serially connected registers are R02, R03, ..., R0f. Each register has the same structure.

[0034] Here, the shift register 201 will be described. Reference numeral 202 denotes a flip-flop that holds 32-bit data. Reference numeral 204 denotes an input signal that transfers data during a shift operation, and since the shift register 201 is the first register in the shift operation, 0 is input as the input signal. There is no particular limit to the number of bits that the flip-flop 201 holds, and it may be 8 bits.

[0035] Reference numeral 205 denotes the output of the flip-flop 202, and during shift operation, the output 205 is taken into a shift register 214 located downstream.

[0036] Reference numeral 210 denotes an input signal from the assignment unit 104 , 211 an input signal from the second processing unit 112 , 212 an input signal from the first processing unit 110 , and 213 an input signal from the CPU 1005 outside the signal processor 1001 .

[0037] Reference numeral 203 denotes a signal for inputting data into the flip-flop 202 and flip-flops in other shift registers during a shift operation. Reference numeral 206 denotes a signal output by the assignment unit 104 when writing an input signal 210 to the flip-flop 202. Reference numeral 207 denotes a signal output by the second processing unit 112 when writing an input signal 211 to the flip-flop 202. Reference numeral 208 denotes a signal output by the first processing unit 212 when writing an input signal 212 to the flip-flop 202. Reference numeral 209 denotes a signal output by the external CPU 1005 when writing an input signal 213 to the flip-flop 202.

[0038] When each of the reference signals 203, 206, 207, 208, and 209 is active, the switches controlled by the reference signals are connected to the lower side of the figure, and output the input signals 210, 211, 212, and 213 to the flip-flop 202. As can be seen from the figure, when the capture signals of the reference signals 203, 206, 207, 208, and 209 are not active, the switches controlled by the reference signals are connected to the upper side of the figure. When each of the reference signals 203, 206, 207, 208, and 209 is not active, the value of the flip-flop 202 returns to the flip-flop 202 again, so that the value of the flip-flop 202 is held even if a clock is provided.

[0039] Also, there is no circuit that arbitrates the data fetching process by the reference signals 203, 206, 207, 208, and 209. Therefore, when multiple reference signals become active at the same time, the data selected by the switch closest to the flip-flop 202 is fetched into the flip-flop 202. However, in the embodiment, multiple processing units are prevented from writing to one shift register at the same time.

[0040] In this way, in addition to the CPU 1005, the assignment unit 104, the first processing unit 110, and the second processing unit 112 can assign data to the shift register 201, and the shift register 201 can also reference (read out) data.

[0041] For example, when the assignment unit 104 assigns (writes) data from the FIFO memory 103 to the register 201, the assignment unit 104 can output the data to be written to the signal line 210 and activate the signal 206 for writing. When the first signal processing unit writes data to the register 201, the assignment unit 104 outputs the data to be written to the signal line 212 and activates the signal 208.

[0042] FIG. 3 is an internal block diagram of the first processing unit 110. Reference numeral 301 denotes a program counter indicating the position of an instruction to be executed in the program memory 111. Reference numeral 303 denotes an instruction selection multiplexer that extracts an instruction at the position indicated by the program counter 301 in the program memory 111. Reference numeral 304 denotes a register bus that bundles together the output signals of all the flip-flops of the register file unit 106. Reference numeral 305 denotes a multiplexer that selects a first source signal of an arithmetic input from the register bus 304. Reference numeral 306 denotes a multiplexer that selects a second source signal of an arithmetic input from the register bus 304. Reference numeral 307 denotes an ALU (Arithmetic and Logic Unit) that performs an arithmetic operation on two values ​​of the first source and the second source according to an execution instruction 308. Reference numeral 310 denotes an arithmetic operation result from the ALU 307. Reference numeral 309 denotes a binary decoder that activates a load signal 311 to any of the registers in the register file unit 106 in accordance with the execution instruction 308. Reference numeral 312 denotes an END detector that detects the last instruction "END" stored in the program memory 111. The END detector 312 outputs a signal 313 when it detects the END instruction. Upon receiving this signal 313, the program counter 301 stops updating the program counter (address) and sets the value to the beginning of the program. The END detector 312 also outputs the signal 313 to the shift control unit 117.

[0043] The above describes the structure of the first processing unit 110. The second processing unit 112 has the same structure as the first processing unit 110, and a description thereof will be omitted. The assignment unit 104 also has a program counter and an END detector, and when the END detector detects an END instruction, it stops updating the address and sets the value at the beginning of the program.

[0044] The shift control unit 117 can receive detection signals of the END command from each of the assignment unit 104, the first processing unit 110, and the second processing unit 112. When the shift control unit 117 receives detection signals of the END command from all of the assignment unit 104, the first processing unit 110, and the second processing unit 112, it activates a signal 203 (see FIG. 2) and issues a shift cycle signal. As a result, each shift register constituting the general-purpose shift register group 107 transfers data to the shift register one step downstream. In addition, the assignment unit 104, the first processing unit 110, and the second processing unit 112 receive the shift cycle signal and start processing the next cycle.

[0045] 4(a) and (b) show program instructions stored in the memory 111 of the first processing unit 110 (or the memory 113 of the second processing unit).

[0046] FIG. 4(a) shows the role of bits in a 19-bit instruction. An instruction is made up of four fields. The four bits from bit 18 to bit 15 form an opcode field 401 indicating the type of operation (addition, subtraction, logical operation, etc.), i.e., the operator. The five bits from bit 14 to bit 10 form a destination field 402 indicating the storage destination for storing the operation result. The five bits from bit 9 to bit 5 form a second source field 403 indicating the second source. And, bit 4 to bit 0 form a first source field 404 indicating the first source. In this way, each bit in an instruction has a determined role.

[0047] Figure 4(b) shows examples of assembler mnemonic notation for these instructions.

[0048] In the figure, reference numeral 405 denotes an opcode indicating "addition", reference numeral 406 denotes a destination register, and 407 and 408 denote reference registers. The opcode 405 corresponds to the opcode field 401 in FIG. 4(a), and the destination register 406 corresponds to the destination field 402. The reference register 407 corresponds to the second source field 403, and the reference register 408 corresponds to the first source field 404.

[0049] The operation when the first processing unit 110 executes the instruction of FIG. 4(b) is as follows. The first processing unit 110 causes the multiplexer 305 to select the shift register R00 in the register file unit 106, and the multiplexer 306 to select the shift register R01. Then, the ALU 307 adds the values ​​selected by the multiplexers 305 and 306 (the values ​​read from the shift registers R00 and R01, respectively) according to the execution instruction 308 (opcode "ADD" in FIG. 4(b)), and outputs the operation result 310. At this time, the decoder 309 selects the shift register "R02" of the destination field 402, and activates the signal 208. As a result, the addition result of the two values ​​held in the shift registers R00 and R01 is written to the shift register R02.

[0050] FIG. 9 shows the memory space of the IO bus 106 of the signal processor 1001 in this embodiment.

[0051] The data bus of the IO bus 106 is 32 bits and can be read and written from the CPU 1005. The shift registers R00, R01, ... constituting the general-purpose shift register group 107 are assigned from the base address 0x0000 indicated by reference numeral 901 to the address 0x003C indicated by reference numeral 902 (addresses beginning with "0x" are in hexadecimal notation). The registers R10, R11, ... constituting the general-purpose register group 108 are assigned from the address 0x0040 indicated by reference numeral 903. The output register 109 is assigned to the address 0x007C indicated by reference numeral 904.

[0052] Addresses 0x0100 to 0x01ff indicated by reference numeral 905 are assigned to the memory 111 of the first processing unit 110. Addresses 0x0200 to 0x02ff indicated by reference numeral 906 are assigned to the memory 113 in the second processing unit 112. Addresses 0x0300 to 0x03ff indicated by reference numeral 907 are assigned to the memory 102 of the readout unit 101. Addresses 0x0400 to 0x04ff indicated by reference numeral 908 are assigned to the memory 105 of the assignment unit 104. And, reference numeral 909 is assigned to a control register of the signal processor 1001. The control register includes, for example, a register that describes the timing of ending a series of processes. For example, it is possible to set the signal processor 1001 to stop when the processing results for the number of pixels in the horizontal direction are obtained.

[0053] 9, each register in the register file unit 106 and the memory storing the program of each program execution unit are mapped to different areas in the memory space of the IO bus 106, so that the CPU 1005 can freely access them. For example, when the CPU 1005 writes a program for the reading unit 101 in the signal processor 1001, the program can be written from the address 0x0300. The CPU 1005 can also write any value to any register in the register file unit 106 by specifying the corresponding address.

[0054] The memories 102, 105, 111, and 113 of the signal processor 1001 are SRAMs that can be accessed at high speed. SRAMs can be accessed at higher speeds than DRAMs, but their cost per unit capacity is higher than that of DRAMs. Since the memories 102, 105, 111, and 113 each store a relatively small program, the capacity of each memory can be small. Therefore, in this embodiment, SRAMs are used as the memories 102, 105, 111, and 113, enabling high-speed reading and writing of programs while suppressing an increase in cost. The memories 102, 105, 111, and 113 may be configured with flip-flops.

[0055] FIG. 11 shows a processing procedure relating to the start-up of the signal processor 1001 among the programs executed by the CPU 1005.

[0056] In step S1101, the CPU 1005 starts processing.

[0057] In S1102, the CPU 1005 reads out the programs (four in the embodiment) executed by the reading unit 101, the assignment unit 104, the first processing unit 110, and the second processing unit 112 in the signal processor 1001 from the memories 1003a to 1003d via the memory bus 1002. Then, the CPU 1005 writes the program executed by the reading unit 101 into the IO address space with 0x0300 as the top address according to the memory map of FIG. 9 shown above. That is, the CPU 1005 stores the program executed by the reading unit 101 into the memory 102. Similarly, the CPU 1005 writes the program for the assignment unit 104 into the memory 105. The CPU 1005 also writes the program for the first processing unit 110 into the memory 111. Then, the CPU 1005 writes the program for the second processing unit 112 into the memory 113. Furthermore, the CPU 1005 writes various data to the control registers and general-purpose registers of the signal processor 1001 as necessary.

[0058] In S1103, the CPU 1005 sets, to the reading unit 101, the start address of the image data in the address space of the memory bus 1002, and sets, to the writing unit 114, the start address at which post-filter processing pixel data is written.

[0059] In S1104, the CPU 1005 starts up the signal processor 1001.

[0060] Then, in S1105, the CPU 1005 waits for an interrupt signal indicating the end of processing from the signal processor 1001. When the CPU 1005 receives the interrupt signal indicating the end of processing, in S1106, the CPU 1005 ends this processing.

[0061] In this way, once the CPU 1005 starts the processing of the signal processor 1001, it does not need to adjust synchronization in the internal parallel operations, and it is sufficient to simply wait for the processing to finish. Note that the CPU 1005 may perform other processing during the waiting period.

[0062] Fig. 6 shows an example of a program executed by the signal processor 1001. Fig. 5 shows an equivalent circuit diagram when the signal processor 1001 executes the program shown in Fig. 6.

[0063] In addition, the image to be processed in this embodiment is composed of 640 pixels horizontally by 480 pixels vertically, and each pixel is represented by 32 bits. Usually, image data is often represented by 8 bits, and in that case, processing can be performed using 8 bits of the 32 bits. The data of each pixel constituting the image data to be processed is stored in raster scan order, starting from a specific address in the address space on the memory bus 1002.

[0064] 6 shows programs stored in memory 102 of reading unit 101. LIST 602 shows programs stored in memory 105 of assignment unit 104. LIST 603 shows programs stored in memory 111 of first processing unit 110. And LIST 603 shows programs stored in memory 113 of second processing unit 112.

[0065] The reading unit 101 reads image data from the top address position set by the CPU 1005 and stores it in the FIFO memory 103. After that, the reading unit 101 updates the read address by adding a value indicated in the program LIST 601 to the previously used address, and reads image data from the updated address position and stores it in the FIFO memory 103. This address update and storage process of image data in the FIFO memory 103 is repeated until the FIFO memory 103 becomes full. Free space is generated in the FIFO memory 103 when the downstream assignment unit 104 performs a process of acquiring data from the FIFO memory 103.

[0066] As described above, in the embodiment, image data is stored in raster scan order from a predetermined address position in the memories 1003a to 1003d. The size of the image data is 640 pixels in the horizontal direction. Therefore, the reading unit 101 updates the read address by offsetting the initial address by 640, 640, and -1279 as shown in LIST 601 and sequentially adding the offsets. Therefore, the reading unit 101 reads three pixels aligned in the vertical direction including the pixel in the upper left corner of the image data and stores them in the FIFO memory 103. Next, the reading unit 101 reads data of three pixels in the vertical direction at a position shifted by one pixel in the horizontal right direction and stores them in the FIFO memory 103. After that, the reading unit 101 repeats this process.

[0067] The assignment unit 104 inputs image data from the FIFO memory 103 and stores it in a shift register specified by the program shown in LIST 602. It should be understood that "LOAD R00" in LIST 602 is a command to "load (store) image data read from the FIFO memory 103 in (shift) register R00." In other words, the assignment unit 104 performs a process to store the top pixel data of three pixels aligned vertically in shift register R00, the middle pixel data in shift register R02, and the bottom pixel data in shift register R04. Then, since step 03 of LIST 602 is "END," the assignment unit 104 stops the program counter, returns the program counter to the top address of the program, and waits for a shift cycle signal to be issued. Then, by continuing the above process, the assignment unit 104 inputs data up to the right end of the image while maintaining the relationship between the three pixels in the vertical direction. When the assignment unit 104 reads image data from the FIFO memory 103, an empty area is generated in the FIFO memory 103. Therefore, when the FIFO memory 103 is full and the reading unit 101 has stopped reading, if the assignment unit 104 reads image data from the FIFO memory 103, the reading unit 101 will resume the reading process of the next image data.

[0068] Here, when the shift control unit 117 issues a shift cycle signal to the general-purpose shift register group 107 indicating that it is shift time, the general-purpose register file unit 106 transfers the data held in each shift register that constitutes the general-purpose shift register group 107 to the shift register located one step downstream.

[0069] The first processing unit 110 performs processing according to the program shown in LIST603. It should be understood that "ADD R06, R01, R03" in the program is a command to "add the values ​​of the shift registers R01 and R03, and store the result of the addition in the shift register R06." When the shift control unit 117 issues a shift cycle signal immediately before the first processing unit 110 starts processing according to the program shown in LIST603, the pixel data held in the shift register R00 is transferred to R01, the pixel data held in the shift register R02 is transferred to R03, and the pixel data held in the shift register R04 is transferred to R05. That is, the data of three pixels in the vertical direction stored in the shift registers R00, R02, and R04 by the readout unit 101 is transferred to the shift registers R01, R03, and R05. In addition, the first processing unit 110 re-executes the program of LIT603 from the beginning in response to receiving the shift cycle signal. When steps 00 to 02 of this LIST 603 are executed, ultimately, when the values ​​of three pixels from top to bottom in the vertical direction are P1, P2, and P3, the value P1+2×P2+P3 is stored in the shift register R06.

[0070] Then, since step 03 of LIST 603 is "END", the first processing unit 110 returns the processing to the beginning and waits for a shift cycle signal to be issued.

[0071] The second processing unit 112 performs processing according to the program shown in LIST 604. "SHIFT R1f, R10, R1d" in step 03 of LIST 604 should be understood as a command to "shift the value of register R10 by the number of bits indicated by the value stored in register R1d, and store the shift result in register R1f." The direction of bit shifting depends on the value of the third operand R1d; if it is positive, it is a left shift (shift toward the higher bits), and if it is negative, it is a right shift (shift toward the lower bits).

[0072] In the embodiment, the CPU 1005 stores the program of LSIT 604 in the memory 113 of the second processing unit 112 and assigns a value of "-4" to the shift register R1e. In other words, since it is not necessary to describe an instruction to store "-4" in the register R1e in LIST 604, it is possible to reduce the size of the program and improve the processing throughput.

[0073] When the shift control unit 117 issues a shift cycle signal immediately before the second processing unit 112 starts processing according to the program shown in LIST 604, the value stored in the shift register R06 by the first processing unit 110 in the immediately previous cycle is transferred to the shift register R07. The value stored in the shift register R06 by the first processing unit 110 two cycles ago is stored in the shift register R08. And the value stored in the shift register R06 by the first processing unit 110 three cycles ago is stored in the shift register R09.

[0074] Therefore, when the second processing unit 112 executes steps 00 to 02 of LIST604, the shift register R10 stores the value of P1+2×P2+P3 when the three pixels in the horizontal direction are P1, P2, and P3. The first processing unit 110 performed the calculation for three pixels arranged in the vertical direction. Therefore, when the second processing unit 112 executes the instruction of step 03 of LIST604, the output register R1f stores the value of the shift register R10 shifted right by four bits (equivalent to division by 16). That is, the output register R1f stores pixel data after the filter process at the center position of the 3×3 pixel block. Since step 04 of LIST604 is "END", the second processing unit 112 returns the program counter to the beginning and waits for the next shift cycle signal to be issued.

[0075] Now, with reference to the timing chart of FIG. 7, the synchronization between the readout unit 101, the assignment unit 104, the twelfth processing unit 110, and the second processing unit 112 in the signal processor 1001 will be described.

[0076] Reference numeral 701 in the figure indicates the position of a pixel corresponding to the image data read by the reading unit 101 and stored in the FIFO memory 103. 0, 640, 1280, ... shown in the figure indicate offset addresses of the image data rather than values ​​of the image data.

[0077] Reference numeral 702 denotes image data written by the assignment unit 104. In the figure, R00, R02, and R04 denote destination registers for the image data, rather than image data values. One cycle later than the reading unit 101, the assignment unit 104 writes data at offset address 0 to shift register R00, data at offset address 640 to shift register R02, and data at offset address 1280 to shift register R04.

[0078] Reference numeral 704 denotes each step performed by the first processing unit 110, and reference numeral 705 denotes each step performed by the second processing unit 112.

[0079] Also, reference numeral 703 denotes a shift cycle signal. The shift control unit 117 issues the shift cycle signal 703 when all the program execution units, ie, the assignment unit 104, the first processing unit 110, and the second processing unit 112, return their respective program counters to the top of their own memories (when the END command is detected). As a result, each of the processing units, ie, the assignment unit 104, the first processing unit 110, and the second processing unit 112, executes the program in its own memory again.

[0080] This mechanism enables the assignment unit 104, the first processing unit 110, and the second processing unit 112 to operate in parallel, and also enables data transfer between them without the need for an interface circuit.

[0081] This will be explained in detail with reference to Figure 8. Figure 8(a) is a diagram for explaining the problem that occurs when an assigning program and a referencing program that refers to the assigned data run in parallel.

[0082] If the program that refers to the data overtakes the program that refers to the data after the program that refers to the data substitutes the third pixel into registers 801 and 802, register 803 will end up referring to the information of the second pixel.

[0083] 8(b) shows a configuration in which a double buffer system (registers 801 and 804, registers 802 and 805, registers 803 and 806) is used to avoid such a situation, and the referencing program can only refer to the second pixel until the assigning program finishes assigning the third pixel. After the assigning program finishes writing the third pixel to registers 801 to 803 and the referencing program finishes referencing the second pixel from registers 804 to 806, the data is transferred in one cycle.

[0084] The configuration in FIG. 8(b) is a common configuration, but it requires twice as many flip-flops, resulting in a large circuit scale.

[0085] In this embodiment, a group of shift registers is used as shown in Fig. 8(c). As a result, when the assignment unit 104, the first processing unit 110, and the second processing unit 112 all finish one cycle of processing, the shift control unit 117 issues a shift cycle signal. The group of general-purpose shift registers 107 is used for calculations, and depending on the application, it is possible to allocate more to the interface between programs, or conversely, to circuit calculations. Therefore, it is possible to configure a programmable signal processing circuit with a wide range of application using a small circuit.

[0086] The first embodiment has been described above, and its features can be summarized as follows.

[0087] The programmable signal processor includes a group of serially connected general-purpose shift registers, a plurality of program execution units (represented by an assignment unit, a first processing unit, and a second processing unit) that can be executed in parallel with each other and use the group of general-purpose shift registers, and a shift control unit that generates a shift request to the shift registers that constitute the group of general-purpose shift registers and a shift cycle signal for requesting each program execution unit to start processing for one cycle. In this configuration, when one program execution unit passes data to another program execution unit, the data is stored in a shift register located one upstream of the shift register used by the other program execution unit as an input. When one program execution unit passes multiple pieces of data to another program execution unit, the data is stored in multiple shift registers located at discrete positions, with at least one shift register in between. When each program execution unit detects an instruction indicating the end of processing for one cycle, the program execution unit ends its operation and sets the program counter to an initial position in preparation for the next cycle. Then, the shift control unit issues a shift cycle signal upon detection of an instruction to end processing for one cycle from all program execution units.

[0088] The reading unit 101 does not use the general-purpose shift register group 107. Also, the program executed by the reading unit 101 does not include an END command. This is because the reading unit 101 only needs to stop reading when the FIFO memory 103 is full and to perform reading when the FIFO memory 103 is not full.

[0089] With the above configuration, the program execution units capable of operating in parallel with each other can transfer data to each other and perform given filter processing by utilizing a simple configuration of a group of serially connected shift registers.

[0090] [Second embodiment] The second embodiment will be described below. The device configuration is assumed to be the same as that of the first embodiment shown in Fig. 10, and the description thereof will be omitted.

[0091] 12 is a circuit diagram of a programmable signal processor 1001 according to the second embodiment. The same components as those in FIG. 1 are denoted by the same reference numerals.

[0092] In the first embodiment, four program execution units, namely, the read unit 101, the assign unit 104, the first processing unit 110, and the second processing unit 112, operate in parallel. In general, the more the number of program execution units, the more the processing speed of the signal processor can be improved, but the circuit scale also becomes correspondingly more complex. In the second embodiment, there is no program execution unit corresponding to the second processing unit 112, and the number of program execution units operating in parallel is reduced, thus simplifying the circuit scale. Moreover, the processing of the read unit 101, the assign unit 104, and the write unit 114 is the same as in the first embodiment.

[0093] 13 shows a list of programs stored in the memory 111 of the first processing unit 110 in the second embodiment. The first processing unit 110 in the second embodiment executes programs including those executed by the second processing unit 112 in the first embodiment, so the number of program steps increases. Since the throughput of the signal processor 1001 depends on the maximum number of program steps executed by the program execution unit, the configuration of the second embodiment has a lower throughput than the first embodiment, but the circuit scale can be reduced.

[0094] [Third embodiment] The third embodiment will be described below. The device configuration is assumed to be the same as that of the first embodiment shown in Fig. 10, and the description thereof will be omitted.

[0095] 14 is a circuit diagram of a programmable signal processor 1001 according to the third embodiment. The same components as those in FIG. 1 are denoted by the same reference numerals.

[0096] A feature of the signal processor 1001 in the third embodiment is that an interrupt signal generating register 1400 is added as shown in Fig. 14. This interrupt signal generating register 1400 may be newly added, but here, a register R1E, ​​which is one of the general-purpose registers 108, is used. When "1" is written to bit 0 of the register R1E, ​​the signal processor 1001 outputs an interrupt signal to the outside (for example, the CPU 1005) and transitions to a suspended state, stopping until it is resumed from the outside by a control register.

[0097] The CPU 1005 starts interrupt processing in response to this interrupt signal. As a result, exceptionally complex processing by the signal processor 1001 can be offloaded to the CPU 1005, dramatically increasing the range of application.

[0098] In addition, when the CPU 1005 executes processing such as scanning a data string that is not signal processing-related and detecting a marker code, there is no overhead such as a loop jump instruction or a counter increment instruction, so that the total throughput including the signal processor 1001 can be structured.

[0099] [Fourth embodiment] A fourth embodiment will be described. Fig. 15 is a block diagram of an apparatus according to the fourth embodiment. In Fig. 15, a GPGPU unit (General-purpose computing on GPU (Graphics Processing Units)) 1501 is added to the configuration of Fig. 10.

[0100] In recent years, it has become common for image capture devices, such as digital cameras, to perform autofocus using signals obtained from the image capture surface. However, there is a problem in that converting a pupil-split signal to a defocus amount depends on the optical characteristics of the sensor and the optical characteristics of the lens, and the characteristics change significantly depending on the coordinates on the image. In the case of an image capture device with interchangeable lenses, the conditions become even more complicated.

[0101] Such complicated optical calculations use floating-point polynomials. In view of this, in the fourth embodiment, the defocus amount conversion coefficient is calculated by the GPGPU unit 1501. However, the side that uses the defocus amount conversion coefficient only requires information of about 8 integer bits, and the data needs to be processed in order to reduce the amount of data transfer.

[0102] Therefore, in the fourth embodiment, a programmable signal processor 1001 is used to process the floating-point arithmetic results.

[0103] Figure 17 shows the bit assignment of floating-point numbers. Bits are assigned to the sign part (positive and negative signs), exponent part, and mantissa part, respectively, and in this state it is difficult to perform autofocus calculations at high speed.

[0104] 18 is an equivalent circuit diagram of the signal processor 1001 in the fourth embodiment. Reference numeral 1800 corresponds to the assignment unit 104. A register 1802 holds floating-point data assigned by the assignment unit 104. A shift operation transfers data from the register 1802 to the register 1803. The CPU 1005 starts the signal processor 1001 after assigning fixed values ​​to the registers R10 to R16 in advance. The values ​​set in the registers R10 to R16 are as shown in the program list in the figure. Note that the registers R10 to R1F do not have a shift function and therefore continue to hold the same values.

[0105] FIG. 16 shows a list of programs stored in the memory 111 of the first processing unit 110 that simulates the circuit diagram of FIG.

[0106] In step 00, the first processing unit 110 shifts the floating point of R01. Then, in step 01, the first processing unit 110 extracts the exponent part by performing a mask process. Then, in step 02, the first processing unit 110 subtracts an offset for shifting the mantissa part.

[0107] In step 03, the first processing unit 110 masks the floating-point data held in the register R01 with a value preset in the register R13, and extracts the mantissa part. Then, in step 04, the first processing unit 110 adds the most significant bit of the mantissa part. This is because the most significant bit of the mantissa part must always be 1 in IEEE754.

[0108] In step 05, the first processing unit 110 converts the mantissa into an integer by shifting it by the value held in the register R02. Then, in step 06, the first processing unit 110 subtracts the register R03 from the value "0" preset in the register R15 to create a negative value.

[0109] In step 07, the first processing unit 110 obtains the sign of the floating point by performing a logical AND on the register R01 holding the floating point data and the value previously held in the register R16. Then, in step 09, the first processing unit 110 selects a positive value or a negative value depending on the sign, and writes the result to the output register R1f.

[0110] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention.

[0111] (Other embodiments) The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

[0112] The disclosure of this specification includes the following programmable signal processing circuit and a program for the circuit. (Item 1) 1. A programmable signal processing circuit, comprising: A plurality of execution units each having a program memory and operable in parallel by executing a program stored in the program memory; a register file unit including a plurality of serially connected registers, each of which is available to the plurality of execution units, and which transfers data held in each of the plurality of registers to a register located one register downstream in the serial connection in response to receiving a shift signal; a control unit that writes the program into a program memory in the plurality of execution units, and issues the shift signal to the plurality of execution units when each of the plurality of execution units has finished executing one cycle of the program; the control unit stores the program in a program memory in the execution units, the program including an instruction for specifying, when each of the execution units passes data to another execution unit, a register located immediately upstream of a register referenced when the other execution unit inputs data, as a storage destination for data resulting from the processing in one cycle; Each of the execution units executes the program again in response to receiving the shift signal after executing the program for one cycle. 1. A programmable signal processing circuit comprising: (Item 2) When the execution unit passes a plurality of data to another execution unit, the data is stored in a plurality of registers at discrete positions with at least one register between them in the series connection relationship. 2. The programmable signal processing circuit according to item 1, (Item 3) One of the plurality of execution units is an assignment unit that assigns data to be processed to the register, the programmable signal processing circuit further includes a read unit that executes a program to read the data to be processed from a memory and supplies the read data to the assignment unit via a FIFO memory; The control unit stores a program that specifies the data to be processed in a program memory included in the reading unit. 3. The programmable signal processing circuit according to item 1 or 2. (Item 4) The reading unit stops reading from the external memory when the FIFO memory becomes full, and resumes reading when a free space is generated in the FIFO memory. 4. The programmable signal processing circuit according to item 3, (Item 5) Each of the execution units has a detector for detecting an instruction indicating the end of one cycle of processing; When the detector detects an instruction indicating an end, the detector stops a program counter and sets a start address of the program in the program counter; The control unit issues the shift signal in response to all of the detectors of the plurality of execution units detecting an instruction indicating an end. 5. A programmable signal processing circuit according to any one of items 1 to 4. (Item 6) the register file unit includes an output register for storing final processing result data; The control unit stores a program including an instruction to store processed data in the output register in a program memory of one of the execution units. 6. A programmable signal processing circuit according to any one of items 1 to 5. (Item 7) The program memories of each of the execution units and the registers of the register file unit are mapped to different address spaces in the memory space of a predetermined IO bus. 7. A programmable signal processing circuit according to any one of items 1 to 6. (Item 8) the register file unit has a plurality of general-purpose registers that hold data regardless of receipt of the shift signal; When a preset bit position of one of the plurality of general-purpose registers is set to "1", the programmable signal processing circuit is put into a suspend state and an interrupt signal is output to the outside. 8. A programmable signal processing circuit according to any one of items 1 to 7. (Item 9) The control unit reads out a program to be executed by each of the plurality of execution units from a memory in which the program is stored, and stores the program in the program memory of each of the plurality of execution units via an IO bus. 9. A programmable processing circuit according to any one of items 1 to 8. (Item 10) A plurality of execution units capable of executing programs; wherein each of the plurality of execution units re-executes the program in response to receiving a predetermined shift signal; a register file unit having a plurality of serially connected registers, each of the plurality of execution units being available; Here, in response to receiving the shift signal, the register file unit transfers the data held in each of the plurality of registers to a register located one register downstream in the series connection; an issuing unit that issues the shift signal when each of the plurality of execution units has completed execution of one cycle of the program; A program for a programmable signal processing circuit having The program executed by each of the plurality of execution units is In order to pass data to another execution unit, the instruction includes an instruction to store the data resulting from the processing in the one cycle in a register located one step upstream of the register that the other execution unit references when inputting data. A program for a programmable signal processing circuit.

[0113] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0114] 1001...signal processor, 1002...memory bus, 1003a to 1003d...memories, 1005...CPU, 1006...IO bus, 101...reading unit, 104...assignment unit, 106...register file unit, 107...general-purpose shift register group, 108...general-purpose register group, 109...output register, 110...first processing unit, 112...second processing unit, 114...writing unit, 102, 105, 111, 113...memories, 117...shift control unit

Claims

1. A processing device having a CPU and a programmable signal processing circuit, wherein the programmable signal processing circuit has a program memory respectively, and a plurality of execution units capable of operating in parallel by executing a program stored in the program memory; a register file unit having a plurality of registers that are respectively available to the plurality of execution units and are serially connected, and in response to a shift cycle signal, transferring data held by each of the plurality of registers to a register located one downstream in the serial connection; and a shift control unit that issues the shift cycle signal when each of the plurality of execution units finishes executing a one-cycle program, and the CPU stores, in the program memory in the plurality of execution units, a program including an instruction to use, as a storage destination for data resulting from execution of the one-cycle program by the first execution unit among the plurality of execution units, a register located upstream of a register that the second execution unit refers to for inputting data so that the data resulting from execution of the one-cycle program by the first execution unit is transferred to the second execution unit among the plurality of execution units using the register file unit. A processing device characterized by the above.

2. Each of the plurality of execution units stops program execution when it finishes executing the one-cycle program, and resumes program execution in response to the shift cycle signal. The processing device according to claim 1, characterized by the above.

3. One of the plurality of execution units is an assignment unit that assigns data to be processed to the register, and the programmable signal processing circuit has a reading unit having a program memory. The reading unit reads the data to be processed from an external memory by executing a program stored in the program memory of the reading unit, and supplies the read data to the substitution unit. The CPU stores a program for designating the data to be processed in the program memory of the reading unit. The processing apparatus according to claim 1, characterized in that.

4. The reading unit supplies the data to be processed to the substitution unit via a FIFO memory. When the FIFO memory becomes full, the reading unit stops reading the data to be processed from the external memory, and resumes reading when a free area occurs in the FIFO memory. The processing apparatus according to claim 3, characterized in that.

5. Each of the plurality of execution units has a program counter and a detector for detecting an instruction indicating the end of a program for one cycle. The program counter stops updating the count value in response to a detection signal from the detector and sets the start address of the program. The shift control unit issues the shift cycle signal when receiving the detection signals from all the detectors of the plurality of execution units. The processing apparatus according to claim 1, characterized in that.

6. The register file unit has an output register for storing data of a result of a predetermined process by the programmable signal processing circuit. The CPU stores a program including an instruction for storing the data obtained by executing the program in the output register in the program memory of the execution unit that outputs the data of the result of the predetermined process among the plurality of execution units. The processing apparatus according to claim 1, characterized in that.

7. The program memory of each of the plurality of execution units and each register of the register file unit are mapped to different address spaces in the memory space of a predetermined IO bus. The processing device according to claim 1, characterized in that.

8. The register file unit has a plurality of general-purpose registers that hold data regardless of the reception of the shift cycle signal. When a predetermined value is set in a predetermined one of the plurality of general-purpose registers, the programmable signal processing circuit shifts to a suspended state, and an interrupt signal is output to the outside of the programmable signal processing circuit. The processing device according to claim 1, characterized in that.

9. The register file unit has a plurality of general-purpose registers that hold data regardless of the shift cycle signal. The CPU sets a predetermined value in the general-purpose register and stores a program including an instruction that uses the general-purpose register in the program memory of the plurality of execution units. The processing device according to claim 1, characterized in that.

10. The CPU reads out the programs to be executed by each of the plurality of execution units from the memory in which the programs are stored, and stores the read programs in the program memories of each of the plurality of execution units. The processing device according to any one of claims 1 to 9, characterized in that.

11. The CPU executes the program stored in the memory and controls the processing device. The processing device according to claim 10, characterized in that.

12. A CPU and A programmable signal processing circuit, A plurality of execution units each having a program memory and capable of operating in parallel by executing the program stored in the program memory. A register file unit having a plurality of serially connected registers that can be used by each of the plurality of execution units, the register file unit transferring data held by each of the plurality of registers to a register located one downstream in the serial connection in response to a shift cycle signal. A shift control unit that issues the shift cycle signal when each of the plurality of execution units finishes executing a one-cycle program. A programmable signal circuit having the above. A program for a processing device having the above. The CPU is caused to store, in the program memory in the plurality of execution units, a program including an instruction to use, as a storage destination for data resulting from execution of the one-cycle program by the first execution unit among the plurality of execution units, a register located upstream of a register that the second execution unit among the plurality of execution units refers to for inputting data, such that the data resulting from execution of the one-cycle program by the first execution unit is transferred to the second execution unit using the register file unit. A program for a processing device for operating as described above.