Data processing apparatus and data processing program

The data processing apparatus and program synchronize SIMD type microprocessor elements using flags and counters to manage varying loop sizes, preventing idle states and enhancing processing efficiency.

JP2025112894APending Publication Date: 2025-08-01DENSO CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024007419
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-22
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In SIMD type microprocessors, when the sizes of target data processed by each processing element differ, some elements enter an idle state waiting for others to complete processing, leading to inefficiency.

Method used

A data processing apparatus and program that includes an execution counter and flag generation unit to manage processing elements based on loop sizes, allowing elements to start processing next data even if loop sizes differ, using flags to synchronize operations.

Benefits of technology

This approach prevents processing elements from idling, ensuring continuous operation and reducing overall processing time by aligning start times and maintaining parallel processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025112894000001_ABST
    Figure 2025112894000001_ABST
Patent Text Reader

Abstract

To provide a data processing apparatus which can prevent a processing element from falling into an idle state.SOLUTION: A microcomputer board or an SoC board functioning as the data processing apparatus comprises a plurality of processing elements 70. Processing executed by the processing elements 70 is counted by an execution counter 82. When loop sizes LS of target data being processing targets of the processing elements 70 are different by the processing elements 70, flags 84 associated with the processing elements 70 respectively are generated on the basis of the loop sizes LS and the execution counter 82.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosure according to this specification relates to data processing technology.

Background Art

[0002] Patent Document 1 discloses a SIMD (Single Instruction, Multiple Data) type microprocessor having a plurality of processing elements.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a SIMD type microprocessor as disclosed in Patent Document 1, a plurality of processing elements execute parallel processing integrally. Therefore, when the sizes of the target data processed by each processing element are different from each other, some of the processing elements that have finished processing the target data are likely to be in an idle state in order to wait for the completion of the processing of other processing elements.

[0005] An object of the present disclosure is to provide a data processing apparatus and a data processing program capable of suppressing a processing element from entering an idle state.

Means for Solving the Problems

[0006] To achieve the above object, one disclosed aspect is a data processing apparatus including a plurality of processing elements (70), an execution counter (82) that counts the processes executed by the processing elements, and a flag generation unit (81) that generates a flag (84) associated with each of the processing elements based on the loop size (LS) of target data (90) to be processed by the processing elements when the loop size is different for each processing element.

[0007] Another disclosed aspect is a data processing program including instructions to be executed by a processor (60) including a plurality of processing elements (70), the instructions including a counter process that counts the processes executed by the processing elements, and a flag generation process that generates a flag (84) associated with each of the processing elements based on the count by the counter process and the loop size when the loop size (LS) of target data (90) to be processed by the processing elements is different for each processing element.

[0008] In these aspects, a flag associated with each processing element is generated based on the loop size of the target data and the execution counter or the count in the counter process. Therefore, even if the loop sizes of the target data processed by each processing element are different from each other, based on the flag, the processing element that has finished processing the target data can be made to start processing the next target data. As a result, it is possible to suppress the processing elements from entering an idle state.

[0009] Note that the reference numerals in parentheses in the above and in the claims etc. merely show an example of the correspondence with the specific configurations in the embodiments described later, and do not limit the technical scope in any way. Also, combinations of claims not explicitly stated in the claims are possible as long as there is no problem with the combination.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Mode for Carrying Out the Invention

[0011] The acoustic sensing system 10 according to the present disclosure is used in a moving body such as a vehicle Ve shown in FIG. 1. The acoustic sensing system 10 recognizes the driving environment around the vehicle Ve by measuring the external sound arriving at the vehicle Ve. The acoustic sensing system 10 is mounted on a vehicle Ve equipped with an automatic driving function together with other in-vehicle sensing systems such as a camera system, a radar system, and a sonar system.

[0012] The acoustic sensing system 10 detects the approach of an emergency vehicle, including a police vehicle, a fire truck, an ambulance, etc., to the vehicle Ve by voice recognition technology. The emergency vehicle recognized by the acoustic sensing system 10 is an emergency vehicle that is traveling for an emergency mission and sounding a siren. In addition, the acoustic sensing system 10 may be able to recognize abnormal sounds emitted by a breakdown vehicle, noises emitted by an illegally modified vehicle, etc., voices of children, screams, strange sounds, gunshots related to a crime, and area broadcasts by the nationwide instantaneous warning system (J-Alert).

[0013] <Configuration of the Acoustic Sensing System> As shown in FIGS. 1 to 3, the acoustic sensing system 10 includes a microphone 20 and at least one of a microcontroller board 30 and a SoC (System on a Chip) board 50. The microcontroller board 30 and the SoC board 50 have the function of a sound data processing device that can particularly process sound data.

[0014] The microphone 20 is, for example, a condenser microphone, and outputs an electrostatic capacitance change generated by a thin diaphragm vibrating due to sound pressure as an electrical signal. The microphone 20 is provided with, for example, MEMS (Micro Electro Mechanical Systems) or the like as a microphone element that converts air vibration into an electrical signal. A piezo element may be adopted as a microphone element in place of MEMS for the microphone 20.

[0015] The microphone 20 is held by the external structure of the vehicle Ve with the sound collecting surface of the microphone element facing the external structure. The microphone 20 functions as an external acoustic sensor. The microphone 20 is provided on the front, rear, left and right side surfaces, and upper surface of the vehicle Ve (see Fig. 1). The microphone mounting position Ps1 on the front of the vehicle is, for example, the front emblem, the front camera module, and the front millimeter-wave radar, etc. The microphone mounting position Ps2 on the rear of the vehicle is, for example, the rear camera module, the back door, and the rear surface of the rear hatch, etc. The microphone mounting position Ps3 on the side surface of the vehicle is, for example, the side mirror, the door, each pillar, and the side camera, etc. The microphone mounting position Ps4 on the upper surface of the vehicle is, for example, the roof panel, the upper surface of the trunk lid, the sunroof, the roof rail, the ADAS sensor module, and the antenna module, etc. The microphone 20 may be installed at a plurality of locations among these microphone mounting positions Ps1 to Ps4, or may be installed at only one location.

[0016] The microcomputer board 30 (see Fig. 2) is an arithmetic device electrically connected to the microphone 20. The microcomputer board 30 is arranged in the vicinity of the microphone 20 in the vehicle Ve. The microcomputer board 30 may be integrally configured with the microphone 20. A microcontroller or a microcomputer (hereinafter, the microcomputer 31) and a memory 32 are mounted on the microcomputer board 30. The microcomputer 31 is an arithmetic processing unit combined with the memory 32. The memory 32 functions as a RAM and is used as an arithmetic area of the microcomputer 31. The electrical signal generated by the microphone 20, that is, the sound data of the external sound measured by the microphone 20 is sequentially input to the microcomputer 31. The microcomputer 31 acquires the digitized sound data. The circuit part (AD converter) for digitizing the sound data may be provided in the microphone 20, or may be provided in the microcomputer board 30 or the microcomputer 31.

[0017] The SoC board 50 (see Fig. 3) is an arithmetic unit electrically connected to the microphone 20. The SoC board 50 may be connected to a plurality of microphones 20. The SoC board 50 may be provided as an ECU (Electronic Control Unit) dedicated to acoustic sensing in the vehicle Ve, or may be provided as a part of another in-vehicle ECU. Other in-vehicle ECUs are, for example, a driving support ECU, an autonomous driving ECU, an HMI (Human Machine Interface)-ECU, a zone ECU, a central ECU, and a gateway ECU, etc. An SoC 50a is mounted on the SoC board 50, and an arithmetic processing circuit mainly composed of the SoC 50a is formed.

[0018] The SoC 50a includes a CPU 51, a DSP 52, a GPU 53, a memory interface 54, an external interface 55, and a bus 56 connecting these. The CPU (Central Processing Unit) 51, the DSP (Digital Signal Processor) 52, and the GPU (Graphics Processing Unit) 52 function as arithmetic processing units. The memory interface 54 is connected to a memory (RAM) mounted on the SoC board 50. The memory interface 54 causes the memory to function as an arithmetic area for the CPU 51, the DSP 52, and the GPU 53. The external interface 55 is a functional unit that relays data between external devices of the SoC board 50 and each arithmetic processing unit. Sound data of external sound measured by the microphone 20 is sequentially input to the external interface 55. The external interface 55 may have a circuit section for digitizing the sound data. The external interface 55 provides the digitized sound data to at least one of the CPU 51 and the DSP 52 via the bus 56.

[0019] <Configuration of SIMD-Type Processor> In at least one of the above-mentioned microcomputer 31, CPU 51, and DSP 52, an SIMD (Single Instruction, Multiple Data) type processor 60 shown in FIG. 4 is adopted. The SIMD type processor 60 can parallel-process a plurality of data series with one instruction. The SIMD type processor 60 is composed of, for example, a global processor 61, a register 63, an arithmetic array 69, and the like.

[0020] The global processor 61 is an SISD (Single Instruction, Single Data) type processor with a built-in program RAM and data RAM. In at least one of the program RAM and the data RAM, a program (data processing program) for executing voice recognition processing described later is stored. The global processor 61 performs various arithmetic processes using built-in general-purpose registers and an ALU (Arithmetic and Logic Unit) and the like by executing instructions based on the decoded program. In addition, the global processor 61 generates various control signals to be supplied to the register 63 and the arithmetic array 69.

[0021] The register 63 holds data to be processed by the arithmetic array 69. The register 63 includes general-purpose registers and accumulator registers. The register 63 is an aggregate of a large number of general-purpose registers and accumulator registers. The data read from the register 63 is provided to the arithmetic array 69. The arithmetic result in the arithmetic array 69 is written into the register 63. The reading of data from the register 63 and the writing of data into the register 63 are controlled by the global processor 61.

[0022] The arithmetic array 69 consists of a plurality (a large number) of processing elements 70. The arithmetic processing by each processing element 70 is controlled by the global processor 61. The number of data stored in the general-purpose register and the accumulator register provided in the register 63 corresponds to the number of processing elements 70 included in the arithmetic array 69. In other words, individual general-purpose registers and accumulator registers are individually associated with the individual processing elements 70. Based on an instruction, the processing element 70 reads the data to be processed stored in the corresponding general-purpose register and accumulator register, and executes an operation using the read value as an operand. The value resulting from the operation is written into the general-purpose register or the accumulator register.

[0023] <Application of SIMD-Type Processor to Acoustic Recognition Processing> The SIMD-type processor 60 performs the acoustic recognition processing shown in FIG. 5. The SIMD-type processor 60 processes audio data by a plurality of processing elements 70 in the acoustic recognition processing. The SIMD-type processor 60 starts the acoustic recognition processing based on the activation of the microcontroller board 30 or the SoC board 50. The SIMD-type processor 60 repeatedly performs the acoustic recognition processing until the operation of the microcontroller board 30 or the SoC board 50 stops.

[0024] At S10, the SIMD-type processor 60 applies a discrete Fourier transform (DFT) or a fast Fourier transform (FFT) to the audio data. The SIMD-type processor 60 decomposes the input audio data into a series of frequency components and calculates the strength of each component. As a result, the audio data, which was a signal in the time domain, is converted into data in the frequency domain. Specifically, the SIMD-type processor 60 generates a complex matrix F (see FIG. 6) based on the audio data.

[0025] The SIMD-type processor 60 generates a spectrogram at S20. The SIMD-type processor 60 calculates in parallel the absolute values of a plurality (e.g., 4 in parallel) of elements arranged in the column direction of the complex matrix F by a plurality (4) of processing elements 70. The number of processing elements 70 to be parallelized may be appropriately changed according to the configuration of the SIMD-type processor 60. Each parallelized processing element 70 repeats the calculation of the absolute value for the SIMD-parallel (4) elements while moving the calculation target one row at a time (see arrow 1 in FIG. 6). When each processing element 70 finishes the calculation of the last row of the complex matrix F, it moves the calculation target only by the SIMD parallel in the column direction. After each processing element 70 shifts the calculation target in the column direction, it repeats the calculation of the absolute value again while moving the calculation target one row at a time from the first row (see arrow 2 in FIG. 6). The SIMD-type processor 60 generates a mel spectrogram matrix S (see FIG. 7) corresponding to the spectrogram of the audio data by the process of calculating the absolute value of each element of the complex matrix F.

[0026] The SIMD-type processor 60 generates a mel spectrogram at S30. The mel spectrogram is obtained by applying a mel scale that absorbs the difference between the actual sound and the human pitch perception to the spectrogram. In the mel spectrogram, the horizontal axis is the time axis, the vertical axis is the frequency, and the value is the power (energy) of the sound. The SIMD-type processor 60 multiplies each column of the mel spectrogram matrix S generated at S20 by the 0th to C-1st rows of the transformation matrix T in order (see FIG. 7). The transformation matrix T corresponds to a filter for applying the mel scale. Each row of the transformation matrix T has coefficients such that the low-frequency components are sampled more densely and the high-frequency components are sampled more sparsely. By applying the mel scale, the human characteristic of being sensitive to the frequency difference in the low-frequency region and insensitive to the frequency difference in the high-frequency region is reflected in the acoustic recognition.

[0027] The SIMD-type processor 60 detects a recognition target using a Deep Neural Network (DNN) at S40. The DNN is pre-generated by deep learning (deep neural network learning) using a large amount of sample data of the recognition target. The SIMD-type processor 60 inputs the mel spectrogram generated at S30 into the DNN and obtains an identification result of a preset recognition target. The identification result of the recognition target by the DNN is output from the microcontroller board 30 or the SoC board 50 to other in-vehicle ECUs.

[0028] <Details of Application of SIMD-Type Processor to Mel Spectrogram Generation> In the speech recognition process described so far, in the inference using DFT or FFT, spectrogram generation, and DNN, it is necessary to apply the same operation to a large number of data and process it at high speed. Therefore, it is useful to use the SIMD-type processor 60 capable of performing parallel processing using a plurality of processing elements 70.

[0029] On the other hand, in the process of generating a mel spectrogram, as described above, the sampling widths for the data of specific columns of the mel spectrogram matrix S are different. That is, when parallelizing the number of rows for SIMD parallel processing, the loop size LS of the target data 90 to be processed by each processing element 70 differs for each processing element 70. Specifically, the loop size LS of the 0th row is the smallest (shortest), and as the row number increases, such as the 1st row, 2nd row, 3rd row..., the loop size LS gradually becomes larger (longer) (see Fig. 7).

[0030] Here, it is substantially impossible to individually control the processing of a plurality of parallelized processing elements 70. In other words, each processing element 70 cannot start processing on the next target data 90 (lines 4 to 7) until all the processing elements 70 have completed the processing on the target data 90 (lines 0 to 3) assigned to each of them (see Fig. 8). Therefore, the processing element 70 that has finished processing the target data 90 on line 0 with a short loop size LS will be in an idle state at least until the other processing elements 70 complete the processing on the target data 90 on line 3 with a long loop size LS.

[0031] To suppress such an idle state of the processing element 70, the SIMD type processor 60 sets a plurality of first coefficient columns 91 and a plurality of second coefficient columns 92 corresponding to each SIMD parallel part as a set of target data 90 for parallel processing (see Figs. 9 and 10). After one processing element 70 finishes processing the first coefficient column 91, it starts processing the assigned second coefficient column 92 without waiting for the other processing elements 70 to finish processing the first coefficient column 91.

[0032] As an example, as shown in Fig. 9, for the processing element 70 with PE number [0], the data on line 0 is assigned as the first coefficient column 91, and the data on line 4 is assigned as the second coefficient column 92. For the processing element 70 with PE number [1], the data on line 1 is assigned as the first coefficient column 91, and the data on line 5 is assigned as the second coefficient column 92. For the processing element 70 with PE number [2], the data on line 2 is assigned as the first coefficient column 91, and the data on line 6 is assigned as the second coefficient column 92. And for the processing element 70 with PE number [3], the data on line 3 is assigned as the first coefficient column 91, and the data on line 7 is assigned as the second coefficient column 92. In the above aspect, the order of assigning the plurality of second coefficient columns 92 to each processing element 70 is the same as the order of assigning the plurality of first coefficient columns 91 to each processing element 70.

[0033] As another example, in the embodiment shown in FIG. 10, the plurality of second coefficient columns 92 are assigned to each processing element 70 in the reverse order of the order in which the plurality of first coefficient columns 91 are assigned to each processing element 70. That is, in the processing element 70 with the PE number [3], each second coefficient column 92 is allocated to each processing element 70 in a folded-back manner. As described above, in the processing element 70 with the PE number [0], the data in the 0th row is assigned as the first coefficient column 91, and the data in the 7th row is assigned as the second coefficient column 92. In the processing element 70 with the PE number [1], the data in the 1st row is assigned as the first coefficient column 91, and the data in the 6th row is assigned as the second coefficient column 92. In the processing element 70 with the PE number [2], the data in the 2nd row is assigned as the first coefficient column 91, and the data in the 5th row is assigned as the second coefficient column 92. Then, in the processing element 70 with the PE number [3], the data in the 3rd row is assigned as the first coefficient column 91, and the data in the 4th row is assigned as the second coefficient column 92.

[0034] In order to execute the above-described consecutive processing of the first coefficient column 91 and the second coefficient column 92 in parallel, the SIMD type processor 60 has a flag generation instruction. In the following description, for convenience, in the SIMD type processor 60, the configuration (functional unit) related to the execution of the flag generation instruction is referred to as a "flag generation unit 81", and the process related to the flag generation instruction is referred to as a "flag generation process".

[0035] The flag generation unit 81 shown in FIG. 11 identifies a plurality of parallelized processing elements 70 by the PE numbers associated with each processing element 70. When four processing elements 70 are SIMD-parallelized, these processing elements 70 are assigned PE numbers from [0] to [3]. The flag generation unit 81 generates a flag 84 associated with each PE number by executing a flag generation process based on an instruction. The flag 84 takes a value of 0 or 1.

[0036] The flag generation unit 81 acquires the value of the execution counter 82: N, and the value of the loop size LS (also refer to FIGS. 9 and 10, etc.): S. The execution counter 82 counts the processes executed by the processing element 70 through the execution of the counter process based on the instruction. The flag generation unit 81 generates a flag 84 associated with each of the processing elements 70 based on the loop size LS and the execution counter 82.

[0037] Specifically, the flag generation unit 81 compares the value of the execution counter 82 with the value of the loop size LS of the target data 90 assigned to each processing element 70, and determines the value of the flag 84. Specifically, when the value of the loop size LS of the target data 90 assigned to the PE number [0]: S[0] is the same as the value of the execution counter: N, the value of the flag [0] associated with the PE number [0] is changed from "0" to "1". Similarly, when the value of the loop size LS of the target data 90 assigned to the PE number [1]: S[1] is the same as the value of the execution counter: N, the value of the flag [1] associated with the PE number [1] is changed from "0" to "1". For the flags [2] and [3] associated with the PE numbers [2] and [3], similar flag generation is performed.

[0038] The register 63 shown in FIG. 12 includes a general-purpose register and an accumulator register individually associated with each processing element 70 as described above. The general-purpose register and the accumulator register each have a pair of banks, namely an even bank 64 and an odd bank 65. In other words, the general-purpose register and the accumulator register are composed of two banks, an even bank and an odd bank.

[0039] The first coefficient column 91 and the second coefficient column 92 are stored in the even bank 64 and the odd bank 65 of the general-purpose register, respectively. As an example, when a plurality of first coefficient columns 91 are loaded into the even bank 64, a plurality of second coefficient columns 92 are loaded into the odd bank 65. Or, when a plurality of second coefficient columns 92 are loaded into the even bank 64, a plurality of first coefficient columns 91 are loaded into the odd bank 65.

[0040] The general-purpose register is provided with two selectors 66 and 67. The selector 66 is input with a destination number and the value of the flag 84. The selectors 66 and 67 provide a value selected from the even bank 64 and the odd bank 65 to the processing element 70 based on the value of the flag 84 for an instruction that refers to the flag 84. In contrast, for an instruction that does not refer to the flag 84, the selectors 66 and 67 provide a value selected based on the least significant bit (LSB) of the register number.

[0041] Here, an instruction that refers to the flag 84 is an instruction related to the generation of the mel spectrogram (Figure 4 S30) and an instruction that processes the data of the loop size LS for each processing element 70. In contrast, an instruction that does not refer to the flag 84 is an instruction related to DFT or FTT (Figure 4 S10), the generation of the spectrogram (Figure 4 S20), and the inference of the DNN (Figure 4 S40), etc., and an instruction that processes the data of the same loop size LS in parallel.

[0042] In the instruction for generating the mel spectrogram, two values are loaded into the general-purpose register. Specifically, the value of each element of the mel spectrogram matrix S is loaded as the sound data to be processed into the general-purpose register (vload sound data). In addition, the value of each element of the transformation matrix T is loaded into the general-purpose register. A plurality of first coefficient columns 91 and a plurality of second coefficient columns 92 that are sequentially processed by each processing element 70 are loaded into the even bank 64 and the odd bank 65 of the general-purpose register, respectively (vload weight0, vload weight1).

[0043] The coefficient sequence loaded into the general-purpose register is a zero-extended coefficient sequence. By zero-extension, both the start timing and the end timing of parallel processing are aligned. Specifically, a value of "0" is set between each first coefficient sequence 91 and each second coefficient sequence 92 (see FIGS. 9 and 10). In addition, a value of "0" is also set on the front side of the first coefficient sequence 91 associated with the PE numbers [1] to [3] and on the rear side of the second coefficient sequence 92 associated with the PE numbers [0] to [2]. As described above, the start timing and the end timing of the calculation become uniform in the plurality of processing elements 70.

[0044] The processing element 70 executes a multiply-accumulate instruction (vmacf opr1, opr2, acc) using two values read from the general-purpose register based on the flag 84 and the value read from the accumulator register. Specifically, the processing element 70 multiplies the mel spectrogram matrix S (sound data) read from the general-purpose register by the coefficient sequence of the transformation matrix T, and uses the multiplication value obtained by the multiplication as one operand. Further, the processing element 70 uses the read value read from the accumulator register as the other operand, and calculates an addition value obtained by adding the multiplication value to the read value. The processing element 70 uses the accumulator register as the destination and stores the addition value in the accumulator register.

[0045] After executing the flag generation instruction, the processing element 70 repeatedly executes the multiply-accumulate instruction. That is, the processing element 70 repeatedly performs a multiply-accumulate process of adding the value obtained by multiplying the processing element 70, the sound data, and the coefficient sequence to the value read from the accumulator register, and storing the added value in the accumulator register. Further, the value provided from the general-purpose register to the processing element 70 is switched from the first coefficient sequence 91 to the second coefficient sequence 92 at different timings for each processing element 70 based on the value of the flag 84 generated by the flag generation instruction. As a result, each processing element 70 continuously processes the first coefficient sequence 91 to the second coefficient sequence 92 in one set of parallel processing.

[0046] (Summary of Embodiment) In the present embodiment described so far, based on the loop size LS of the target data 90 and the execution counter 82 or the count in the counter process, a flag 84 associated with each of the processing elements 70 is generated. Therefore, even if the loop sizes LS of the target data 90 processed by each processing element 70 are different from each other, the processing element 70 that has finished processing the target data 90 based on the flag 84 can be made to start processing the next target data 90. As a result, it is possible to suppress the processing element 70 from entering an idle state.

[0047] In addition, the general-purpose register and the accumulator register of the present embodiment each have an even bank 64 and an odd bank 65 as a pair of banks. Then, for an instruction that refers to the flag 84, the general-purpose register and the accumulator register provide a value selected from the even bank 64 and the odd bank 65 based on the flag 84. Further, for an instruction that does not refer to the flag 84, the general-purpose register and the accumulator register provide a value selected based on the least significant bit of the register number. According to such a bank configuration, while adopting a register 63 with a simplified configuration of the selectors 66, 67, it is possible to appropriately process both an instruction that refers to the flag 84 and an instruction that does not refer to the flag 84 in the processing element 70.

[0048] Also, the processing element 70 of the present embodiment uses, as one operand, a multiplication value of a value read from a general-purpose register based on the flag 84. Further, the processing element 70 uses, as the other operand, a read value read from the accumulator register, and calculates an addition value obtained by adding the multiplication value to the read value. Then, the processing element 70 executes a multiply-add instruction with the accumulator register as the destination. According to such arithmetic processing, it is possible to surely perform continuous processing of the target data 90 based on the flag 84.

[0049] Furthermore, in the present embodiment, two coefficient columns, i.e., a first coefficient column 91 and a second coefficient column 92, are loaded into the general-purpose register. Then, after an instruction for generating the flag 84 is executed, the processing element 70 executes the multiply-add instruction in a loop. According to the above, each processing element 70 can individually switch the target data 90 from the first coefficient column 91 to the second coefficient column 92 by referring to the generated flag 84. As a result, the idle state of the processing element 70 can be suppressed.

[0050] As described above, in the SIMD type processor 60, when tasks with different processing loads are executed by each processing element 70, the overall processing time can be shortened by continuously executing the next process in the processing element 70 where the process finishes quickly. On the other hand, since the processing start timing of each row assigned to each processing element 70 is the timing when the data (the above-mentioned sound data) required for the process is loaded, there is a deviation in the timing of starting the process for each row. Therefore, for a process that loops the same process, it is necessary to make changes that cause different behaviors in each processing element 70, which may lead to an increase in the number of code lines and an increase in processing time.

[0051] Therefore, in the present embodiment, an extended coefficient column is loaded into the general-purpose register. As a result, a "0" value is inserted into the weights for the number of cycles until each row can be processed, so that it is possible to tolerate the deviation in the processing start timing of each row without changing the processing within the loop. In addition, by inserting a "0" value, a situation where an incorrect value is output can also be avoided.

[0052] Also, in the general-purpose register of this embodiment, a plurality of first coefficient columns 91 and a plurality of second coefficient columns 92 that are sequentially processed by each processing element 70 are loaded. And the plurality of second coefficient columns 92 are assigned to each processing element 70 in a reverse order to the order in which the plurality of first coefficient columns 91 are assigned to each processing element 70 (see FIG. 10). As described above, in the conversion matrix T (see FIG. 7), as the number of rows progresses, the loop size LS tends to become longer as a tendency. As described above, by assigning the second coefficient column 92 in reverse order to the first coefficient column 91, it is possible to further shorten the processing time for one set.

[0053] In addition, in the above embodiment, the microcomputer board 30 and the SoC board 50 correspond to the "data processing device", the SIMD type processor 60 corresponds to the "processor", and the even bank 64 and the odd bank 65 correspond to the "pair of banks".

[0054] (Other embodiments) As described above, one embodiment according to the present disclosure has been described. However, the present disclosure is not construed as being limited to the above embodiment, and can be applied to various embodiments and combinations without departing from the gist of the present disclosure.

[0055] In the above embodiment, a configuration in which four processing elements 70 are parallelized has been illustrated. However, the number of processing elements 70 to be parallelized may be appropriately changed. For example, a very large number (e.g., 64) of processing elements 70 may be parallelized. As the number of parallelized processing elements 70 increases, the effect of suppressing the idle state by using the flag 84 and thus shortening the processing time can be further improved.

[0056] The configuration of the SIMD type processor 60 is not limited to the configuration exemplified in the above embodiment (see FIG. 4). For example, various CPUs with SIMD extensions (such as ARM's NEON, etc.) and microcontrollers can be used as the SIMD type processor 60. Also, the CPU may have any configuration of the ARM series, x86 series, and MIPS series.

[0057] The configuration of the register 63 of the SIMD type processor 60 may also be changed as appropriate. Specifically, the general-purpose register may have a two-bank configuration, and the accumulator register may have a one-bank configuration. Also, the accumulator register may have a two-bank configuration, and the general-purpose register may have a one-bank configuration. Furthermore, both the general-purpose register and the accumulator register may have a one-bank configuration.

[0058] The arithmetic processing using the flag 84 in the above embodiment is applicable not only to the generation of the Mel spectrogram. For example, in any of DFT or FFT, spectrogram generation, and inference using DNN, which are each steps of the speech recognition process, when tasks with different processing loads are executed by each processing element 70, the arithmetic processing using the flag 84 may be applied. Furthermore, in data processing for purposes other than speech recognition, such as image recognition and language processing, when tasks with different processing loads are executed by each processing element 70, the data processing method according to the present disclosure may be applied. For example, image data and text data, etc. can be the target data 90.

[0059] The microcontroller board 30 and the SoC board 50 may be mounted on various vehicles such as motorcycles, driverless vehicles for mobility services, construction machinery, agricultural machinery, railway vehicles, trams, and DMVs (Dual Mode Vehicles), and may be used for recognizing sound data. Also, the microcontroller board 30 and the SoC board 50 may be mounted on moving bodies other than vehicles such as ships, and electric aircraft such as drones and eVTOLs, and may be used for recognizing sound data. In addition, the microcontroller board 30 and the SoC board 50 may be used for recognizing not only the sound outside the moving body but also the sound inside the moving body (indoor sound). Furthermore, the microcontroller board 30 and the SoC board 50 may be installed together with a surveillance camera or the like at various facilities or along roadsides, etc., and may be used for recognizing sound data and the like.

[0060] The form of the storage medium (persistent tangible computer-readable medium, non-transitory tangible storage medium) that stores the data processing program and the like in the above embodiment may be appropriately changed. Furthermore, the storage medium may be an optical disk, a hard disk drive, a solid state drive, etc. that serve as the copy source or distribution source of the program.

[0061] The control unit and its method described in the present disclosure may be realized by a dedicated computer configured with a processor programmed to execute one or more functions embodied by a computer program. Alternatively, the device and its method described in the present disclosure may be realized by dedicated hardware logic circuits. Or, the device and its method described in the present disclosure may be realized by one or more dedicated computers configured by a combination of a processor that executes a computer program and one or more hardware logic circuits. Also, the computer program may be stored in a computer-readable non-transitory tangible recording medium as instructions to be executed by a computer.

Explanation of Reference Numerals

[0062] 30 Microcontroller board (data processing device), 50 SoC board (data processing device), 60 SIMD type processor (processor), 64 Even bank (pair of banks), 65 Odd bank (pair of banks), 70 Processing element, 81 Flag generation unit, 82 Execution counter, 84 Flag, 90 Target data, 91 First coefficient column, 92 Second coefficient column, LS Loop size

Claims

1. A plurality of processing elements (70); An execution counter (82) for counting the processes executed by the processing elements; A flag generation unit (81) that generates a flag (84) associated with each of the processing elements based on the loop size (LS) of target data (90) to be processed by the processing elements and the execution counter when the loop size of the target data to be processed by the processing elements is different for each processing element; A data processing apparatus comprising the above.

2. Further comprising a general-purpose register and an accumulator register each having a pair of banks (64, 65), The general-purpose register and the accumulator register, For an instruction that refers to the flag, provide a value selected from the banks based on the flag, The data processing apparatus according to claim 1, wherein for an instruction that does not refer to the flag, a value selected based on the least significant bit of the register number is provided.

3. The processing element, Uses the multiplication value of two values read from the general-purpose register based on the flag as one operand, Uses the read value read from the accumulator register as the other operand, calculates an addition value obtained by adding the multiplication value to the read value, The data processing apparatus according to claim 2, which executes a multiply-add instruction with the accumulator register as the destination.

4. Two coefficient columns are loaded into the general-purpose register, The data processing apparatus according to claim 3, wherein the processing element loop-executes the multiply-add instruction after the instruction for generating the flag is executed.

5. The data processing apparatus according to any one of claims 2 to 4, wherein a coefficient column extended with 0 is loaded into the general-purpose register.

6. A plurality of first coefficient columns (91) and a plurality of second coefficient columns (92) to be sequentially processed by each of the processing elements are loaded into the general-purpose register, The data processing apparatus according to any one of claims 2 to 4, wherein the plurality of second coefficient columns are assigned to each of the processing elements in a reverse order to the order in which the plurality of first coefficient columns are assigned to each of the processing elements.

7. A data processing program including instructions to be executed by a processor (60) comprising a plurality of processing elements (70), wherein The instructions are A counter process that counts the processes executed by the processing element, When the loop size (LS) of the target data (90) to be processed by the processing element is different for each processing element, a flag generation process that generates a flag (84) associated with each of the processing elements based on the count by the counter process and the loop size, A data processing program including the above.

Citation Information

Patent Citations

  • SIMD microprocessor

    JP2008071037A