Gradient Direction Quantization Device, Image Recognition Device, Gradient Direction Quantization Program, and Image Recognition Program

The gradient direction quantization device addresses the challenge of high-precision gradient direction calculations on limited resources by quantizing the gradient direction into 36 directions with reduced branch conditions, achieving efficient and precise image recognition.

JP7698835B2Active Publication Date: 2025-06-26AISIN CORP +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2020217071
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-12-25
Publication Date
2025-06-26
Estimated Expiration
2040-12-25

AI Technical Summary

Technical Problem

Existing image recognition algorithms face challenges in performing high-precision gradient direction calculations while minimizing computational resources, especially when dealing with limited calculation resources such as FPGA devices.

Method used

A gradient direction quantization device is developed, which includes a luminance acquisition unit, a tangent component acquisition unit, a quantization unit, and an output unit. The quantization unit uses a quadrant selection unit, a rotation selection unit, an offset selection unit, a tanθ table, and a multiplication unit to quantify the gradient direction into 36 directions with reduced branch conditions.

Benefits of technology

The solution enables high-precision gradient direction calculations with a small number of branch conditions, effectively suppressing calculation resources and improving efficiency on devices with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698835000001
    Figure 0007698835000001
  • Figure 0007698835000002
    Figure 0007698835000002
  • Figure 0007698835000003
    Figure 0007698835000003
Patent Text Reader

Abstract

To calculate a gradient direction with high accuracy under less branch conditions while suppressing a calculation resource.SOLUTION: When a luminance gradient in an x direction of a target pixel is fx, and a luminance gradient in a y direction is fy, a range of fy / fx can be caused to correspond to a range of θ. When θ is sectioned into 9 directions at each 10°, the range of fy / fx can be sectioned according to that and fy / fx can be quantized in 9 directions. If quantization in 9 directions is performed for each quadrant out of four quadrants, the gradient direction can be quantized in 36 directions (quadrant decision 4 conditions×angle decision 9 conditions=36 directions). When fxtanθ=fy and a range is variably set with fxtanθ as a boundary value, fy can be caused to correspond to quantized θ and division by fy / fx can be avoided. When a boundary value in the section where tanθ=fy / fx is approximated by a fixed point, accuracy can be adjusted with a bit width of a decimal part.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a gradient direction quantization device, an image recognition device, a gradient direction quantization program, and an image recognition program, and relates to quantizing the gradient direction of the luminance of pixels constituting an image.

Background Art

[0002] Among algorithms for image recognition of objects such as pedestrians, there are those that extract feature amounts of an object by focusing on the gradient direction of the luminance of pixels constituting an image, such as HOG, CoHOG, and MRCoHOG. For the gradient direction θ, assuming that the gradients of luminance in the x and y directions are fx and fy, it can be obtained by θ = arctan(fy / fx) from the relationship tanθ = fy / fx.

[0003] For example, in the technique of Non-Patent Document 1, the gradient direction is calculated by the arctangent of θ and quantized into 8 directions. By the way, since the calculation of the arctangent uses a lot of calculation resources, when implementing on a device with limited calculation resources such as an FPGA, it is desirable to avoid the calculation of trigonometric functions by imposing branch conditions on fx and fy to quantize the gradient direction. However, when the quantization direction is extended to many directions such as 36 directions, many branch conditions occur depending on the accuracy of the gradient direction to be obtained, leading to a problem of increased resources.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] The object of the present invention is to perform high-precision calculation of the gradient direction with fewer branch conditions while suppressing computational resources.

Means for Solving the Problems

[0006] (1) In the invention according to claim 1, there are provided a luminance acquisition means for acquiring the luminance of each pixel arranged on an image, a tangent component acquisition means for acquiring an fx which is a denominator component of the tangent of the luminance gradient and an fy which is a numerator component for each pixel based on the acquired luminance array, a quantization means for quantizing the gradient direction of the luminance by applying the acquired fx and fy to a range of tangent values corresponding to a predetermined number of quantized gradient directions, and an output means for outputting the quantized gradient direction. The quantization means receives the inputs of the fx and fy, and includes a quadrant selection unit that determines and outputs the quadrant of the gradient direction corresponding to the signs of the fx and fy among the first quadrant to the fourth quadrant; a rotation selection unit that receives the fx, fy, -fx and -fy whose signs are inverted, and the quadrant of the gradient direction output by the quadrant selection unit, and when the received quadrant is the first quadrant, selects fx or -fx that is 0 or more and fy or -fy that is 0 or more as the rotated fx and fy and outputs them, and when the received quadrant is in the second quadrant, third quadrant, or fourth quadrant, selects the fx and fy after being rotated 90°, 180°, or 270° clockwise respectively from fx, -fx, fy, -fy as the rotated fx and fy and outputs them; an offset selection unit that receives the input of offset values 0, 9, 18, and 27 corresponding to the first quadrant to the fourth quadrant respectively, and selects and outputs one of the offset values 0, 9, 18, and 27 corresponding to the quadrant determined by the quadrant selection unit; a tanθ table that stores approximate values of tanθ every 10° from 0° to 90°; a multiplication unit that multiplies each approximate value of tanθ every 10° by the rotated fx and calculates and outputs fxtanθ every 10°; a direction selection unit that outputs a numerical value (n - 1) of the direction corresponding to the nth range where the rotated fy is located among the first to ninth ranges with the fxtanθ every 10° as the boundary value; and an addition unit that adds the offset value output by the offset selection unit to the numerical value output by the direction selection unit. The approximate value of the tanθ is an approximation with a fixed decimal point where the bit width of the fractional part is 4 bits, 5 bits, or 6 bits, and the quantization means is configured by programming an FPGA (field-programmable gate array) composed of a number of LUTs (look-up tables), FFs (flip-flops), etc. The output means outputs, as the gradient direction obtained by quantizing the numerical value after addition in the addition unit, in 36 directions, and provides a gradient direction quantization device characterized by this. ( 2 ) In the invention according to claim 2 , there are provided an image acquisition means for acquiring an image, a gradient direction acquisition means for acquiring the gradient direction of the luminance of each pixel constituting the acquired image by the gradient direction quantization device according to claim to 1 , a feature amount acquisition means for acquiring a feature amount of the acquired image based on the distribution of the acquired gradient direction, and an image recognition means for performing image recognition using the acquired feature amount, and provides an image recognition device characterized by including these. ( 3 ) In the invention according to claim 3 , the feature amount acquisition means acquires the feature amount by the burden rate representing the co-occurrence distribution of the gradient direction by a mixture Gaussian distribution, and provides the image recognition device according to claim 2 . ( 4 ) In claim 4In the invention described in [reference], there is a gradient direction quantization program that realizes, by a computer, a luminance acquisition function for acquiring the luminance of each pixel arranged on an image, a tangent component acquisition function for acquiring, for each pixel, an fx which is the denominator component of the tangent of the luminance gradient and an fy which is the numerator component based on the acquired luminance array, a quantization function for quantizing the gradient direction of the luminance by applying the acquired fx and fy to a range of tangent values corresponding to a predetermined number of quantized gradient directions, and an output function for outputting the quantized gradient direction. The quantization function receives the input of fx and fy, determines and outputs the quadrant of the gradient direction corresponding to the signs of fx and fy among the first quadrant to the fourth quadrant (quadrant selection function), receives fx, fy, -fx and -fy whose signs are inverted, and the quadrant of the gradient direction output by the quadrant selection function. When the received quadrant is the first quadrant, it selects either fx or -fx that is 0 or more and either fy or -fy that is 0 or more and outputs them as the rotated fx and fy. When the received quadrant is in the second quadrant, the third quadrant, or the fourth quadrant, it selects, from fx, -fx, fy, -fy, the fx and fy after being rotated 90°, 180°, and 270° clockwise respectively and outputs them as the rotated fx and fy (rotation selection function), receives the input of offset values 0, 9, 18, 27 corresponding to the first quadrant to the fourth quadrant respectively, selects and outputs one of the offset values 0, 9, 18, 27 corresponding to the quadrant determined by the quadrant selection function (offset selection function), multiplies each approximation value of tanθ for every 10° in a tanθ table storing the approximation values of tanθ from 0° to 90° for every 10° by the rotated fx to calculate and output fxtanθ for every 10° (multiplication function), outputs a numerical value (n - 1) of the direction corresponding to the nth range where the rotated fy is located among the first to ninth ranges with the fxtanθ for every 10° as the boundary value (direction selection function), and includes an addition function for adding the offset value output by the offset selection function to the numerical value output by the direction selection function. The approximate value of the tanθ is an approximation with a fixed decimal point where the bit width of the fractional part is 4 bits, 5 bits, or 6 bits, and the quantization function is configured by programming an FPGA (field-programmable gate array) composed of a number of LUTs (look-up tables), FFs (flip-flops), etc. The output function outputs the numerical value added by the addition function as the gradient direction quantized into 36 directions. A gradient direction quantization program characterized by this is provided. ( 5 )Claim 5In the invention described in to 1 a gradient direction acquisition function for acquiring the gradient direction of the luminance of each pixel constituting the acquired image by the gradient direction quantization device described in claim

Effect of the Invention

[0007] According to the present invention, by using the value of the tangent corresponding to the quantized gradient direction, it is possible to perform high-precision calculation of the gradient direction with a small number of branch conditions while suppressing calculation resources.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Mode for Carrying Out the Invention

[0009] (1) Outline of the Embodiment When the luminance gradient in the x direction of the target pixel is fx and the luminance gradient in the y direction is fy, the gradient direction θ is the arctangent of fy / fx, that is, θ = arctan(fy / fx). The calculation process of this θ is difficult. In this embodiment, since fy / fx is the tangent of θ, fy / fx = tanθ is used. As shown in FIG. 5, tanθ is a monotonically increasing function, and the increase and decrease relationship between θ and tanθ is the same. For this reason, the range of fy / fx and the range of θ can be made to correspond. Therefore, when θ is divided into 9 directions every 10°, the range of fy / fx can be divided accordingly, and fy / fx can be quantized into 9 directions. When this quantization in 9 directions is performed for each of the 4 quadrants, the gradient direction can be quantized into 36 directions: quadrant determination 4 conditions × angle determination 9 conditions = 36 directions.

[0010] Also, by setting \(f_x\tan\theta = f_y\) and variably setting the range with \(f_x\tan\theta\) as the boundary value as shown in FIG. 6, \(f_y\) can be made to correspond to the quantized \(\theta\), and division by \(f_y / f_x\) can be avoided and calculation can be performed by multiplication. Furthermore, if the boundary value of the division of \(\tan\theta = f_y / f_x\) is approximated in fixed-point as shown in FIG. 4(c), its accuracy can be adjusted by the bit width of the fractional part. In this way, by utilizing the correlation between \(f_y / f_x\) and \(\theta\), high-precision calculation of the gradient direction can be performed with few branch conditions while suppressing calculation resources.

[0011] (2) Details of the Embodiment First, the HOG feature amount, CoHOG feature amount, and MRCoHOG feature amount will be described. FIG. 1 is a diagram for explaining the concept of the HOG feature amount. The HOG feature amount is extracted from an image by the following procedure. The image 101 shown in the left diagram of FIG. 1(a) is set as a target image region such as an observation window for observing the target. First, the image 101 is divided into rectangular cells 102a, 102b, ···. Next, as shown in the right diagram of FIG. 1(a), for each cell 102, the gradient direction (the direction from low luminance to high luminance) of the luminance of each pixel is quantized into a plurality of directions for each predetermined angle. Hereinafter, the gradient direction of the luminance will simply be referred to as the gradient direction. Also, the same quantized gradient direction is used for the CoHOG feature amount and the MRCoHOG feature amount.

[0012] Next, as shown in FIG. 1(b), by generating a histogram with the quantized gradient direction as the class and the number of occurrences as the frequency, a histogram 106 of the gradient direction included in the cell 102 is created for each cell 102. Then, normalization is performed so that the total frequency of the histogram 106 becomes 1 in units of blocks formed by collecting several cells 102.

[0013] In the example of Fig. 1(a) left figure, one block is formed from cells 102a, 102b, 102c, and 102d. The histogram 107 of the image 101 is the histogram in which the histograms 106a, 106b,... normalized in this way are arranged in a row as shown in Fig. 1(c).

[0014] Fig. 2 is a diagram for explaining the CoHOG feature amount. The CoHOG feature amount is a feature amount focusing on the gradient pair between two pixels in a local area, and is extracted from an image by the following procedure. As shown in Fig. 2(a), the image 101 is divided into rectangular cells 102a, 102b,.... Note that a cell is also called a block.

[0015] In the CoHOG feature amount, a target pixel 110 is set in cells 102a, 102b,..., and a co-occurrence matrix (histogram regarding the target pixel 110) is created by the combination of the gradient direction of the target pixel 110 and the gradient directions of the pixels at distances 1 to 4 from the target pixel 110. Note that the pixels related to the combination with the target pixel 110 are called offsets.

[0016] For example, the distance from the target pixel 110 is represented by a mathematical formula. When the formula is applied, as shown in Fig. 2(a), pixels 1a to 1d adjacent to the target pixel 110 are obtained as the pixels at distance 1. The reason why the pixels above and to the left of the target pixel 110 are not included in the combination is that the target pixels 110 are set and processed in order from the left end of the top pixel row in the right direction, and the processing has already been completed.

[0017] Next, observe the luminance gradient directions of the target pixel 110 and the pixel 1a. In the figure, the gradient direction is shown by an arrow line. The luminance gradient direction of the target pixel 110 is in the right direction, and the gradient direction of the pixel 1a is in the upper right direction. Therefore, in the co-occurrence matrix 113 of Fig. 2(b), one vote is cast for the element at (row number, column number) = (right direction, upper right direction). In the example of Fig. 2(b), as a combination of the gradient directions of the target pixel 110 and the pixel 1a, 1 is added to the element in the row with the rightward arrow described as the row number and the column with the upper-rightward arrow described as the column number, and as a result, the value of the element has become 10.

[0018] Note that originally, the co-occurrence matrix 113 should be drawn as a three-dimensional histogram and the number of votes should be represented by a bar graph in the height direction, but for simplicity of the figure, the number of votes is represented numerically. Hereinafter, similarly, voting (counting) is performed by combining the target pixel 110 with the pixels 1b, 1c, and 1d.

[0019] As shown in Fig. 2(c), centering on the target pixel 110, the pixels at a distance of 2 are the pixels 2a to 2f on the outer periphery of the pixels 1a to 1d, the pixels at a distance of 3 are the pixels 3a to 3h on the further outer periphery thereof, and the pixels at a distance of 4 are the pixels 4a to 4l on the further outer periphery thereof. Similarly for these, voting is performed on the co-occurrence matrix 113 in combination with the target pixel 110.

[0020] The above voting process is performed for all the pixels constituting the cell 102, and a co-occurrence matrix for each pixel is obtained. Furthermore, this is performed for all the cells 102, and a histogram in which the components of all the co-occurrence matrices are arranged in a column as shown in Fig. 2(d) is the CoHOG feature 117 of the image 101.

[0021] Fig. 3 is a diagram for explaining the MRCoHOG feature. The MRCoHOG feature significantly reduces the number of offsets by looking at co-occurrence between different resolutions of the same image. First, as shown in Fig. 3(a), by generating images with different resolutions (image sizes) from the original image, a high-resolution image 120 (original image), a medium-resolution image 121, and a low-resolution image 122 are obtained. The squares in the image represent pixels. Although not shown, cells (also called blocks) are also set in each of these resolution images. Then, the gradient direction is calculated for each pixel of the high-resolution image 120, the medium-resolution image 121, and the low-resolution image 122.

[0022] For the extraction of MRCoHOG features, the medium-resolution image 121 and the low-resolution image 122 are used. For the sake of clarity, as shown in Fig. 3(b), the medium-resolution image 121 and the low-resolution image 122 are stretched into the medium-resolution image 121a and the low-resolution image 122a to have the same size as the high-resolution image 120.

[0023] Next, as shown in Fig. 3(c), similar to the CoHOG feature, the co-occurrence (combination of luminance gradient directions) between the gradient direction at the target pixel 125 in the high-resolution image 120 and the gradient directions of the pixels 1a to 1d in the surrounding high-resolution image 120 is obtained, and votes are cast on a co-occurrence matrix (not shown).

[0024] Next, votes are cast on the co-occurrence matrix according to the co-occurrence between the target pixel 125 in the high-resolution image 120 and the pixels 2a to 2d in the medium-resolution image 121a on the outer periphery of the pixels 1a to 1d. Furthermore, votes are cast on the co-occurrence matrix according to the co-occurrence between the target pixel 125 and the pixels 3a to 3d in the low-resolution image 122a on the outer periphery of the pixels 2a to 2d.

[0025] In this way, for the target pixel 125 in the high-resolution image 120, a co-occurrence matrix is obtained by taking co-occurrences in combinations within the high-resolution image 120, in combination with the medium-resolution image 121a, and in combination with the low-resolution image 122a. This process is performed for each pixel within the cell of the high-resolution image 120, and further for all cells. Thereby, a co-occurrence matrix for each pixel of the high-resolution image 120 is obtained.

[0026] Similarly, further, the co-occurrence matrix with each resolution image when a target pixel is set in the medium-resolution image 121a and the co-occurrence matrix with each resolution image when a target pixel is set in the low-resolution image 122a are calculated, and a histogram in which the components of all co-occurrence matrices are arranged in a column as shown in Fig. 3(d) is the MRCoHOG feature 127 of the high-resolution image 120.

[0027] Note that in this example, the histogram obtained by concatenating the co-occurrence matrices when the target pixel is set in the high-resolution image 120, the co-occurrence matrix when the target pixel is set in the medium-resolution image 121a, and the co-occurrence matrix when the target pixel is set in the low-resolution image 122a is used as the MRCoHOG feature amount. However, it is also possible to use, for example, the histogram based on the co-occurrence matrix when the target pixel is set in the high-resolution image 120 as the MRCoHOG feature amount. Also, any two of them may be combined, or further, the resolution may be increased to perform co-occurrence with four or more types of resolution images. The MRCoHOG feature amount can significantly reduce the feature amount compared to CoHOG, while having the feature that its robustness is better than that of CoHOG.

[0028] In the above algorithm, the feature amount is extracted based on the gradient direction of each pixel. These gradient directions can be obtained by the gradient direction quantization device 7 described with reference to FIG. 7. Here, the calculation method of the angle of the gradient direction performed by the gradient direction quantization device 7 will be described.

[0029] FIG. 4 is a diagram for explaining a method of approximating tanθ in fixed-point decimal. As described later, in the present embodiment, the tangent value is used to quantize the gradient direction, and this is approximated by a fixed-point decimal suitable for implementation in hardware.

[0030] The fx and fy shown in FIG. 4(a) represent the gradient intensities of the luminance in the x direction (horizontal direction) and the y direction (vertical direction), respectively. Mathematically, fx and fy are obtained by partially differentiating the luminance in the x direction and the y direction. However, in the present embodiment, as shown in Equation (41), fx is represented by the difference in luminance of the pixels adjacent to both sides in the horizontal direction (left and right horizontal direction) of the target pixel, and fy is represented by the difference in luminance of the pixels adjacent to both sides in the vertical direction (up and down vertical direction) of the target pixel. In Equation (41), Y represents the luminance, and x and y represent the coordinate values of the target pixel.

[0031] As a result, the luminance gradient is represented by a vector from point T to point S, and the gradient direction θ is represented by the angle formed by the vector and fx. As shown in Equation (42) of FIG. 4(b), the gradient direction θ is represented by the arctangent of the value obtained by dividing fy by fx. Thus, as shown in Equation (43), tanθ = fy / fx. Furthermore, as shown in Equation (44), the relationship fx × tanθ = fy also holds. Equations (43) and (44) are respectively used to quantize the gradient direction in FIGS. 5 and 6.

[0032] In this embodiment, as shown in FIG. 4(c), the boundary values of the range of tanθ used for quantization are approximately represented using a fixed-point number with a 3-bit integer part bit width and a 0- to 7-bit decimal part bit width. By using a fixed-point number, calculation resources can be saved, and high-speed operations such as shift operations become possible. In this way, the gradient direction quantization device 7 approximates the boundary values of the tangent range corresponding to quantization with a fixed-point number whose decimal part bit width is any one of 0 bits to 7 bits.

[0033] The larger the bit width of the decimal part, the higher the approximation accuracy, but the more calculation resources are used accordingly. Therefore, an appropriate bit width is set considering the balance between accuracy and the calculation resources used. In this embodiment, it is set to 4 bits based on the experiments described later. Hereinafter, the case where the bit width of the decimal part is 4 bits will be described, and the value with the 4-bit width will be simply referred to as an approximate value. In this way, as an example, the gradient direction quantization device 7 has a 4-bit decimal part bit width for the fixed-point number.

[0034] FIG. 4(d) shows the approximate values of the tanθ values every 10°. As shown in the figure, the values of tanθ are 0.176, 0.364, 0.577, 0.839, 1.19, 1.73, 2.75, 5.67 for every 10° from 10° to 80°, and the corresponding approximate values are 0.1875, 0.375, 0.5625, 0.8425, 1.1875, 1.75, 2.75, 5.6875.

[0035] Figure 5 is a diagram for explaining a quantization method of θ using the approximate value of tanθ. The graph shown in the figure represents the relationship between tanθ = fy / fx and θ in the range of 0° ≤ θ < 90°. As shown in the figure, since tanθ is a monotonically increasing function, tanθ = fy / fx in the formula (43) shown in Fig. 4(b) can be made to correspond to θ without duplication. Therefore, by dividing tanθ = fy / fx into a plurality of ranges and making these correspond to the ranges of θ, these ranges can be used as units of quantization.

[0036] In this embodiment, the tanθ axis is divided in ranges with approximate values for every 10° of θ as boundary values, and these are made to correspond to θ for every 10°. Specifically, the tanθ axis is divided into ranges of 0 ≤ tanθ < 0.1875, 0.1875 ≤ tanθ < 0.375, ···, 5.6875 ≤ θ, and these are made to correspond to the first θ, the second θ, ···, the ninth θ respectively.

[0037] The nth θ corresponds to the nth direction (n is an integer from 1 to 9), and thereby, the angles from 0° to 90° can be quantized in 9 directions. For example, as shown in the figure, when tanθ = fy / fx is in the range of 1.75 ≤ tanθ < 2.75, it is quantized to the seventh θ, that is, the seventh direction. The above is the case when θ is in the first quadrant. However, by the method described later, for the other three quadrants as well, it is quantized to the tenth θ, the eleventh θ, ···, the thirty-sixth θ, and the gradient direction is quantized in a total of 36 directions.

[0038] Thus, the gradient direction quantization device 7 includes quantization means for quantizing the gradient direction θ of luminance by applying the denominator component fx and the numerator component fy of the tangent of the gradient direction of luminance to the range of tangent values corresponding to a predetermined number of quantized gradient directions. The quantization means can quantize the gradient direction of luminance by applying fy / fx, which is the value of the tangent composed of the denominator component and the numerator component, to the range.

[0039] FIG. 6 is a diagram for explaining a method of avoiding division by fy / fx. The gradient direction may be quantized according to the example of FIG. 5, but it is necessary to calculate fy / fx. In the example of FIG. 6, the division of fy by fx is avoided using the relationship fx×tanθ = fy in equation (d), and θ is quantized by multiplication. By avoiding division, the use of computing resources can be saved.

[0040] In this example, the value obtained by multiplying the approximate value of tanθ by fx is used as the boundary value, and the range thereby obtained is made to correspond to the quantized θ. Thereby, fy can be made to correspond to the quantized θ. For example, as shown in the figure, when 1.75fx ≤ fy < 2.75fx, this is quantized to the 7th θ. By variably setting the range according to fx on the tanθ axis, fy can be made to correspond to the quantized θ, and the gradient direction can be quantized by multiplication.

[0041] The quantization means in this example sets the boundary value of the quantization range according to fx, which is the denominator component of the tangent of the gradient direction, and applies fy, which is the numerator component, to the range of the set boundary value to quantize the gradient direction of luminance.

[0042] FIG. 7 is a diagram for explaining the circuit configuration of the gradient direction quantization device 7. The gradient direction quantization device 7 can be configured, for example, by programming a field-programmable gate array (FPGA). An FPGA is an integrated circuit that allows users to program the circuit configuration and is composed of a large number of LUTs (look-up tables), FFs (flip-flops), etc.

[0043] First, the gradient direction quantization device 7 receives the inputs of fx and fy. Although the method shown in FIG. 5 is also possible, here, the gradient direction quantization device 7 quantizes the gradient direction using the method shown in FIG. 6. This avoids calculations such as division and arctangent that require a large number of computing resources, and the gradient direction can be quantized into 36 directions at 10° intervals with few conditional branches.

[0044] The quadrant selection unit 74 receives the inputs of fx and fy, determines the quadrant of the gradient direction, and outputs the determination result to the rotation selection unit 71 and the offset selection unit 75. Since the relationship in FIG. 6 is quantization when the gradient direction is in the first quadrant (0 ≤ θ < 90°), when the gradient direction is in the second to fourth quadrants, the gradient direction quantization device 7 rotates this to map it to the first quadrant and quantizes it, and then offsets it in the reverse rotation direction to calculate the gradient directions in the second to fourth quadrants. The quadrant selection unit 74 performs quadrant determination for this purpose.

[0045] More specifically, the quadrant selection unit 74 receives inputs of numerical values 0 to 3 corresponding to the first to fourth quadrants respectively, and selects and outputs the numerical value of the corresponding quadrant according to the values of fx and fy. The correspondence between fx, fy and the quadrant of the gradient direction is based on the signs of fx and fy. For example, when fx > 0 and fy ≥ 0, it is the first quadrant; when fx ≤ 0 and fy > 0, it is the second quadrant; when fx < 0 and fy ≤ 0, it is the third quadrant; when fx ≥ 0 and fy < 0, it is the fourth quadrant.

[0046] When the gradient direction is in the second to fourth quadrants, the rotation selection unit 71 rotates it by an angle corresponding to the quadrant to map it to the first quadrant. The mapped gradient direction becomes the quantization target. More specifically, the rotation selection unit 71 receives the inputs of fx, fy, -fx, and -fy with inverted signs, and also receives the quadrant of the gradient direction as a numerical value from 0 to 3 from the quadrant selection unit 74. For example, when the quadrant received from the quadrant selection unit 74 is the first quadrant, the rotation selection unit 71 selects the value that is 0 or greater out of fx and -fx (since fx takes positive and negative values, the selected value is the absolute value of fx), and also selects the value that is 0 or greater out of fy and -fy, and outputs these as the rotated fx and fy. In this case, the rotation angle is 0°.

[0047] When the gradient direction is in the second quadrant, third quadrant, or fourth quadrant, the rotated fx and fy that rotate these by 90°, 180°, and 270° clockwise respectively are selected from fx, -fx, fy, and -fy and output. Specifically, when the gradient direction is in the second quadrant or fourth quadrant, the value that is 0 or greater out of fy and -fy is selected as the rotated fx, and the value that is 0 or greater out of fx and -fx is selected as fy. When the gradient direction is in the third quadrant, the value that is 0 or greater out of fx and -fx is selected as the rotated fx, and the value that is 0 or greater out of fy and -fy is selected as fy.

[0048] In this way, the gradient direction quantization device 7 includes a rotation means that rotates the gradient direction of the pixel luminance by a predetermined angle according to the quadrant of the gradient direction. Further, the gradient direction quantization device 7 includes a tangent component acquisition means that acquires the denominator component and the numerator component of the tangent of the luminance gradient for each pixel based on the acquired luminance array, and the rotation means rotates the gradient direction to a predetermined quadrant for quantization.

[0049] The offset selection unit 75 receives the input of the quadrant of the gradient direction from the quadrant selection unit 74, selects an offset value corresponding thereto, and outputs it to the addition unit 77. More specifically, the offset selection unit 75 receives the input of offset values 0, 9, 18, and 27 corresponding to the first quadrant to the fourth quadrant respectively, selects the offset value corresponding to the quadrant determined by the quadrant selection unit 74, and outputs it to the addition unit 77.

[0050] In order to quantize the gradient direction into nine directions in one quadrant, when the gradient direction is in the second quadrant, if 9 is added as an offset value to the numerical value representing the direction quantized by the rotated fx and fy, it will be the direction after rotating counterclockwise by 90° and returning to the second quadrant. Similarly, when the gradient direction is in the first quadrant, the third quadrant, or the fourth quadrant, if 0, 18, and 27 are added as offset values to the numerical value of the quantized direction respectively, it will be the quantized direction in the original quadrant. Thus, the gradient direction quantization device 7 includes correction means for correcting the quantized gradient direction to a direction corresponding to the quadrant of the gradient direction before rotation, and the correction means performs correction to rotate the quantized gradient direction in the reverse direction by a predetermined angle of rotation.

[0051] The tanθ table 72 stores the approximate values of tanθ every 10° shown in FIG. 5 and outputs them to the multiplication unit 73. The multiplication unit 73 multiplies the approximate value output by the tanθ table 72 by the rotated fx to calculate fxtanθ in FIG. 6 and outputs it to the direction selection unit 76.

[0052] The direction selection unit 76 quantizes the gradient direction mapped to the first quadrant by rotation, selects the numerical value corresponding to that direction, and outputs it to the addition unit 77. More specifically, the direction selection unit 76 receives the input of the numerical values 0 to 8 corresponding to the first θ to the ninth θ respectively, and also receives the input of fxtanθ and the rotated fy from the tanθ table 72 and the rotation selection unit 71 respectively. Then, the direction selection unit 76 selects and outputs the numerical value of the direction corresponding to the range where fy is located among the ranges with fxtanθ as the boundary value. For example, when fy is in the range of the first θ, the direction selection unit 76 outputs 0.

[0053] The addition unit 77 adds the offset value output by the offset selection unit 75 to the numerical value output by the direction selection unit 76, and outputs the numerical value corresponding to the gradient direction quantized into 36 directions. As described above, the gradient direction quantization device 7 outputs the gradient direction quantized in 36 directions as integer values from 0 to 35 corresponding to the 1st θ to the 36th θ. In this way, the gradient direction quantization device 7 includes output means for outputting the gradient direction corrected by rotation to return to the original quadrant.

[0054] FIG. 8 is a flowchart for explaining the procedure by which the gradient direction quantization device 7 quantizes the gradient direction. First, the gradient direction quantization device 7 receives the inputs of fx and fy (step 305). Next, the quadrant selection unit 74 selects and outputs a numerical value corresponding to the quadrant of the gradient direction from the input fx and fy (step 310). Next, the offset selection unit 75 selects and outputs an offset value according to the numerical value selected by the quadrant selection unit 74 (step 315).

[0055] On the other hand, the rotation selection unit 71 selects and outputs the rotated fx and fy from fx, fy, -fx, and -fy (step 320). Next, the multiplication unit 73 multiplies the rotated fx by the approximate value of the tan θ table to generate and output the range of fxtanθ (step 325).

[0056] Next, based on the range of fxtanθ output by the multiplication unit 73 and the rotated fy output by the rotation selection unit 71, the direction selection unit 76 selects and outputs a numerical value corresponding to the quantized direction (step 330). Then, the addition unit 77 adds the numerical value output by the direction selection unit 76 and the offset value output by the offset selection unit 75 to output a numerical value corresponding to the quantized gradient direction (step 335), and returns to the main routine (FIG. 19) described later.

[0057] FIG. 9 is a diagram for explaining the resource usage amount and the like when the gradient direction quantization device 7 is implemented on an FPGA. FIG. 9(a) shows the number of resources of LUT and FF when the bit width of the fractional part is changed from 0 to 7. For the LUT, when the bit width is 0, there are 97, and as the bit width increases up to 297 when the bit width is 6, it increases as the bit width increases. When the bit width is 7, it decreases to 284, which is considered to be because part of the calculation is taken over by the use of the DSP described below. For the FF as well, it increases from 52 when the bit width is 0 to 183 when the bit width is 6, and decreases to 175 when the bit width is 7.

[0058] Incidentally, when calculating the arctangent to a high order, about 10,000 LUTs and about 6000 FFs are required. For comparison, experiments were also conducted on the case of branching fx and fy by a large number of branch conditions and quantizing in 36 directions. This quantization was performed by simply classifying 36 directions according to 16 branch conditions based on the combination of the ranges of the magnitudes of fx and fy and the ranges of the sum and difference of fx and fy. In this example, for each condition, 3,087 LUTs and 76 FFs were used. In this way, in the branch using the approximate value of tanθ of the gradient direction quantization device 7, a large amount of calculation resources can be saved.

[0059] Figure 9(b) shows the number of DSP (Digital Signal Processor) resources used when the bit width of the fractional part is changed from 0 to 7. The DSP is a circuit that performs arithmetic operations on digital signals. As shown in the figure, when the bit width of the fraction is from 0 to 6, no DSP is used, and 1 is used when it is 7 bits.

[0060] Figure 9(c) is a graph evaluating the accuracy of the 36-direction angle calculation by the gradient direction quantization device 7. Here, the coincidence rate for all 261,121 angles was verified. The angle coincidence rate was 46.7%, 57.6%, 79.5%, 90.1%, 96%, 97.1%, 98.7%, and 98.8% in order from the case where the bit width of the fractional part is 0 to the case where it is 7. Also, in the experiment for the above comparison, it was 91%. Also, the maximum error in angle (the number of cases classified as the angle adjacent to the original quantized angle) was 2 when the bit width was 0, but 1 when it was from 1 to 7.

[0061] In this way, increasing the bit width of the fractional part improves the accuracy, but when the bit width is 4 or more, it is almost flat. Considering the resource usage, a bit width of 4 bits seems appropriate.

[0062] Next, an example of applying the gradient direction quantization device 7 to an image recognition device will be described. In this example, an MRCoHOG feature amount using the appearance frequency of co-occurrence of gradient directions across different resolutions of the same image as a feature amount is extracted by a Gaussian Mixture Model (hereinafter referred to as GMM). In the present embodiment, since the gradient direction is quantized in 36 directions, when voting for the co-occurrence matrix, the dimension becomes large and a large amount of computational resources are required. Therefore, the method using GMM described below is adopted so that it can be implemented on a small device such as an FPGA. First, a method for creating a GMM serving as a reference for image recognition from such gradient directions will be described.

[0063] FIG. 10 is a diagram for explaining a method for creating a reference GMM. As shown in FIG. 10(a), the image processing device 8 (FIG. 17) receives an input of the image 2 for creating a reference GMM and divides it into a plurality of block regions 3A, 3B,... of the same rectangular shape. The image 2 is, for example, an image of a pedestrian who is the object of image recognition. In the figure, it is divided into 4×4 for easy illustration, but a standard value is, for example, 4×8. Note that when the block regions 3A, 3B,... are not particularly distinguished, they are simply denoted as the block region 3.

[0064] The image processing apparatus 8 divides the image 2 into block regions 3, and converts the resolution of the image 2 to generate high-resolution images 11, medium-resolution images 12, and low-resolution images 13 with different resolutions (image sizes) as shown in FIG. 10(b). When the resolution of the image 2 is appropriate, the image 2 is used as the high-resolution image as it is. In the figure, the high-resolution image 11, medium-resolution image 12, and low-resolution image 13 of the portion of the block region 3A are shown, and the grids schematically represent pixels.

[0065] Then, the image processing apparatus 8 uses the gradient direction quantization device 7 to quantize and calculate the gradient direction of each pixel of the high-resolution image 11, medium-resolution image 12, and low-resolution image 13 in 36 directions every 10°. When the image processing apparatus 8 calculates the gradient direction in this way, it obtains the co-occurrence of the gradient direction between the target pixel and the pixels at positions away from it (hereinafter referred to as offset pixels) as follows.

[0066] First, as shown in FIG. 10(c), the image processing apparatus 8 sets the target pixel 5 in the high-resolution image 11, and focuses on the offset pixels 1a to 1d at an offset distance 1 (i.e., adjacent in the high resolution) from the target pixel 5 in the high-resolution image 11. Note that the distance for n pixels is referred to as the offset distance n.

[0067] Then, the image processing apparatus 8 obtains the co-occurrence (combination of gradient directions) of the gradient directions between the target pixel 5 and the offset pixels 1a to 3d, and plots the corresponding points as co-occurrence corresponding points 51, 51,... on the feature surfaces 15(1a) to 15(3d) shown in FIG. 10(d). In the example of FIG. 3, voting was performed on the co-occurrence matrix, but in the method using GMM, voting is performed by plotting on the feature surface 15. Note that the image processing apparatus 8 creates the 12 feature surfaces 15(1a) to 15(3d) shown in FIG. 10(d) for each of the block regions 3A, 3B,... divided in FIG. 10(a). Hereinafter, when referring to the entire plurality of feature surfaces, it is referred to as the feature surface 15.

[0068] For example, when plotting the co-occurrence of the target pixel 5 and the offset pixel 1a in FIG. 10(c), if the gradient direction of the target pixel 5 is the third θ (the numerical value output by the gradient direction quantization device 7 is 2), and the gradient direction of the offset pixel 1a is the 14th θ (the numerical value output by the gradient direction quantization device 7 is 13), then the image processing device 8 plots a co-occurrence corresponding point 51 at a position where the horizontal axis (x-axis) of the feature surface 15(1a) for the offset pixel 1a is the third θ and the vertical axis (y-axis) is the 14th θ.

[0069] Note that the x-axis and y-axis of the feature surface 15 are quantized into 36 intervals at every 10°, and the feature surface 15 is composed of 36×36 grids. Although the plotted points are represented as a dense group in the figure, the image processing device 8 plots and votes for each of these quantized grids.

[0070] Then, while sequentially moving the target pixel 5 within the high-resolution image 11, the image processing device 8 takes the co-occurrence of the target pixel 5 and the offset pixel 1a and plots it on the feature surface 15(1a). In this way, the feature surface 15 represents the occurrence frequency of pairs of two gradient directions having a specific offset (relative position from the target pixel 5) in the image.

[0071] Note that in FIG. 10(c), observing the co-occurrence of the pixel on the right side of the target pixel 5 facing the drawing surface is to first set a movement path where the target pixel 5 moves from the pixel at the upper left end to the pixels in the right direction sequentially, and when reaching the right end, it moves in the right direction from the pixel at the left end one row below, so as not to obtain overlapping co-occurrence combinations as the target pixel 5 moves.

[0072] Also, the movement of the target pixel 5 is performed within the block region 3A (within the same block region), but the selection of the offset pixel is also performed even when it exceeds the block region 3A. At the edge of the image 2, the gradient direction cannot be calculated, and this is processed by an appropriate arbitrary method.

[0073] Next, the image processing apparatus 8 obtains the co-occurrence in the gradient direction between the target pixel 5 and the offset pixel 1b (see FIG. 10(c)), and plots the corresponding co-occurrence corresponding point 51 on the feature surface 15(1b). Note that the image processing apparatus 8 prepares a new feature surface 15 different from the feature surface 15(1a) previously used for the target pixel 5 and the offset pixel 1a, and votes on this. In this way, the image processing apparatus 8 generates a feature surface 15 for each combination of the relative positional relationship between the target pixel 5 and the offset pixel. Then, while sequentially moving the target pixel 5 within the high-resolution image 11, the image processing apparatus 8 obtains the co-occurrence between the target pixel 5 and the offset pixel 1b, and plots the co-occurrence corresponding point 51 on the feature surface 15(1b).

[0074] Similarly, the image processing apparatus 8 also prepares individual feature surfaces 15(1c) and 15(1d) for the combination of the target pixel 5 and the offset pixel 1c and the combination of the target pixel 5 and the offset pixel 1d, respectively, and plots the co-occurrence in the gradient direction.

[0075] In this way, when the image processing apparatus 8 generates four feature surfaces 15 for the target pixel 5 and the offset pixels 1a to 1d at an offset distance of 1 from the target pixel 5, next, the image processing apparatus 8 focuses on the target pixel 5 in the high-resolution image 11 and the offset pixels 2a to 2d of the medium-resolution image 12 at an offset distance of 2.

[0076] Then, by the same method as the above method, a feature surface 15(2a) for the combination of the target pixel 5 and the offset pixel 2a, and similarly, feature surfaces 15(2b) to 15(2d) for the combinations of the offset pixels 2b, 2c, and 2d are created.

[0077] Similarly, for the target pixel 5 in the high-resolution image 11 and the offset pixels 3a to 3d of the low-resolution image 13 at an offset distance of 3, the image processing apparatus 8 also generates feature surfaces 15(3a) to 15(3d) for each combination of the relative positional relationship between the target pixel 5 and the offset pixels 3a to 3d. The image processing apparatus 8 also performs the above processing on the block areas 3B, 3C, ···, and generates a plurality of feature surfaces 15 that extract the features of the image 2. In this way, the image processing apparatus 8 generates a plurality of feature surfaces 15(1a) to 15(3d) for each of the block areas 3A, 3B, 3C, ···.

[0078] Then, for each of these feature surfaces 15, the image processing apparatus 8 generates a GMM as follows. Here, for the sake of simplicity, a GMM is generated from the feature surface 15 created from the image 2. More specifically, a GMM is generated for the superimposed feature surfaces 15 obtained from a large number of training images.

[0079] Fig. 10(e) shows one of these plurality of feature surfaces 15. First, the image processing apparatus 8 clusters the co-occurrence corresponding points 51 into K clusters (groups) by combining those that are close to each other. The number of mixtures represents the number of Gaussian distributions to be mixed when generating a GMM. When this is appropriately specified, the image processing apparatus 8 automatically clusters the co-occurrence corresponding points 51 into the specified number.

[0080] The number of mixtures can be set to various values such as K = 6, K = 16, K = 32, K = 64, etc., and the value of the number of mixtures is determined by experiments or the like. In the example of Fig. 10(e), for the sake of simplicity, K = 3, and the co-occurrence corresponding points 51 are clustered into clusters 60-1 to 60-3. The co-occurrence corresponding points 51, 51, ··· plotted on the feature surface 15 tend to gather according to the features of the image, and the clusters 60-1, 60-2, ··· reflect the features of the image.

[0081] As shown in FIG. 10(f), after clustering the co-occurrence corresponding points 51, the image processing apparatus 8 represents the probability density function 53 of the co-occurrence corresponding points 51 on the feature surface 15 by a probability density function p(x|θ) obtained by linearly superposing K Gaussian distributions (Gaussian distributions 54-1, 54-2, 54-3). In this way, the Gaussian distribution is used as a basis function (a function that is the target of linear summation and is an element constituting the GMM), and the probability density function 53 represented by the linear summation thereof is the GMM. The image processing apparatus 8 uses the probability density function 53 as a reference GMM55 for determining whether the learned object and the subject are the same or different.

[0082] The specific mathematical formula of the probability density function p(x|θ) is as shown in FIG. 10(g). Here, x is a vector quantity representing the distribution of the co-occurrence corresponding points 51, and θ is a vector quantity representing parameters (μj, Σj) (where j = 1, 2, ···, K). πj is called a mixing coefficient and represents the probability of selecting the j-th Gaussian distribution. μj and Σj represent the mean value and the covariance matrix of the j-th Gaussian distribution, respectively. The probability density function 53, that is, the reference GMM55, is uniquely determined by πj and θ.

[0083] z is a latent parameter used for calculating the EM algorithm and the burden rate, and z1, z2, ···, zK are used corresponding to the K Gaussian distributions to be mixed. The burden rate is obtained by calculating the probability of z from the distribution of x. Although the explanation of the EM algorithm is omitted, it is an algorithm for estimating πj and parameters (μj, Σj) that maximize the likelihood. The image processing apparatus 8 determines πj and θ by applying the EM algorithm, and thereby obtains p(x|θ).

[0084] The reference GMM55 is formed by mixing Gaussian distributions 54-1, 54-2, 54-3 (not shown) located at the positions of clusters 60-1, 60-2, 60-3 as basis functions. Then, using the reference GMM 55, the burden rates for the Gaussian distributions 54-1, 54-2, and 54-3 of each co-occurrence corresponding point 51 are calculated, and the sum for each Gaussian distribution 54 voted for by these Gaussian distributions 54-1, 54-2, and 54-3 becomes the MRCoHOG feature amount. Note that hereinafter, when the Gaussian distributions 54-1, 54-2, and 54-3 are not particularly distinguished, they will simply be referred to as the Gaussian distribution 54, and the same will apply to other components.

[0085] Image recognition is performed using the MRCoHOG feature amount generated in this way. However, when directly applying the reference GMM 55 to calculate the burden rate, a computer with high computing power is required. Therefore, when implementing on a device with limited computing resources, conventionally, a burden rate table prepared in advance using the reference GMM 55 was prepared in the memory, and the burden rate for each Gaussian distribution 54 was obtained by referring to this table. This requires a large amount of memory resources and is not suitable for implementing an image recognition device with a small and inexpensive semiconductor device such as an FPGA or an IC chip.

[0086] Therefore, in this embodiment, by implementing an approximation formula of the reference GMM 55 that is easy to calculate in the image processing device 8, the burden rate can be calculated by a simple hardware-oriented calculation using a small number of parameters without referring to the burden rate table. The method will be described below.

[0087] Each figure in FIG. 11 is a figure for explaining the approximation of the reference GMM 55. The ellipses 62-1, 62-2, and 62-3 in FIG. 11(a) are obtained by slicing the Gaussian distributions 54-1, 54-2, and 54-3, which are the original basis functions of the reference GMM 55, at an appropriate height (p(x|θ)) and projecting them onto the xy plane, which is the domain of the reference GMM 55.

[0088] These ellipses 62-1, 62-2, and 62-3 are formed corresponding to the positions of the clusters 60-1, 60-2, and 60-3. These ellipses 62 may be obtained from the Gaussian distribution 54, or alternatively, a shape that encloses the clusters 60 well-balanced may be appropriately set.

[0089] Since the Gaussian distribution 54 is a two-variable normal distribution, the width of the line sliced by a predetermined p(x|θ) reflects the width of the standard deviation of these two variables, and the ellipse 62 is formed such that the major axis (long axis) and the minor axis (short axis) are perpendicular and rotated in an arbitrary direction.

[0090] The reference GMM 55 of the present embodiment uses, as a basis function, an approximation of the Gaussian distribution 54 by a combination of the ellipse 62 and a calculation formula described later. Then, when the parameters defining the individual ellipses 62 formed on the xy plane are substituted into the calculation formula, individual basis functions that approximate the individual Gaussian distributions 54 are formed. This facilitates the calculation of the burden rate using the reference GMM 55.

[0091] The ellipse 62 is represented by Equation (1), and the parameters that the image processing apparatus 8 should store to identify the ellipse 62 are only the coefficients A, B, C for each ellipse 62 and the coordinate values (x0, y0) of the center of the ellipse 62. The required memory is 5×64 = 320 bits per one ellipse 62, and the total memory required for image recognition is as small as about 39.4 KB. Note that the subscript 0 such as x0 is represented by a full-width character to prevent character misconversion. The same applies to other mathematical formulas below.

[0092] Although it is also possible to calculate the burden rate using the ellipse 62 in a state where the major axis is rotated by an arbitrary angle from the coordinate axis of the reference GMM 55, since the calculation becomes complicated, in the present embodiment, as shown in FIG. 11(b), the ellipses 62-1, 62-2, 62-3 are rotated so that the direction of the maximum width (the direction of the major axis) is parallel or perpendicular to the coordinate axis of the reference GMM 55, and the ellipses 63-1, 63-2, 63-3 are set, and the basis function of the reference GMM 55 is configured based on this.

[0093] Whether the maximum width direction is parallel to the x-axis or the y-axis is determined according to the direction with the smaller rotation angle, but the rotation direction may also be determined by experiments. In addition, with rotation, it is also possible to appropriately shape the ellipse 63, such as enlarging or flattening it. According to experiments, there is no significant difference in image recognition accuracy between the case of using the ellipse 62 and the case of using the ellipse 63, and it was confirmed that the ellipse 63 can be used. Thus, the direction of the maximum width of the ellipse used in this embodiment is parallel or perpendicular to the orthogonal coordinate axes that define the mixture Gaussian model.

[0094] The ellipse 63 is represented by Equation (2), and the parameters that the image processing apparatus 8 should store to identify the ellipse 63 are only the coefficients A and B for each ellipse 63 and the coordinate values (x0, y0) of the center of the ellipse 62. The required memory is 4×64 = 256 bits per ellipse 63, and the total memory required for image recognition is about 31.5 KB. Note that the parameters actually used for the calculation of the burden rate are the major axis radius (the width of the Gaussian distribution in the major axis direction), the minor axis radius (the width of the Gaussian distribution in the minor axis direction), and the coordinate values of the center, as will be described later. In this case as well, since the number of parameters to be stored is 4, the memory consumption is the same.

[0095] Next, the calculation formula used for the basis function and the calculation method of the burden rate will be described. Each figure in FIG. 12 is a figure for explaining the parameters and variables used for the calculation of the burden rate. In FIG. 12(a), the center of the ellipse 63-i (the i-th ellipse 63, which is any one of the ellipses 63-1, 63-2, ···, and the same applies to other components hereinafter) is denoted as wi, and the distance between the co-occurrence corresponding point 51 and wi is represented by the distance di_x in the x-axis direction and the distance di_y in the y-axis direction. The distance measured along such coordinate axes is called the Manhattan distance, and the calculation in hardware is easier compared to the Euclidean distance.

[0096] Also, as shown in Fig. 12(b), the radius (width) in the x-axis direction and the radius (width) in the y-axis direction of the ellipse 63-i are expressed as the nth power of 2 (n is an integer of 0 or more, which can also be said to be a power of 2 by an integer of 0 or more, or a power of 2 including the 0th power), and each width is quantized as the ri_xth power of 2 and the ri_yth power of 2. ri_x and ri_y are integers of 0 or more such as 0, 1, 2, ···.

[0097] This quantization is obtained by approximating according to the width quantization table in Fig. 12(c). For example, the radius of the ellipse corresponds to the standard deviation σ which is the width of the Gaussian distribution. When 1 < σ ≤ 2, it is approximated as the first power of 2, when 2 < σ ≤ 4, it is approximated as the second power of 2, ···, and so on. By approximating and quantizing the radius of the ellipse 63 as the nth power of 2 in this way, operations by bit shift become possible.

[0098] Fig. 13 is a diagram for explaining the calculation formula of the burden rate. The burden rate is the posterior distribution of the latent variable z (the distribution of z when the co-occurrence corresponding point 51 is given), and is represented by p(kz = 1|x). The distribution of the co-occurrence corresponding point 51 contributes to the formation of the Gaussian distributions 54-1, 54-2, ···, and since the GMM is a linear sum of Gaussian distributions, the probability density function 53 of the reference GMM55 is constituted by these being stacked (as the total). At that time, the probability (the contributing ratio) that a certain co-occurrence corresponding point 51 belongs to the Gaussian distributions 54-1, 54-2, ··· becomes the burden rate of the corresponding co-occurrence corresponding point 51 with respect to each Gaussian distribution 54.

[0099] Here, in order to facilitate computer calculation, the Gaussian distributions constituting the mixture Gaussian distribution are approximated by si_x,i_y defined by the formula (3) shown in Fig. 13, and the burden rate is approximated by the calculation formula using zi in the formula (4). That is, si_x,i_y defined by the parameters of the ellipse 63-i corresponds to the basis function, and zi corresponds to the calculation formula of the feature amount corresponding to the basis function. These formulas are newly devised to facilitate implementation on hardware. By substituting the co-occurrence distribution and parameters into the formula for zi, the burden rate, which is a feature quantity of an image using a mixture Gaussian model, can be approximately and easily calculated.

[0100] FIG. 14 is a diagram for explaining the basis function in more detail. The expressions (3) and (4) shown in FIG. 13 are the ones that combine the expressions for two variables in the x-axis and y-axis directions into one. For clarity, they are made into expressions for one variable as expressions (5) and (6) in FIG. 14(a).

[0101] As shown in the graph of the figure, zi becomes 1 when the distance di between the co-occurrence corresponding point 51 and the center of the ellipse 63-i is 0, and gradually decreases as di moves away from the center. And zi becomes 1 / 2 when si is 1 (that is, when di = 2 to the power of (ri - log2a)), and gradually approaches 0 as di becomes larger. The spread of zi is defined by the radius ri of the ellipse 63-i, and the smaller ri is, the steeper the shape becomes.

[0102] Note that the term a in log2a with base 2 is a term that defines the calculation accuracy. When hardware implementation is performed, usually a = 8 bits or 16 bits is set. If this term is ignored, zi becomes 1 / 2 when di is equal to the width of the ellipse 63.

[0103] In this way, zi has properties similar to a Gaussian distribution, and the Gaussian distribution can be preferably approximated by the calculation formula. Also, in si, di is divided by 2 to the power of (ri - log2a), but division by 2 to the power of n can be extremely easily performed in hardware by bit shift. Therefore, approximation of the Gaussian distribution can be performed by bit shift by using zi. Therefore, the Gaussian distribution 54-i is approximated by zi, and zi that approximately represents the probability belonging to the Gaussian distribution 54-i is adopted as the burden rate.

[0104] The above defines the calculation formula for the burden rate by Equation (4) in FIG. 13, but it is not limited thereto, and any function that can allocate the ratio of the co-occurrence corresponding point 51 belonging to the Gaussian distribution 54 based on the ellipse 63 can be applied as a basis function.

[0105] For example, as shown in FIG. 14(b), a function where \(z_i = 1\) for \(r_i^n\) with \(0 \leq d_i \lt 2\) and \(z_i = 0\) for \(1 \leq r_i\) (in the case of two dimensions, it is an elliptic cylinder with \(r_{i\_x}^n\) and \(r_{i\_y}^n\) where the radius width is 2), or as shown in FIG. 14(c), a function where \(z_i\) decreases linearly as \(d_i\) increases from 0 to \(r_i\) and \(z_i = 0\) for \(1 \leq 2^{r_i}\) (in two dimensions, it is an elliptic cone with an ellipse having a bottom radius width of \(2^{r_{i\_x}}\) and \(2^{r_{i\_y}}\)), or other functions such as wavelet-type or Gabor-type functions localized on the ellipse 63 can be used. How much these basis functions can be utilized in image recognition is verified by experiments.

[0106] Each figure in FIG. 15 is a figure for explaining the specific calculation of the burden rate. As shown in FIG. 15(a), consider the co-occurrence corresponding point 51 inside the ellipse 63 - i, and obtain the burden rate of this point with respect to the ellipse 63 - i. As shown in FIG. 15(b), let the radius \(2^{r_{i\_x}}\) of the ellipse 63 - i in the x-axis direction be \(2^5\), and the radius \(2^{r_{i\_y}}\) in the y-axis direction be \(2^3\). Also, let the coordinate value of the center \(w_i\) of the ellipse 63 - i be \((10, 25)\), and the coordinate value of the co-occurrence corresponding point 51 be \((25, 20)\).

[0107] As shown in FIG. 15(c), regarding the x-axis direction, \(d_{i\_x} = 15\) and \(r_{i\_x} = 5\). Substituting these into Equation (3) in FIG. 13 and calculating gives \(s_{i\_x} = 3.75\). On the other hand, as shown in the figure, represent \(d_{i\_x}\) by the bit string \((000000001111)\), and shift it by -2 (i.e., shift it 2 bits to the right) to divide it by \(2^2\), then a bit string \((000000000011)\) corresponding to \(s_{i\_x}\) is obtained. When the value represented by this bit string is converted to a decimal number, it becomes 3 as shown in the figure, which is the value obtained by truncating the decimal part of the previously calculated value. Note that the error in the decimal part is ignored.

[0108] As shown in FIG. 15(d), with respect to the y-axis direction, di_y = 5 and ri_y = 3. Substituting these into Equation (3) of FIG. 13 and calculating gives si_y = 5. On the other hand, as shown in the figure, di_y is represented by the bit string (000000000101). When this is shifted right by 0 (i.e., not shifted) to divide by 2 to the power of 0, the bit string (000000000101) corresponding to si_y is obtained. When the value represented by this bit string is converted to a decimal number, it becomes 5 as shown in the figure, which is equal to the previously calculated value.

[0109] Therefore, as shown in FIG. 15(e), the burden rate zi for the Gaussian distribution 54-i of the co-occurrence corresponding point 51 (the Gaussian distribution 54-i corresponding to the ellipse 63-i) is approximated to 0.1406... by adding zi_x and zi_y. Similarly, Equation (4) of FIG. 13 can be applied to the ellipse 63-(i + 1) and other ellipses 63 to calculate the burden rate (approximate value) for these Gaussian distributions 54 of the co-occurrence corresponding point 51.

[0110] In this way, the burden rate for each Gaussian distribution 54 of a certain co-occurrence corresponding point 51 can be calculated. By aggregating (voting) the burden rates obtained by calculating this for all co-occurrence corresponding points 51 for each Gaussian distribution 54, performing this for all feature planes 15 and connecting them, and further normalizing, the MRCoHOG feature amount is obtained.

[0111] FIG. 16 is a diagram for explaining the quantization of the burden rate. After calculating the burden rate by Equation (4) of FIG. 13, the image processing apparatus 8 further saves memory consumption by quantizing this to the nth power of 2. FIG. 16(a) shows an example of the burden rate for the Gaussian distribution 54-i without quantization. In this example, the mixed number K = 6, and i takes values from 1 to 6. When the burden rate is not quantized, for example, the burden rate in the Gaussian distribution 54-1 is 0.4, the burden rate in the Gaussian distribution 54-2 is 0.15, ···, and so on, which results in a 64-bit representation.

[0112] Figure 16(b) shows an example of the quantization table 21 for the burden rate. The quantization table 21 divides the 64-bit representation of the burden rate into eight levels: when it is 0.875 or more, when it is 0.75 or more and less than 0.875, when it is 0.625 or more and less than 0.75, ···, and approximates these to 3-bit representations by shift addition ((addition of 2 to the power of n)) such as (2 to the power of 0) + (2 to the power of -3), (2 to the power of -1) + (2 to the power of -2), ···.

[0113] When the image processing apparatus 8 calculates the burden rate, it refers to the quantization table 21 and approximates it to a 3-bit representation to save memory consumption. According to a preliminary calculation, for example, in the case of a 64-bit representation, 20412 KB of memory is consumed, while in the case of a 3-bit representation, the memory consumption is 319 KB. Also, when the burden rate is quantized into the form of shift addition, subsequent hardware calculations become easier.

[0114] Above, the method of extracting MRCoHOG feature amounts from an image according to the burden rate has been described. However, the feature amounts can be input into an existing classifier such as a neural network that has learned the target in advance to perform image recognition.

[0115] Figure 17 is a diagram showing an example of the hardware configuration of the image processing apparatus 8. The image processing apparatus 8 shown in Figure 17(a) is mounted on a vehicle, for example, and recognizes pedestrians and the like in an image. In this example, the MRCoHOG feature amount is extracted by the semiconductor device 85, and the CPU 81 performs image recognition of the subject. This is an example. Various forms are possible, such as performing image recognition using the semiconductor device 85, or configuring the whole with a small single-board computer equipped with a GPU (Graphics Processing Unit).

[0116] The image processing apparatus 8 is configured by connecting a CPU 81, a ROM 82, a RAM 83, a storage device 84, a semiconductor device 85, an input unit 86, and an output unit 87 via bus lines. The CPU 81 is a central processing unit, operates according to an image recognition program stored in the storage device 84, and performs image recognition processing using the feature amounts extracted by the semiconductor device 85.

[0117] The ROM 82 is a read-only memory and stores basic programs and parameters for operating the CPU 81. The RAM 83 is a readable and writable memory and provides a working memory when the CPU 81 performs image recognition processing.

[0118] The storage device 84 is configured using a large-capacity storage medium such as a hard disk or a semiconductor storage device, and stores an image recognition program. The input unit 86 includes an input device such as accepting an input from an operator, and accepts various operations on the image processing apparatus 8. The output unit 87 includes output devices such as a display and a speaker for presenting various information to the operator, and outputs an operation screen of the image processing apparatus 8, an image recognition result, etc.

[0119] The semiconductor device 85 is configured using, for example, an FPGA, accepts an input of image data (video data), and extracts and outputs feature amounts by MRCoHOG using GMM. As shown in Fig. 17(b), the semiconductor device 85 is composed of an image input unit 91, a plot unit 92, a load factor calculation unit 93, and a feature amount generation unit 94.

[0120] Hereinafter, the procedure of the image recognition process performed by the image processing apparatus 8 will be described using a flowchart. FIG. 18 is a flowchart for explaining the procedure of the image recognition process performed by the image processing apparatus 8. In the following processes, steps 150 to 170 for extracting feature amounts are performed by the semiconductor device 85, and the subsequent image recognition processes (S175 to S185) are performed by the CPU 81.

[0121] First, the image input unit 91 acquires a frame image from the moving image data transmitted from the camera (step 150). In this way, the image processing apparatus 8 includes image acquisition means for acquiring an image. Next, the plotting unit 92 performs the following plotting process on the image, calculates the gradient direction by the gradient direction quantization device 7, etc., and extracts the feature amount due to the co-occurrence of the gradient direction from the frame image (step 155). In this way, the image processing apparatus 8 includes gradient direction acquisition means for acquiring the gradient direction of the luminance of each pixel constituting the image by the gradient direction quantization device 7.

[0122] Next, the load rate calculation unit 93 calculates the load rate for each feature plane 15 of the image (step 160). Then, the feature amount generation unit 94 concatenates the load rates calculated for each feature plane 15 for all the feature planes 15 to obtain a feature amount representing the features of the entire target image (step 165), and normalizes and outputs this (step 170). In this way, the image processing apparatus 8 includes feature amount acquisition means for acquiring the feature amount of the image based on the distribution of the gradient direction, and the feature amount acquisition means acquires the feature amount by the load rate representing the distribution of the co-occurrence of the gradient direction by a mixture Gaussian distribution.

[0123] Next, the CPU 81 inputs the normalized feature amount to a discriminator configured by a neural network or other discrimination mechanism, makes an analog determination of the pedestrian in the frame image from the output value (step 175), and outputs the recognition result by this analog determination to the vehicle control system or the like (step 180). As described above, the image processing apparatus 8 includes image recognition means for recognizing an image using feature amounts. When the CPU 81 continues to track the object to be recognized (step 185; Y), it returns to step 150 and performs recognition using the next feature amount output by the semiconductor device 85. When it does not continue to track (step 185; N), the process ends.

[0124] FIG. 19 is a flowchart for explaining the plot processing procedure in step 155 in FIG. 18. First, the plot unit 92 reads the image processed by the image input unit 91 (step 5). Next, the plot unit 92 divides the image into block regions 3 and stores the positions of the divisions (step 10).

[0125] Next, the plot unit 92 selects one of the block regions 3 of the divided high-resolution image 11 (step 15), and generates and stores pixels of the high-resolution image 11, the middle-resolution image 12, and the low-resolution image 13 for which co-occurrence is to be determined (step 20). When using the image as the high-resolution image 11 as it is, the pixels of the image are used as the pixels of the high-resolution image 11 without resolution conversion.

[0126] Next, the plot unit 92 calculates (see FIG. 8) the gradient direction for each pixel of the generated high-resolution image 11, middle-resolution image 12, and low-resolution image 13 using the gradient direction quantization device 7 and stores it (step 25). Next, the plot unit 92 takes co-occurrence of the gradient direction within the high-resolution image 11, between the high-resolution image 11 and the middle-resolution image 12, and between the high-resolution image 11 and the low-resolution image 13, and plots and stores it on the feature surface 15 (step 30). Thereby, the feature surface 15 by the block region 3A is obtained.

[0127] Next, the plot unit 92 determines whether plotting has been performed for all pixels (step 35). If there are still pixels for which plotting has not been performed (step 35; N), the plotting unit 92 returns to step 20 to select the next pixel and perform plotting on the feature surface 15 for this pixel.

[0128] On the other hand, when plotting has been performed for all the pixels in the block area 3 (step 35; Y), the plotting unit 92 determines whether plotting has been performed for all the block areas 3 (step 40). If there is still a block area 3 for which plotting has not been performed (step 40; N), the plotting unit 92 returns to step 15 to select the next block area 3 and perform plotting on the feature surface 15 for this block area 3. On the other hand, when plotting has been performed for all the block areas 3 (step 40; Y), the plotting unit 92 outputs the feature surface 15 generated for each offset pixel for each of all the block areas 3 (step 45) and returns to the main routine in FIG. 18.

[0129] FIG. 20 is a flowchart for explaining the load rate calculation process in step 160 in FIG. 18. First, the load rate calculation unit 93 selects the feature surface 15 to be processed (step 205). Next, the load rate calculation unit 93 selects the co-occurrence corresponding points 51 from the selected feature surface 15 and stores their coordinate values (step 210).

[0130] Next, the load rate calculation unit 93 initializes the parameter i for counting the ellipse 63-i to 1 and stores it (step 215). Next, the load rate calculation unit 93 reads the coordinate values of the co-occurrence corresponding points 51 stored in step 210 and also reads the parameters of the ellipse 63-i (center coordinate values (x0, y0) and ri_x and ri_y defining the widths of the major and minor axes), substitutes these into equations (3) and (4) in FIG. 13, and calculates an approximate value of the load rate in the Gaussian distribution 54-i (Gaussian distribution corresponding to the ellipse 63-i) of the co-occurrence corresponding point 51. Furthermore, the image processing apparatus 8 quantizes the approximate value of the load factor with reference to the quantization table 21 and stores it as the final load factor (step 220).

[0131] Next, the load factor calculation unit 93 adds the load factor to the total value of the load factors of the Gaussian distributions 54-i and stores it, thereby voting the load factor for the Gaussian distribution 54-i (step 225). Next, the load factor calculation unit 93 increments i by 1 and stores it (step 230), and determines whether the stored i is less than or equal to the number of mixtures K (step 235).

[0132] If i is less than or equal to K (step 235; Y), the load factor calculation unit 93 returns to step 220 and repeats the same process for the next Gaussian distribution 54-i. On the other hand, if i is greater than K (step 235; N), since voting has been performed for all Gaussian distributions 54 with respect to the co-occurrence corresponding point 51, the load factor calculation unit 93 determines whether the load factor has been calculated for all co-occurrence corresponding points 51 on the feature surface 15 (step 240).

[0133] If there is still a co-occurrence corresponding point 51 for which the load factor has not been calculated (step 240; N), the load factor calculation unit 93 returns to step 210 and selects the next co-occurrence corresponding point 51. On the other hand, if the load factor has been calculated for all co-occurrence corresponding points 51 (step 240; Y), the load factor calculation unit 93 outputs the load factor based on the feature surface 15 to the feature amount generation unit 94 (step 243), and further determines whether the voting process for each Gaussian distribution 54 based on the load factor has been performed for all feature surfaces 15 (step 245).

[0134] If there is still a feature surface 15 for which the process has not been performed (step 245; N), the load factor calculation unit 93 returns to step 205 and selects the next feature surface 15. On the other hand, if the process has been performed for all feature surfaces 15 (step 245; Y), the plotting process is terminated and the routine returns to the main routine of FIG. 18.

[0135] FIG. 21 is a graph showing the experimental results of image recognition according to this embodiment. FIG. 21(a) and (b) represent the cases where the number of mixtures K = 16 and 32, respectively. The vertical axis represents the positive detection rate in image recognition, and the horizontal axis represents the false detection rate. Curve A is the case where arctanθ is calculated and no approximation of the gradient direction is performed. Curve B is the case where the fractional bit width is set to 6 by the image processing apparatus 8. Curve C represents the result when 36 directions are classified according to a combination of the ranges of the magnitudes of fx and fy and the ranges of the sum and difference of fx and fy, with many branches. It shows that the higher the correct answer rate (the higher the curve is located), the higher the performance. In any number of mixtures, when the bit width is 6, it shows a positive detection rate close to that when no approximation is performed, and shows better results than when multiple branches are performed. As described above, the image processing apparatus 8 is suitable for implementation on small-scale hardware and can exhibit high image recognition ability.

[0136] As described above, the gradient direction quantization device 7 can perform high-precision gradient direction calculation with fewer branch conditions while suppressing calculation resources by utilizing the correlation between y / x and θ. In addition, the accuracy of the constant (approximate value of tanθ) used for the boundary condition of the gradient direction can be controlled by the number of bits used for quantization, and the amount of calculation resources used can also be controlled by the number of bits.

[0137] In the above example, as an example, quantization is performed in 36 directions every 10°, but it is also possible to perform quantization in more or fewer directions. Since the approximate value of tanθ is used as the boundary, expansion to other angles is easy. Also, when there are 36 directions, the fractional bit width can be set to 4, and when there are 72 directions, the fractional bit width can be set to 7, etc., so that the bit width can be changed in conjunction with the quantization direction. Furthermore, quantization can also be performed in unequal angular ranges such as 0° to 10°, 10° to 30°, 30° to 45°, ···

[0138] Since the calculation algorithm of the gradient direction by the gradient direction quantization device 7 can be implemented by avoiding trigonometric functions and divisions that require a large amount of calculation resources in the calculation of the gradient direction (angle), it can be implemented on a device with limited calculation resources such as an FPGA.

Explanation of Signs

[0139] 2 Images 3 Block Areas 5 Target Pixels 7 Gradient Direction Quantization Device 8 Image Processing Device 11 High-Resolution Image 12 Medium-Resolution Image 13 Low-Resolution Image 15 Feature Surface 21 Quantization Table 51 Co-Occurrence Corresponding Points 53 Probability Density Function 54 Gaussian Distribution 55 Reference GMM 60 Cluster 62 Ellipse 63 Ellipse 81 CPU 82 ROM 83 RAM 84 Storage Device 85 Semiconductor Device 86 Input Section 87 Output Section 71 Rotation Selection Section 72 tanθ Table 73 Multiplication Section 74 Quadrant Selection Section 75 Offset Selection Section 76 Direction Selection Section 77 Addition Section 91 Image Input Section 92 Plotting Section 93 Load Factor Calculation Section 94 Feature Quantity Generation Section 101 Image 102 Cell 106 Histogram 107 HOG feature 110 Target pixel 113 Co-occurrence matrix 117 CoHOG feature 120 High-resolution image 121 Medium-resolution image 122 Low-resolution image 125 Target pixel 127 MRCoHOG feature

Claims

1. Luminance acquisition means for acquiring the luminance of each pixel arranged on the image, Tangent component acquisition means for acquiring fx, which is the denominator component of the tangent of the luminance gradient, and fy, which is the numerator component, for each of the pixels based on the acquired luminance array, Quantization means for quantizing the luminance gradient direction by applying the acquired fx and fy to the range of tangent values corresponding to a predetermined number of quantized gradient directions, Output means for outputting the quantized gradient direction, and comprising, The quantization means, A quadrant selection unit that receives the input of fx and fy, determines and outputs the quadrant of the gradient direction corresponding to the signs of fx and fy among the first quadrant to the fourth quadrant, Receiving fx, fy, -fx and -fy with their signs inverted, and the quadrant of the gradient direction output by the quadrant selection unit, and when the received quadrant is the first quadrant, selecting fx or -fx that is 0 or more and fy or -fy that is 0 or more and outputting them as the rotated fx and fy, and when the received quadrant is in the second quadrant, the third quadrant, or the fourth quadrant, selecting the fx and fy after rotating 90°, 180°, or 270° clockwise respectively from fx, -fx, fy, -fy and outputting them as the rotated fx and fy, a rotation selection unit, Receiving the offset values 0, 9, 18, 27 corresponding to the first quadrant to the fourth quadrant respectively, and an offset selection unit that selects and outputs one of the offset values 0, 9, 18, 27 corresponding to the quadrant determined by the quadrant selection unit, A tanθ table storing the approximate values of tanθ every 10° from 0° to 90°, A multiplication unit that multiplies each approximate value of tanθ every 10° by the rotated fx to calculate and output fxtanθ every 10°, A direction selection unit that outputs the numerical value (n - 1) of the direction corresponding to the nth range where the rotated fy is located among the first to ninth ranges with the fxtanθ every 10° as the boundary value, An addition unit that adds the offset value output by the offset selection unit to the numerical value output by the direction selection unit, And comprising, The approximate value of tanθ is an approximate value with a fixed decimal point where the bit width of the fractional part is 4 bits, 5 bits, or 6 bits, The quantization means is configured by programming a field - programmable gate array (FPGA) composed of a number of look - up tables (LUTs), flip - flops (FFs), etc. The output means outputs the value after addition in the addition unit as a gradient direction quantized in 36 directions. A gradient direction quantization device characterized by the above. **Claim 2** An image acquisition means for acquiring an image, Gradient direction acquisition means for acquiring the gradient direction of the luminance of each pixel constituting the acquired image by the gradient direction quantization device according to Claim 1, Feature amount acquisition means for acquiring a feature amount of the acquired image based on the distribution of the acquired gradient directions, Image recognition means for performing image recognition using the acquired feature amount, An image recognition device characterized by comprising the above. **Claim 3** The feature amount acquisition means acquires the feature amount by a burden rate representing the co-occurrence distribution of the gradient directions by a mixture Gaussian distribution. The image recognition device according to Claim 2, characterized by this. **Claim 4** A luminance acquisition function for acquiring the luminance of each pixel arranged on the image, A tangent component acquisition function for acquiring fx, which is the denominator component of the tangent of the luminance gradient, and fy, which is the numerator component, for each pixel based on the acquired luminance array, A quantization function for quantizing the gradient direction of the luminance by applying the acquired fx and fy to a range of tangent values corresponding to a predetermined number of quantized gradient directions, An output function for outputting the quantized gradient direction, A gradient direction quantization program realized by a computer, comprising: The quantization function: A quadrant selection function that receives the inputs of fx and fy, determines and outputs the quadrant of the gradient direction corresponding to the signs of fx and fy among the first quadrant to the fourth quadrant, Receives fx, fy, -fx and -fy with their signs inverted, and the quadrant of the gradient direction output by the quadrant selection function. When the received quadrant is the first quadrant, it selects fx or -fx that is 0 or more and fy or -fy that is 0 or more and outputs them as the rotated fx and fy. When the received quadrant is in the second quadrant, the third quadrant, or the fourth quadrant, it selects the fx and fy after rotating 90°, 180°, or 270° clockwise, respectively, from fx, -fx, fy, -fy and outputs them as the rotated fx and fy. A rotation selection function, Receives the offset values 0, 9, 18, 27 corresponding to the first quadrant to the fourth quadrant respectively, and selects and outputs one of the offset values 0, 9, 18, 27 corresponding to the quadrant determined by the quadrant selection function. An offset selection function, In a tanθ table storing approximate values of tanθ at 10° intervals from 0° to 90°, a multiplication function that multiplies each approximate value of tanθ at each 10° interval by the rotated fx and calculates and outputs fxtanθ at each 10° interval. A direction selection function that outputs a numerical value (n - 1) in the direction corresponding to the nth range in which the rotated fy is located, among the first to ninth ranges, with the fxtanθ at each 10° interval as a boundary value. An addition function that adds the offset value output by the offset selection function to the numerical value output by the direction selection function. Comprising The approximate value of tanθ is an approximate value with a fixed decimal point where the bit width of the fractional part is 4 bits, 5 bits, or 6 bits. The quantization function is configured by programming a field-programmable gate array (FPGA) composed of a number of look-up tables (LUTs), flip-flops (FFs), etc. The output function outputs the numerical value after addition by the addition function as a gradient direction quantized in 36 directions. A gradient direction quantization program, characterized in that.

5. An image acquisition function for acquiring an image. A gradient direction acquisition function for acquiring the gradient direction of the luminance of each pixel constituting the acquired image by the gradient direction quantization device according to claim 1. A feature amount acquisition function for acquiring the feature amount of the acquired image based on the distribution of the acquired gradient direction. An image recognition function for performing image recognition using the acquired feature amount. An image recognition program, characterized in that it is realized by a computer.

Citation Information

Patent Citations

  • Image processing apparatus

    JP2012133418A

  • Image processing device, and image processing program

    JP2020166342A