Image processing device, image processing system, image processing method and program

The image processing device addresses the trade-off between speed and accuracy by acquiring correspondence information and switching between calculation units based on settings, allowing for efficient detection of object types and ranges.

JP7689181B2Active Publication Date: 2025-06-05MAXELL LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023525895
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-02
Filing Date
2022-06-01
Publication Date
2025-06-05
Estimated Expiration
2042-06-01

AI Technical Summary

Technical Problem

There is a trade-off between processing speed and object detection accuracy in existing image processing technologies, where higher resolution images and more objects detected lead to longer processing times and decreased accuracy.

Method used

An image processing device that detects object types and ranges by acquiring correspondence information associating position coordinates with class likelihoods, and switches between calculation units based on setting information to prioritize either accuracy or processing speed.

Benefits of technology

Enables detection of object types and ranges with appropriate processing speed and accuracy by adjusting the number of classes calculated and using different calculation units depending on the setting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007689181000001
    Figure 0007689181000001
  • Figure 0007689181000002
    Figure 0007689181000002
  • Figure 0007689181000003
    Figure 0007689181000003
Patent Text Reader

Abstract

This image processing device uses image processing to detect the type of an object included in an image and the position coordinates where the object is present, and comprises: an associated information acquisition unit that acquires a first associated information group including a plurality of items of associated information in which position coordinates and the likelihood of a class among a predetermined plurality of classes are associated, said position coordinates indicating a range in which an object is expected to be present in the image, and said class being associated with said range; a configuration information acquisition unit that acquires configuration information relating to the information processing; an extraction unit that, on the basis of the acquired first associated information group and the acquired configuration information, extracts a second associated information group including at least one of a plausible class and position information corresponding to the plausible class; and an output unit that outputs the extracted second associated information group.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an image processing device, an image processing system, an image processing method, and a program. This application claims priority based on Japanese patent application No. 2021-092985, filed on June 2, 2021, and incorporates all of the contents of that application by reference. [Background technology]

[0002] Conventionally, in the technical field of detecting an object contained in an image, there has been a technology that detects the type of object present in the image and the range in the image in which the object exists by image processing. In such a technical field, for example, a technology for improving the speed of object detection is known (for example, see Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2020-205039 A Summary of the Invention [Problem to be solved by the invention]

[0004] It is known that there is a trade-off between processing speed and object detection accuracy. That is, the higher the resolution of the image to be processed, the longer the processing time. Also, the more objects that can be detected, the longer the processing time. According to the above-mentioned technology, there is a problem that the accuracy of object detection decreases as the processing speed increases, and there is also a problem that the processing speed decreases as the accuracy of object detection increases.

[0005] SUMMARY OF THE PRESENT EMBODIMENTS An object of the present invention is to provide an image processing technique that is capable of detecting the type and range of an object contained in an image with appropriate processing speed and accuracy. [Means for solving the problem]

[0006] An image processing device according to one aspect of the present invention is an image processing device that detects the type of object included in an image and the position coordinates where the object exists by image processing, and includes a correspondence information acquisition unit that acquires a first correspondence information group including a plurality of pieces of correspondence information in which position coordinates indicating a range in which the object is expected to exist in the image are associated with a likelihood of a class that is associated with the range among a plurality of predetermined classes, a setting information acquisition unit that acquires setting information related to the image processing, an extraction unit that extracts a second correspondence information group including at least one likely class and position information corresponding to the likely class based on the acquired first correspondence information group and the acquired setting information, and an output unit that outputs the extracted second correspondence information group. and the setting information includes at least information on whether the setting is a first setting that prioritizes accuracy of the class and position coordinates extracted by the extraction unit, or a second setting that prioritizes a processing speed of the extraction unit. .

[0008] Furthermore, in an image processing device according to one embodiment of the present invention, in the process in which the extraction unit extracts the second correspondence information group, the number of classes to be calculated when the setting information is the second setting is smaller than the number of classes to be calculated when the setting information is the first setting.

[0009] In addition, in an image processing device according to one embodiment of the present invention, the extraction unit further includes a switching unit that switches between a first calculation unit that performs calculations to extract the second correspondence information group when the setting information is the first setting, and a second calculation unit that performs calculations to extract the second correspondence information group when the setting information is the second setting, based on the setting information.

[0010] In addition, in an image processing device according to one embodiment of the present invention, the extraction unit further includes a compression unit that compresses the first correspondence information group into a specific class from among the classes included in the first correspondence information group using a predetermined method, and the first calculation unit or the second calculation unit performs calculations to extract the second correspondence information group based on the compressed correspondence information.

[0011] In addition, in an image processing device according to one embodiment of the present invention, the compression unit compresses the correspondence information included in the first correspondence information group when the number of classes in which the likelihood of the correspondence information included in the first correspondence information group is greater than or equal to a predetermined value is less than or equal to a predetermined value.

[0012] In the image processing device according to an aspect of the present invention, the switching unit switches based on the setting information when the image processing device is started.

[0013] In the image processing device according to an aspect of the present invention, the setting information acquisition unit acquires the setting information from a setting file.

[0014] In addition, in the image processing device according to an aspect of the present invention, the setting information acquisition unit acquires the setting information based on the first correspondence information group acquired by the correspondence information acquisition unit.

[0015] Moreover, an image processing system according to one embodiment of the present invention includes a pre-processing device that calculates a first correspondence information group including a plurality of pieces of correspondence information that correspond to position coordinates indicating a range in which an object is expected to exist in the image and a likelihood of a class that corresponds to the range among predetermined classes, and an image processing device according to any one of claims 1 to 9 that acquires the first correspondence information group from the pre-processing device.

[0016] Moreover, an image processing method according to one aspect of the present invention is an image processing method for detecting a type of object included in an image and position coordinates where the object exists by image processing, the image processing method including a correspondence information acquisition step of acquiring a first correspondence information group including a plurality of pieces of correspondence information in which position coordinates indicating a range in which the object is expected to exist in the image are associated with a likelihood of a class that is associated with the range among a plurality of predetermined classes, a setting information acquisition step of acquiring setting information related to the image processing, an extraction step of extracting a second correspondence information group including at least one likely class and position information corresponding to the likely class based on the acquired first correspondence information group and the acquired setting information, and an output step of outputting the extracted second correspondence information group. and the setting information includes at least information on whether a first setting is a setting that prioritizes accuracy of the class and position coordinates extracted by the extraction step, or a second setting that prioritizes a processing speed of the extraction step. .

[0017] Furthermore, a program according to one aspect of the present invention is a program for causing a computer to detect a type of object included in an image and position coordinates where the object exists by image processing, the program including a correspondence information acquisition step of acquiring a first correspondence information group including a plurality of pieces of correspondence information in which position coordinates indicating a range in which an object is expected to exist in the image are associated with a likelihood of a class that is associated with the range among a plurality of predetermined classes, a setting information acquisition step of acquiring setting information related to the image processing, an extraction step of extracting a second correspondence information group including at least one likely class and position information corresponding to the likely class based on the acquired first correspondence information group and the acquired setting information, and an output step of outputting the extracted second correspondence information group. wherein the setting information includes at least information on whether a first setting prioritizes accuracy of the class and position coordinates extracted by the extraction step, or a second setting prioritizes processing speed of the extraction step. . Effect of the Invention

[0018] According to the present invention, the type and range of an object contained in an image can be detected with appropriate processing speed and accuracy. [Brief description of the drawings]

[0019] [Figure 1] FIG. 2 is a diagram for explaining the functional configuration of the image processing system according to the embodiment. [Diagram 2] FIG. 1 is a diagram for explaining an overview of an image processing system according to an embodiment. [Diagram 3] FIG. 2 is a block diagram illustrating an example of a functional configuration of a post-process according to the embodiment. [Figure 4] FIG. 2 is a block diagram illustrating an example of a functional configuration of an extraction unit according to the embodiment. [Diagram 5] 11 is a flowchart illustrating an example of a series of operations of a post-process according to the embodiment. [Figure 6] FIG. 13 is a block diagram for explaining a modified example of the functional configuration of an extraction unit according to the embodiment. [Figure 7] FIG. 1 is a diagram for explaining an overview of an example of an imaging system according to an embodiment. [Figure 8] FIG. 13 is a diagram for explaining an overview of a modified example of an imaging system according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0020] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. The embodiment described below is merely an example, and the embodiment to which the present invention is applied is not limited to the following embodiment. In this embodiment, the description is based on the premise that there is a trade-off relationship between object detection accuracy and processing speed. Here, the object detection accuracy may have a trade-off relationship with power consumption, required resources, and the like in addition to the processing speed. In the following description, an example of the processing speed among the performance indexes having a trade-off with the object detection accuracy will be described, but this example does not limit the present embodiment and includes a plurality of performance indexes having a trade-off relationship with the object detection accuracy.

[0021] [Image processing system overview] 1 is a diagram for explaining the functional configuration of an image processing system according to an embodiment of the present invention, with reference to which an image processing system 1 according to the embodiment of the present invention will be explained. Based on an input image P, the image processing system 1 detects the type of object included in the image P and the position coordinates of the range in which the object exists by image processing. The image processing system 1 outputs an object detection result O as a result of the image processing. The object detection result O includes the type of object included in the image P and the position coordinates of the range in which the object exists. When the image P includes multiple objects, the object detection result O includes the types of multiple objects included in the image P and the position coordinates of the range in which each object exists. The image processing according to the present embodiment includes, as an example, machine learning processing. In particular, as one form, it may include a deep neural network (DNN) that repeatedly performs convolution calculations with predetermined weights in multiple processing layers.

[0022] Here, the type of object included in the image P is also referred to as a class. The types of classes that the image processing system 1 can detect are determined in advance. In this embodiment, the image processing system 1 is described as having learned detectable classes in advance. Specifically, a class may be an animal such as a human or a dog, an object such as a car or a bicycle, or a natural object such as a cloud or the sun.

[0023] The image processing system 1 includes a pre-processing 10 and a post-processing 30. The image processing system 1 calculates candidates for the type of object included in the input image P and candidates for the position coordinates where the object exists by using the DNN included in the pre-processing 10, and extracts the most likely class and position coordinates from the calculated candidates by using the post-processing 30. When the image processing system 1 includes a DNN, it may be a trained model that acquires various parameters by learning. The image processing system 1 can be realized by a processor executing various programs stored in a non-volatile memory, but some of the processing of the pre-processing 10 or the post-processing 30 may be implemented as a hardware accelerator.

[0024] The number of pixels of image P input to image processing system 1 is preferably the number of pixels based on the processing unit in which preprocessor 10 performs processing. The processing unit of preprocessor 10 is also referred to as an element matrix. Preprocessor 10 divides the number of pixels of image P into element matrices and performs processing for each element matrix. For example, if the size of the element matrix is ​​16 x 12 [px (pixels)] and the number of pixels of image P is 256 x 192 [px], preprocessor 10 divides image P into 256 and performs processing for each 16 x 12 [px] element matrix. The number of pixels of image P that can be processed by image processing system 1 does not have to depend on the size of the element matrix. Even if the number of pixels of image P is an arbitrary value, for example, by converting the number of pixels of image P into a number of pixels based on the size of the element matrix by preprocessor 10 or in a predetermined process before input to preprocessor 10, processing by preprocessor 10 becomes possible.

[0025] For example, a case will be described in which image processing is performed by software before image P is input to the pre-processor 10. The software processing before image P is input to the pre-processor 10 broadly includes processing for improving image quality, processing of the image itself, and other data processing. Processing for improving image quality may be brightness / color conversion, black level adjustment, noise reduction, correction of optical aberration, etc. Processing of the image itself may be processing such as image cropping, enlargement / reduction / transformation, etc. Other data processing may be data processing such as gradation reduction, compression encoding / decoding, data duplication, etc.

[0026] The pre-process 10 calculates, for each element matrix, position coordinates indicating the range in which an object is expected to exist and the likelihood of the class corresponding to the position coordinates. The range of position coordinates calculated by the pre-process 10 is larger than the element matrix. That is, the pre-process 10 takes into account the entire image P and calculates position coordinates by associating the range in which an object is expected to exist with each element matrix. The position coordinates are expressed in a format that allows for identification of a range with each element matrix as a reference point. Each element matrix is ​​associated with a likelihood for each class, i.e., each element matrix is ​​associated with a number of likelihoods corresponding to the number of classes to be calculated.

[0027] Information that associates position coordinates indicating a range in an image where an object is expected to exist with the likelihood of a class that corresponds to that range among predetermined classes is also referred to as correspondence information. The pre-process 10 calculates correspondence information in a number corresponding to the number of element matrices based on the image P. The multiple pieces of correspondence information calculated by the pre-process 10 are also referred to as a first correspondence information group RI1. In other words, the pre-process 10 calculates a first correspondence information group RI1 that includes multiple pieces of correspondence information. The preprocessor 10 is also referred to as a preprocessing device.

[0028] All or part of the functions of the preprocessor 10 may be a deep learning accelerator realized by using hardware such as an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA). By being realized by hardware, the functions of the preprocessor 10 can quickly calculate candidates for the type of object contained in the image P and candidates for the position coordinates where the object exists. The DNN calculation process included in the preprocessor 10 requires repeated execution of a large number of calculations according to the number of element matrices for each of the multiple layers included. On the other hand, since the content of the calculations is often limited and has low dependency on the application, it is preferable to apply calculation processing using an accelerator with high processing speed rather than performing program processing on a flexible processor.

[0029] The post-process 30 detects the type of object included in the image and the position coordinates where the object exists by image processing based on the first correspondence information group RI1 calculated by the pre-process 10. Specifically, the post-process 30 first acquires the first correspondence information group RI1 from the pre-process 10. The post-process 30 calculates a second correspondence information group RI2 based on the acquired first correspondence information group RI1. The second correspondence information group RI2 is information that includes at least one plausible class and position information corresponding to the plausible class from the information included in the first correspondence information group RI1. The post-processor 30 is also referred to as an image processing device.

[0030] All or part of the functions of the post-processor 30 may be realized using a storage device such as a central processing unit (CPU) (not shown), a read only memory (ROM) or a random access memory (RAM) connected via a bus. The post-processor 30 functions as a device having the functions of the post-processor 30 by executing an image processing program. The image processing program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, magneto-optical disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into a computer system. The image processing program may be transmitted via an electric communication line. The content of the calculations included in the post-process 30 is highly application-dependent compared to the pre-process 10. Furthermore, since it is necessary to switch processing depending on the user settings and the required application, program processing on a highly flexible processor is preferable. Note that it is not necessary for all processing in the post-process 30 to be program processing, and some processing may be performed on an accelerator.

[0031] Fig. 2 is a diagram for explaining an overview of an image processing system according to an embodiment. Processing of the image processing system 1 according to an embodiment will be explained with reference to the diagram. Fig. 2(A) shows an element matrix at a stage before processing by the pre-processing 10, Fig. 2(B) shows a first correspondence information group RI1 calculated by the pre-processing 10, and Fig. 2(C) shows a second correspondence information group RI2 calculated by the post-processing 30.

[0032] First, with reference to Fig. 2(A), an explanation will be given of the element matrix, which is the stage before being processed by the preprocessor 10. The figure shows an example of a case where an image P is divided into a total of 169 element matrices, 13 vertically and 13 horizontally. In this example, for example, the number of pixels of the input image is 208 x 156 [px], and the size of the element matrix is ​​16 x 12 [px]. The pre-processor 10 performs processing for each element matrix. Based on the pixel information of each element matrix and the pixel information of the entire image P, the pre-processor 10 calculates candidates for the type of object contained in the image P and candidates for the position coordinates indicating the range in which the object exists.

[0033] Next, the first correspondence information group RI1 calculated by the pre-processing 10 will be described with reference to Fig. 2(B). As shown in Fig. 2(B), in the first correspondence information group RI1, a plurality of ranges are indicated by rectangles associated with an element matrix. Each rectangle indicates a candidate range in which some object exists. Also, each rectangle is associated with the likelihood of a class to be calculated. When there are a plurality of classes to be calculated, each rectangle is associated with the likelihood of each of the plurality of classes.

[0034] Next, the second correspondence information group RI2 calculated by the post-process 30 will be described with reference to Fig. 2(C). As shown in Fig. 2(C), in the second correspondence information group RI2, a likely range is identified from among the multiple ranges calculated by the pre-process 10. Furthermore, a specific class is associated with each range. That is, the post-process 30 identifies a likely candidate from among the multiple rectangle candidates included in the first correspondence information group RI1 and the one or more class candidates corresponding to each rectangle.

[0035] [Post-processing function configuration] 3 is a block diagram for explaining an example of the functional configuration of a post-processor according to an embodiment. The functional configuration of the post-processor 30 will be explained with reference to the same figure. The post-processor 30 acquires a setting file SF from an input device ID, in addition to acquiring a first correspondence information group RI1 from the pre-processor 10. The input device ID may be an input device such as a touch panel, a mouse, or a keyboard, or may be an information recording medium such as a USB memory. The setting file SF may be an electronic file that includes predetermined setting information. The post-processor 30 includes a correspondence information acquisition unit 310 , a setting information acquisition unit 320 , an extraction unit 330 , and an output unit 340 . In this embodiment, an example is shown in which an input device ID is used to obtain the setting file SF, but the present invention is not limited to this. For example, the setting file SF may be obtained based on the time or a predetermined cycle, or the setting file SF may be obtained based on the first correspondence information group RI1 or the second correspondence information group RI2.

[0036] The correspondence information acquisition unit 310 acquires a first correspondence information group RI1 from the pre-process 10. The first correspondence information group RI1 includes a plurality of pieces of correspondence information. The correspondence information is information in which position coordinates indicating a range in which an object is expected to exist in an image P are associated with the likelihood of a class, among a plurality of predetermined classes, that is associated with the range in which the object is expected to exist. That is, the post-process 30 acquires a first correspondence information group that includes a plurality of pieces of correspondence information in which position coordinates indicating a range in which an object is expected to exist in an image are associated with the likelihood of a class, among a plurality of predetermined classes, that is associated with the range.

[0037] The setting information acquisition unit 320 acquires the setting information SI from the input device ID. The setting information SI is information included in the setting file SF, and is information related to image processing. That is, the setting information acquisition unit 320 acquires the setting information SI related to image processing included in the setting file SF. The setting information SI also includes information for setting whether to prioritize the detection accuracy of the class and position coordinates (accuracy priority) or the processing speed (speed priority). The setting that prioritizes accuracy is also referred to as the first setting, and the setting that prioritizes speed is also referred to as the second setting. Specifically, the first setting prioritizes the accuracy of the class and position coordinates extracted by the extraction unit 330, and the second setting prioritizes the processing speed of the extraction unit 330. In other words, the setting information includes at least information on whether the setting is the first setting that prioritizes the accuracy of the class and position coordinates extracted by the extraction unit 330, or the second setting that prioritizes the processing speed of the extraction unit 330.

[0038] The setting information SI acquired by the setting information acquisition unit 320 may be derived from the first correspondence information group RI1 calculated by the pre-process 10. For example, when classes with high likelihood are limited among the classes included in the first correspondence information group RI1 calculated by the pre-process 10, the setting information SI may be configured to prioritize speed and to limit it to classes with high likelihood. In this case, by not performing calculations for classes with low likelihood, there is a risk of a decrease in detection accuracy, but the processing speed can be increased. That is, in this example, the setting information acquisition section 320 acquires the setting information SI based on the first corresponding information group RI1 acquired by the corresponding information acquisition section 310.

[0039] The extraction unit 330 acquires the first correspondence information group RI1 from the correspondence information acquisition unit 310, and acquires the setting information SI from the setting information acquisition unit 320. The extraction unit 330 extracts the second correspondence information group RI2 based on the acquired first correspondence information group RI1 and setting information SI. The second correspondence information group RI2 includes at least one plausible class and location information corresponding to the plausible class. That is, the extraction unit 330 extracts the second correspondence information group RI2 including at least one plausible class and location information corresponding to the plausible class, based on the first correspondence information group RI1 acquired by the correspondence information acquisition unit 310 and the setting information acquired by the setting information acquisition unit 320.

[0040] The output unit 340 outputs the second correspondence information group RI2 extracted by the extraction unit 330. The output unit 340 outputs the second correspondence information group RI2 in an image format or in a predetermined file format.

[0041] 4 is a block diagram for explaining an example of the functional configuration of the extraction unit according to the embodiment. The functional configuration of the extraction unit 330 will be explained with reference to the same figure. The extraction unit 330 includes a switching unit 332, a first calculation unit 333, a second calculation unit 334, and a calculation result output unit 335.

[0042] The first calculation unit 333 performs a process of calculating the second correspondence information group RI2, prioritizing the accuracy of the class and the position coordinates. Specifically, the first calculation unit 333 identifies the class with high accuracy by extracting a likely class based on the likelihood of the class included in the first correspondence information group RI1. In addition, the first calculation unit 333 identifies the position coordinates with high accuracy by performing calculation based on the resolution of the acquired first correspondence information group RI1. The first calculation unit 333 performs calculations to extract a second correspondence information group when the setting information SI is the first setting.

[0043] The second calculation unit 334 prioritizes processing speed and performs processing to calculate the second correspondence information group RI2. Specifically, the second calculation unit 334 quickly identifies a class by extracting a likely class by limiting the likelihood of classes included in the first correspondence information group RI1 to a specific class. In addition, the second calculation unit 334 quickly identifies position coordinates by performing calculation based on a resolution lower than the resolution of the acquired first correspondence information group RI1. The second calculation unit 334 performs calculations to extract a second correspondence information group when the setting information SI is the second setting.

[0044] The switching unit 332 switches whether the processing is to be performed by the first calculation unit 333 or the second calculation unit 334. The switching unit 332 switches to the first calculation unit 333 when the setting information SI is the first setting, and switches to the second calculation unit 334 when the setting information SI is the second setting, based on the setting information SI. That is, the switching unit 332 switches between the first calculation unit 333 that performs calculations to extract the second correspondence information group RI2 when the setting information SI is the first setting, and the second calculation unit 334 that performs calculations to extract the second correspondence information group RI2 when the setting information SI is the second setting, based on the setting information SI.

[0045] Note that the first setting prioritizing accuracy may have many classes as the objects of calculation, and the second setting prioritizing speed may have a small number of classes as the objects of calculation. That is, in the process in which the extracting unit 330 extracts the second correspondence information group RI2, the number of classes as the objects of calculation when the setting information SI is in the second setting may be smaller than the number of classes as the objects of calculation when the setting information SI is in the first setting.

[0046] The switching unit 332 switches between the first calculation unit 333 and the second calculation unit 334 based on the setting information SI when the post-process 30 is started. Specifically, when the post-process 30 is realized by software, the setting information SI may be acquired by reading the setting file SF after a reset process, and the switching unit 332 may switch between the first calculation unit 333 and the second calculation unit 334. Alternatively, the switching unit 332 may switch to the first calculation unit 333 or the second calculation unit 334 at any timing. The arbitrary timing may be, for example, the timing at which the detection target is switched.

[0047] The calculation result output section 335 outputs the second correspondence information group RI2 extracted by the first calculation section 333 or the second calculation section 334 to the output section 340 as the calculation result.

[0048] In the present embodiment, an example has been described in which the extraction unit 330 includes two calculation units, the first calculation unit 333 and the second calculation unit 334, but the present invention is not limited to this example, and the extraction unit 330 may include three or more calculation units. As another example, in the case where the extraction unit 330 includes a configuration in which a plurality of calculation units are connected in series, it is also possible to control the extraction unit 330 to bypass and omit some of the connected calculation units. When the extraction unit 330 includes a plurality of calculation units, each calculation unit may have a different setting for calculating the second correspondence information set RI2. For example, each calculation unit may have a different number of classes or types of classes to be calculated depending on whether detection accuracy or processing speed is to be prioritized. In addition, the multiple calculation units may have different calculation methods. For example, a calculation unit that prioritizes speed may integrate multiple calculations or skip some calculations compared to a calculation unit that prioritizes accuracy. A configuration may be used to prioritize accuracy or speed by using different thresholds for calculation.

[0049] Here, the threshold value used in the calculation will be described. Conventionally, since the calculation result for each bounding box can take a value range of (-∞, +∞), the calculation result is multiplied by a sigmoid function to normalize it to a value range of (0, 1) and calculate the likelihood. The calculated likelihood is compared with the likelihood threshold. That is, conventionally, the likelihood is calculated by multiplying each of the multiple calculation results corresponding to each bounding box by a sigmoid function, and the calculated likelihood is compared with the threshold. Therefore, conventionally, the number of calculations is large because each of the multiple calculation results is multiplied by the sigmoid function each time. When the image processing system 1 is applied to an edge device, it is preferable to have a small number of calculations in order to reduce the processing load.

[0050] According to this embodiment, instead of normalizing the calculation result each time, a calculation is performed on a threshold in advance, so that it is not necessary to normalize the calculation result each time. The calculation on the threshold may be, for example, multiplication by an inverse function of a function used for normalization. As a specific example, instead of multiplying the calculation result for each bounding box by a sigmoid function, a likelihood threshold is multiplied in advance by a logit function, which is an inverse function of the sigmoid function, and the likelihood threshold multiplied by the logit function is compared with the calculation result for each bounding box. That is, according to this embodiment, since the threshold for calculating the likelihood can be determined in advance by calculation or the like, by applying a predetermined function value (for example, an inverse function of a function used for normalization) to the threshold, calculation for each of the multiple calculation results corresponding to each bounding box becomes unnecessary. Therefore, according to this embodiment, the processing load can be reduced. In particular, when the pre-process 10 is configured by hardware, the circuit scale can be reduced. Since the circuit scale of the pre-process 10 can be reduced, when the image processing system 1 is applied to an edge device, the processing load can be reduced and the product size can be further reduced. In this embodiment, the calculation for the threshold value is not limited to multiplying the inverse function of the function used for normalization. For example, the threshold value may be multiplied by a predetermined scaling coefficient, or an offset value may be added.

[0051] [Post-processing sequence] 5 is a flowchart for explaining an example of a series of operations of the post-processing according to the embodiment, which will be described with reference to the drawing.

[0052] (Step S110) The correspondence information acquisition unit 310 acquires a first correspondence information group RI1 which is an output result from the pre-processing unit 10. The correspondence information acquisition unit 310 may acquire information obtained by converting the first correspondence information group RI1 into a predetermined format that can be processed by the post-processing unit 30.

[0053] (Step S120) The post-process 30, by using a conversion unit (not shown), converts the acquired first corresponding information group RI1 into a format processable by the post-process 30. For example, the conversion unit performs a process of returning the acquired first corresponding information group RI1 to a high-dimensional API.

[0054] (Step S130) The extraction unit 330 selects likely coordinates based on candidates of position coordinates where an object exists, which are included in the acquired first correspondence information group RI1. Here, the position coordinates where an object exists are also referred to as a bounding box. That is, the first correspondence information group RI1 includes multiple bounding box candidates, and the extraction unit 330 extracts a likely bounding box from the multiple bounding box candidates. The extraction unit 330 extracts a likely bounding box by integrating or deleting the multiple bounding box candidates using a method such as NMS (Non-Maximum Suppression).

[0055] (Step S140) The extraction unit 330 identifies a class corresponding to the extracted bounding box based on the likelihood included in the acquired first correspondence information group RI1. For example, the extraction unit 330 identifies the class corresponding to the bounding box by comparing the likelihood included in the first correspondence information group RI1 with a predetermined threshold, or by ranking the likelihood and then identifying a top-ranked class using a predetermined method.

[0056] (Step S150) The processes in steps S130 and S140 are performed for each element matrix. After steps S130 and S140 are performed for all element matrices in image P, the extraction unit 330 integrates the processes performed for each element matrix. As a result of the integration, the extraction unit 330 generates a bounding box and a likelihood for the entire image P.

[0057] (Step S160) Extraction unit 330 extracts a likely bounding box from the integrated bounding boxes, and extracts a class associated with the extracted boundary. The extraction of the class is performed based on the likelihood after integration.

[0058] (Step S170) The output unit 340 outputs the position coordinates of the extracted bounding box and the class associated with the bounding box.

[0059] [Modification of the extraction part] 6 is a block diagram for explaining a modified example of the functional configuration of the extraction unit according to the embodiment. With reference to the same figure, an extraction unit 330A which is a modified example of the extraction unit 330 will be explained. The extraction unit 330A differs from the extraction unit 330 in that it includes a compression unit 331. The configurations already explained in the extraction unit 330 may be denoted by the same reference numerals and explanations thereof may be omitted.

[0060] The compression unit 331 compresses the size of the element matrix of the first correspondence information group RI1 based on the setting information SI. For example, the compression unit 331 compresses the size of the element matrix of the first correspondence information group RI1 so as to extract a likely class by limiting the likelihood of the classes included in the first correspondence information group RI1 to a specific class or the class with the highest likelihood. At this time, the compression unit 331 uses a method such as Max Pooling to integrate or delete multiple bounding box candidates by a method such as NMS (Non-Maximum Suppression). That is, the compression unit 331 compresses the classes included in the first correspondence information group RI1 to a specific class by a predetermined method. Here, each element matrix is ​​associated with a position coordinate of a bounding box and a class. Information associated with each element matrix is ​​included in a first correspondence information group RI1 as correspondence information RI. The compression unit 331 may compress the correspondence information RI included in the first correspondence information group RI1.

[0061] The first calculation unit 333 or the second calculation unit 334 performs calculations to extract a second correspondence information group RI2 based on the correspondence information RI compressed by the compression unit 331. By performing calculations based on the compressed correspondence information RI, high-speed processing can be achieved. Furthermore, by compressing the first correspondence information group RI1 at a stage prior to the post-processing 30, the overall processing load can be significantly reduced. The compression unit 331 may be included in a conversion unit (not shown) described with reference to FIG.

[0062] In addition to or instead of based on the setting information SI, the compression unit 331 may determine whether to compress the element matrix based on the number of classes in which the likelihood of the correspondence information RI included in the first correspondence information group RI1 is equal to or greater than a predetermined value. For example, when the number of classes in which the likelihood of the plurality of pieces of correspondence information RI included in the first correspondence information group RI1 is equal to or greater than a predetermined value is equal to or less than a predetermined value, the compression unit 331 compresses the correspondence information RI included in the first correspondence information group RI1.

[0063] [Imaging system overview] Next, an example of an imaging system using the image processing system 1 according to this embodiment will be described with reference to Fig. 7 and Fig. 8. The image processing system 1 is configured to process an image captured in real time, for example, and feed back the result of the image processing to hardware.

[0064] The imaging system described with reference to Fig. 7 and Fig. 8 includes an imaging device to capture an image of an object, and analyzes the captured image by an image processing system 1. The imaging system is installed, for example, inside or outside a facility such as a store or public facility, and is installed as a surveillance camera (security camera) that monitors people's behavior. The imaging system may also be installed on the windshield or dashboard of a vehicle such as an automobile, and used as a drive recorder that records the situation when driving or when an accident occurs. The imaging system may also be installed in a moving object such as a drone or an AGV (Automated Guided Vehicle).

[0065] 7 is a diagram for explaining an overview of an example of an imaging system according to an embodiment. With reference to the figure, an example of an imaging system 2 will be explained. The imaging system 2 captures an image of an object using an imaging device, and analyzes the captured image using an image processing system 1. At this time, the image processing system 1 performs image processing further based on predetermined information obtained from the imaging device 50. The imaging system 2 includes an image processing system 1 and an imaging device 50. The imaging device 50 includes a camera 51 and a sensor 52.

[0066] The camera 51 captures an image of a target object. The target object broadly includes animals, objects, and other objects that can be detected by image processing. The sensor 52 acquires information indicating the state of the imaging device 50 itself or information about the surroundings of the imaging device 50. The sensor 52 may be, for example, a battery remaining amount sensor that detects the remaining amount of a battery (not shown) included in the imaging device 50. The sensor 52 may also be an environmental sensor that detects information about the surrounding environment of the imaging device 50. The environmental sensor may be, for example, a temperature sensor, a humidity sensor, an illuminance sensor, an air pressure sensor, a noise sensor, or the like. In addition, when the image processing system 1 is used in a moving object such as a drone, the sensor 52 may be a sensor for detecting the state of the moving object, that is, an acceleration sensor, an altitude sensor, or the like. The sensor 52 outputs the acquired information as detection information DI to the image processing system 1. The detection information DI may be associated with the image P.

[0067] The image processing system 1 acquires an image P captured by a camera 51 and detection information DI detected by a sensor 52. The pre-process 10 calculates a first correspondence information group RI1 based on the image P. The post-process 30 calculates a second correspondence information group RI2 based on the calculated first correspondence information group RI1 and the detection information DI. In this embodiment, the post-process 30 can perform image processing at an appropriate processing speed and accuracy by calculating the second correspondence information group RI2 based on the detection information DI. That is, when the sensor 52 is a battery sensor, the post-process 30 can perform image processing in a mode that does not consume the battery by lowering the accuracy based on the battery capacity when the remaining battery power is low. Also, when the sensor 52 is an environmental sensor, the post-process 30 can perform image processing more efficiently by performing image processing in a mode limited to a class expected according to the situation of the acquired image P. Also, when the sensor 52 is a sensor for detecting the state of a moving object, the post-process 30 can perform image processing more efficiently by performing image processing in a mode limited to a class expected according to the position and direction in which the moving object is facing.

[0068] 8 is a diagram for explaining an overview of a modified example of the imaging system according to the embodiment. An example of the imaging system 3 will be explained with reference to the diagram. The imaging system 3 includes an imaging device to capture an image of an object, analyzes the captured image using an image processing system 1, and controls the imaging device based on the analysis result. The imaging system 3 includes an image processing system 1 and an imaging device 50A. The imaging device 50A includes a camera 51 and a driving device 53.

[0069] The camera 51 captures an image of a target object. The target object broadly includes animals, objects, and other objects that can be detected by image processing. The driving device 53 controls imaging conditions such as the imaging direction, the angle of view, and the imaging magnification of the camera 51. In addition, when the imaging system 3 is used in a moving body such as a drone or an AGV, the driving device 53 controls the movement of the moving body such as the drone or the AGV.

[0070] The image processing system 1 calculates a second correspondence information group RI2 based on the image P captured by the imaging device 50A. The image processing system 1 outputs the calculated second correspondence information group RI2 to the imaging device 50A. The driving device 53 controls the imaging conditions of the camera 51 and the movement of the moving object based on the acquired second correspondence information group RI2. For example, when the imaging system 3 is used in a surveillance camera, if the class and position coordinates of a person suspected to be a criminal are specified by the second correspondence information group RI2, the imaging device 50A can control the imaging direction, angle of view, imaging magnification, etc. of the imaging device 50A so as to track the criminal. Also, when the imaging system 3 is used in a mobile object such as a drone or an AGV, the driving device 53 can control the movement so as to track the person suspected to be a criminal while capturing an image. Also, by displaying the class specified by the second correspondence information group RI2 on a display unit or the like, or by transferring and storing data including the second correspondence information group RI2 to an external server device, it can be utilized for various applications.

[0071] [Summary of the embodiment] According to the embodiment described above, the image processing system 1 includes a pre-processing unit 10 and a post-processing unit 30. The image processing system 1 calculates a plurality of bounding box candidates and the likelihood of the class corresponding to each bounding box by the pre-processing unit 10 realized by hardware such as an FPGA. The image processing system 1 also identifies a likely bounding box from the calculated candidates and a class corresponding to the bounding box by the post-processing unit 30 realized by software. Therefore, according to the present embodiment, the process of extracting bounding box candidates, which is a process with a large amount of processing, is performed by hardware, and the process of identifying a likely bounding box and a class from the extracted candidates is performed by software. Therefore, according to this embodiment, by selecting whether to prioritize accuracy or speed in software processing, it is possible to detect the type of object contained in an image with appropriate processing speed and accuracy. In addition, when the pre-process 10 includes a DNN, its parameters are determined in advance by learning using teacher data. In learning, it is preferable to learn not only the pre-process 10 but also the post-process 30. Therefore, since the post-process 30 of this embodiment has multiple calculation units, learning must be performed in each calculation unit. However, if a lot of time is required for learning, learning may be limited to some of the calculation units. In this embodiment, learning is performed using a calculation unit that prioritizes accuracy, so that the decrease in accuracy when speed is prioritized can be suppressed.

[0072] Furthermore, according to the embodiment described above, the post-process 30 is provided with the correspondence information acquisition unit 310 to acquire the first correspondence information group RI1, and is provided with the setting information acquisition unit 320 to acquire the setting information SI. The post-process 30 is provided with the extraction unit 330 to extract the second correspondence information group RI2 based on the acquired first correspondence information group RI1 and setting information SI. In other words, the extraction unit 330 performs image processing based on the information set by the setting information SI. Therefore, according to this embodiment, the post-process 30 can easily detect the type of object included in an image with appropriate processing speed and accuracy. In particular, it is preferable to apply the image processing method described in this embodiment when the hardware accelerator that executes the pre-processing 10 uses a quantized DNN of 8 bits or less. More specifically, by performing arithmetic processing on the quantized DNN on the accelerator, it is possible to achieve both processing speed and accuracy compared to processing with multi-bit fixed point. However, since the output of the post-processing 30 is subject to further processing at a later stage, it is preferable to process it as a multi-bit fixed point, and these processes become a major problem in edge devices with small processor processing power, and the effect of using an accelerator for processing the pre-processing 10 is reduced. In response to this, the extraction unit 330 performs image processing based on the information set by the setting information SI. Therefore, according to this embodiment, the post-processing 30 can easily detect the type of object contained in the image with appropriate processing speed and accuracy.

[0073] Moreover, according to the embodiment described above, the setting information SI includes at least information on whether the first setting is a setting that prioritizes accuracy or a setting that prioritizes processing speed. Therefore, according to this embodiment, the user of the image processing system 1 can easily set whether to prioritize accuracy or processing speed. Also, according to this embodiment, the user can arbitrarily switch between prioritizing accuracy and processing speed.

[0074] Furthermore, according to the embodiment described above, in the post-process 30, the number of classes to be calculated in the first setting is different from the number of classes to be calculated in the second setting. Furthermore, the number of classes to be calculated in the first setting is greater than the number of classes to be calculated in the second setting. That is, according to this embodiment, by changing the number of classes to be calculated, it is possible to switch between prioritizing accuracy and processing speed. Therefore, according to this embodiment, the post-process 30 can easily switch between prioritizing accuracy and processing speed.

[0075] Furthermore, according to the embodiment described above, the post-process 30 uses different calculation units in the first setting and the second setting. That is, the extraction unit 330 prepares two different calculation units, and the switching unit 332 switches the calculation unit used for the calculation. In other words, the post-process 30 has a program used in the first setting and a program used in the second setting, and the switching unit 332 switches between each program based on the setting information SI. Therefore, according to this embodiment, it is possible to quickly switch between the first setting and the second setting.

[0076] Furthermore, according to the embodiment described above, the extraction unit 330A includes the compression unit 331, and compresses the first correspondence information group RI1 calculated by the pre-process 10 using a technique such as Max Pooling. The calculation unit performs calculations based on the compressed first correspondence information group RI1. Therefore, according to this embodiment, unnecessary processing can be reduced, and the processing speed can be easily increased.

[0077] Furthermore, according to the embodiment described above, when the number of classes is small, the compression unit 331 compresses the first correspondence information group RI1 calculated by the pre-processing 10 using a method such as Max Pooling before the post-processing. Therefore, according to this embodiment, image processing can be performed at high speed.

[0078] According to the embodiment described above, the post-process 30 acquires the setting information SI at the time of startup. Therefore, according to the embodiment, the post-process 30 can easily switch between prioritizing accuracy and processing speed.

[0079] According to the embodiment described above, the setting information acquisition unit 320 acquires the setting information SI from the setting file SF. Therefore, according to the embodiment, the post-processor 30 can easily switch between the accuracy and the processing speed to be prioritized according to the user's settings.

[0080] According to the embodiment described above, the setting information acquisition unit 320 acquires the setting information SI based on the first correspondence information group RI1. Therefore, according to the embodiment, even if the setting information SI is not set by the user, image processing can be performed with appropriate accuracy or processing speed based on the first correspondence information group RI1.

[0081] Furthermore, according to the embodiment described above, the image processing system 1 performs software processing on the image P before the image P is input to the pre-processor 10. The image processing performed by the image processing system 1 includes, for example, processing for improving image quality, processing of the image itself, and other data processing. Here, when the pre-processor 10 is configured with hardware such as an FPGA, the pre-processor 10 may not be able to process the image P depending on the image quality, image size, image format, etc. of the image P. Therefore, according to the present embodiment, by performing software processing on the image P before the image P is input to the pre-processor 10, the image P can be processed by the pre-processor 10 and the post-processor 30 regardless of the image quality, image size, image format, etc. of the image P.

[0082] Here, according to the conventional technology, when there is a change in the image quality, image size, image format, etc. of the input image compared to the situation during learning, the inference accuracy may decrease. However, according to this embodiment, the image processing system 1 performs software processing on the image P before the image P is input to the preprocessor 10, so there is no need to re-learn in response to the changed image quality, image size, image format, etc. of the input image. Therefore, according to this embodiment, even if there is a change in the image quality, image size, image format, etc. of the input image, it is possible to prevent the inference accuracy from decreasing.

[0083] Furthermore, the image processing system 1 may perform software processing on the image P in response to a change in the type of image represented in the image P (for example, a change caused by a change in the object being imaged, a change in the imaging environment, a change in the imaging situation, etc.). In this case, the image processing system 1 may obtain information on a change in the object being imaged, a change in the imaging environment, a change in the imaging situation, etc. from a sensor (not shown) or the like, and perform software processing on the image P in response to the obtained situation. The image processing system 1 can perform inference with even greater accuracy by performing suitable software processing on the image P before the image P is input to the pre-processor 10.

[0084] In this embodiment, the number of classes or the type of classes to be calculated differs depending on whether detection accuracy or processing speed is to be prioritized, but a calculation unit that consumes less power instead of detection accuracy or processing speed may be included as a switching target. In other words, it is preferable to appropriately switch between processes that have a trade-off relationship in order to properly execute the required process. Note that the whole or part of the functions of each unit of the image processing system 1 in the above-mentioned embodiment may be realized by recording a program for realizing these functions on a computer-readable recording medium, reading the program recorded on the recording medium into a computer system, and executing it. Note that the term "computer system" here includes hardware such as an OS and peripheral devices.

[0085] In addition, "computer-readable recording medium" refers to portable media such as magneto-optical disks, ROMs, and CD-ROMs, and storage units such as hard disks built into computer systems. Furthermore, "computer-readable recording medium" may also include those that dynamically hold a program for a short period of time, such as a communication line when transmitting a program over a network such as the Internet, and those that hold a program for a certain period of time, such as volatile memory inside a computer system that serves as a server or client in such a case. Furthermore, the above program may be one that realizes part of the above-mentioned functions, or may be one that can realize the above-mentioned functions in combination with a program already recorded in the computer system.

[0086] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the spirit of the present invention. [Explanation of symbols]

[0087] 1...image processing system, 10...pre-processing, 30...post-processing, 310...correspondence information acquisition unit, 320...setting information acquisition unit, 330...extraction unit, 340...output unit, 331...compression unit, 332...switching unit, 333...first calculation unit, 334...second calculation unit, 335...calculation result output unit, 2...imaging system, 50...imaging device, 51...camera, 52...sensor, 53...driving device, P...image, RI1...first correspondence information group, RI2...second correspondence information group, O...object detection result, ID...input device, SF...setting file, SI...setting information

Claims

1. An image processing device that detects a type of object included in an image and a position coordinate where the object exists by image processing, a correspondence information acquisition unit that acquires a first correspondence information group including a plurality of pieces of correspondence information in which position coordinates indicating a range in which an object is expected to exist in the image are associated with a likelihood of a class corresponding to the range among a plurality of predetermined classes; A setting information acquisition unit that acquires setting information related to the image processing; an extraction unit that extracts a second correspondence information group including at least one likely class and one or more pieces of location information corresponding to the likely class, based on the acquired first correspondence information group and the acquired setting information; an output unit that outputs the extracted second correspondence information group; The setting information includes at least information on whether the setting is a first setting that prioritizes accuracy of the class and position coordinates extracted by the extraction unit, or a second setting that prioritizes processing speed of the extraction unit. Image processing device.

2. In the process of extracting the second correspondence information group by the extraction unit, the number of the classes to be calculated when the setting information is the second setting is smaller than the number of the classes to be calculated when the setting information is the first setting. The image processing device according to claim 1 .

3. The extraction unit further includes a switching unit that switches between a first calculation unit that performs a calculation to extract the second correspondence information group when the setting information is the first setting and a second calculation unit that performs a calculation to extract the second correspondence information group when the setting information is the second setting, based on the setting information.

3. The image processing device according to claim 1 or 2.

4. the extraction unit further includes a compression unit that compresses the first correspondence information group into a specific class from among the classes included in the first correspondence information group by a predetermined method; The first calculation unit or the second calculation unit performs a calculation for extracting the second correspondence information group based on the compressed correspondence information. The image processing device according to claim 3 .

5. The compression unit compresses the correspondence information included in the first correspondence information group when a number of classes in which the likelihood of the plurality of pieces of correspondence information included in the first correspondence information group is equal to or greater than a predetermined value is equal to or less than a predetermined value. The image processing device according to claim 4.

6. The switching unit switches based on the setting information when the image processing device is started. The image processing device according to claim 3 .

7. The setting information acquisition unit acquires the setting information from a setting file.

3. The image processing device according to claim 1 or 2.

8. The setting information acquisition unit acquires the setting information based on the first correspondence information group acquired by the correspondence information acquisition unit.

3. The image processing device according to claim 1 or 2.

9. a pre-processing device that calculates a first correspondence information group including a plurality of pieces of correspondence information in which position coordinates indicating a range in which an object is expected to exist in the image are associated with a likelihood of a class corresponding to the range among predetermined classes; 3. The image processing device according to claim 1, wherein the first correspondence information group is acquired from the preprocessing device. An image processing system comprising:

10. 1. An image processing method for detecting a type of object included in an image and a position coordinate where the object exists by image processing, comprising: a correspondence information acquisition step of acquiring a first correspondence information group including a plurality of pieces of correspondence information in which position coordinates indicating a range in which an object is expected to exist in the image are associated with a likelihood of a class corresponding to the range among a plurality of predetermined classes; A setting information acquisition step of acquiring setting information related to the image processing; an extraction step of extracting a second correspondence information group including at least one likely class and location information corresponding to the likely class based on the acquired first correspondence information group and the acquired setting information; and an output step of outputting the extracted second correspondence information group, The setting information includes at least information on whether the setting is a first setting that prioritizes accuracy of the class and position coordinates extracted by the extraction step, or a second setting that prioritizes processing speed of the extraction step. Image processing methods.

11. On the computer, A program for detecting a type of object included in an image and a position coordinate of the object by image processing, a correspondence information acquisition step of acquiring a first correspondence information group including a plurality of pieces of correspondence information in which position coordinates indicating a range in which an object is expected to exist in the image are associated with a likelihood of a class corresponding to the range among a plurality of predetermined classes; A setting information acquisition step of acquiring setting information related to the image processing; an extraction step of extracting a second correspondence information group including at least one likely class and location information corresponding to the likely class based on the acquired first correspondence information group and the acquired setting information; and an output step of outputting the extracted second correspondence information group, The setting information includes at least information on whether the setting is a first setting that prioritizes accuracy of the class and position coordinates extracted by the extraction step, or a second setting that prioritizes processing speed of the extraction step. program.

Citation Information

Patent Citations

  • Image processing apparatus, image processing method, and program

    JP2015082245A

  • Density measurement device, density measurement method and program

    JP2016095640A

  • Object detection method, object detection device, and image processing apparatus

    JP2020205039A

  • Object detection device, object detection method, program, and recording medium

    WO2020235269A1