Image processor, method, and program

The image processing device enhances object detection throughput by predicting object positions and determining targeted detection areas, addressing the limitations of existing systems in terms of efficiency and cost.

JP2025079452APending Publication Date: 2025-05-22NEC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023192121
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing object detection processing systems face challenges in achieving high throughput, leading to increased costs, power consumption, and the need for multiple devices to prevent oversights.

Method used

An image processing device that includes an object position prediction unit to predict object positions in new input images based on past detections, an object detection target area determination unit to determine areas for object detection, and an object detection unit to perform detection within these areas.

Benefits of technology

The proposed solution improves the throughput of object detection processing, reducing calculation load and delay, thereby enabling faster and more efficient object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025079452000001_ABST
    Figure 2025079452000001_ABST
Patent Text Reader

Abstract

To provide an image processor, a method, and a program which increase the through-put of object detection processing.SOLUTION: The method includes: object position prediction means for predicting the position of an object in a newly input image on the basis of the position of an object detected in a previous input image; object detection target region determination unit for determining an object detection target region as a target of an object detection in a newly input image on the basis of the result of prediction by the object position prediction means; and object detection means for detecting an object on the object detection target region determined by the object detection target region determination means.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to techniques for processing images. [Background technology]

[0002] One of the main tasks using machine learning is the object detection task in an image. The object detection task is a task for generating a list of pairs of positions and classes (types) of target objects present in an image. In recent years, object detection tasks using deep learning, particularly machine learning, have been widely used. For example, Patent Document 1 discloses a technology related to object position detection processing using a neural network.

[0003] In the learning phase of machine learning, the object detection task is given a group of learning images and information on the target object in each image as correct answer data. The information on the target object is selected according to the specifications of the object detection task. For example, the information on the target object includes the coordinates of the four vertices of a rectangular area in which the target object is reflected (bounding box (BB)) and the class of the target object. Note that in the following description, the BB and class are used as an example of the information on the target object. The object detection task uses the group of learning images and the information on the target object to generate a learned model as a result of machine learning using, for example, deep learning.

[0004] In the detection phase, the object detection task applies the trained model to an image including a target object to infer the target object included in the image. The object detection task also outputs a pair of a BB and a class for each target object included in the image. The object detection task may also output an evaluation result (e.g., confidence or score) for the result of object detection together with the BB and class.

[0005] For example, a surveillance system for monitoring people and vehicles can be constructed by using images output from a surveillance camera as input to an object detection task, which detects the positions and classes of people and vehicles appearing in the images. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] JP 2019-036008 A Summary of the Invention [Problem to be solved by the invention]

[0007] Object detection processing requires high throughput. For example, a system that performs object detection processing on images output from a surveillance camera needs to operate at a relatively high frame rate to prevent oversights (i.e., missed detections). If the throughput is low, problems arise such as an increase in the number of devices used for processing, increased costs, and increased power consumption.

[0008] The technology described in Patent Document 1 aims to improve the detection accuracy of the object position by correcting information acquired during the estimation process based on object movement information acquired from a series of images. Therefore, the technology described in Patent Document 1 cannot be expected to have the effect of improving throughput.

[0009] The present disclosure has been made in consideration of these problems, and has an object to provide an image processing device, a method, and a program that improve the throughput of object detection processing. [Means for solving the problem]

[0010] The image processing device according to the present disclosure is characterized by including an object position prediction means for predicting the position of an object in a new input image based on the position of the object detected in a past input image, an object detection target area determination means for determining an object detection target area in the new input image to be subject to object detection based on a prediction result by the object position prediction means, and an object detection means for performing object detection within the object detection target area determined by the object detection target area determination means.

[0011] The image processing method according to the present disclosure is characterized in that it predicts the position of an object in a new input image based on the position of the object detected in a past input image, determines an object detection target area in the new input image to be subject to object detection based on the prediction result, and performs object detection in the determined object detection target area.

[0012] The image processing program according to the present disclosure is characterized in that it causes a computer to execute an object position prediction process that predicts the position of an object in a new input image based on the position of the object detected in a past input image, an object detection target area determination process that determines an object detection target area in the new input image to be subject to object detection based on a prediction result of the object position prediction process, and an object detection process that performs object detection in the object detection target area determined by the object detection target area determination process. Effect of the Invention

[0013] According to the present disclosure, the throughput of object detection processing can be improved. [Brief description of the drawings]

[0014] [Figure 1] 1 is a block diagram showing an example of a configuration of an image processing device according to the present disclosure. [Diagram 2] FIG. 11 is a flow diagram showing an example of the operation of an object detection process in the image processing device according to the present disclosure. [Diagram 3] FIG. 11 is a flow diagram showing an example of the operation of an object detection process in the image processing device according to the present disclosure. [Figure 4] FIG. 11 is a flow diagram showing an example of the operation of an object detection process in the image processing device according to the present disclosure. [Diagram 5] FIG. 11 is a flow diagram showing an example of the operation of an object detection process in the image processing device according to the present disclosure. [Figure 6] FIG. 2 is a block diagram showing an example of a hardware configuration. [Figure 7] 1 is a block diagram showing an example of a configuration of an image processing device according to the present disclosure. [Figure 8] FIG. 11 is a flow diagram showing an example of the operation of an object detection process in the image processing device according to the present disclosure. [Figure 9] 1 is a block diagram showing an overview of an image processing device according to the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] The following documents 1 to 3 disclose techniques related to object detection tasks using deep learning. However, the following documents 1 to 3 do not disclose techniques related to improving the throughput of object detection processing. In other words, the techniques disclosed in the following documents 1 to 3 alone cannot improve the throughput of the object detection task. Reference 1: Shaoqing Ren, Kaiming He, Ross Girshick, Jian Sun, "Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks", [online], January 6, 2016, Cornel University, [Retrieved September 13, 2023], Internet<URL:https: / / arxiv.org / abs / 1506.01497> Reference 2: Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, Alexander C. Berg, "SSD: Single Shot MultiBox Detector", [online], December 29, 2016, Cornel University,[Retrieved September 13, 2023], Internet<URL:https: / / arxiv.org / abs / 1512.02325> Reference 3: Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, Piotr Dollar, "Focal Loss for Dense Object Detection", [online], International Conference on Computer Vision (ICCV), 2017, pp. 2980-2988, [Retrieved September 13, 2023], Internet<URL: https: / / openaccess.thecvf.com / content_ICCV_2017 / papers / Lin_Focal_Loss_for_ICCV_2017_paper.pdf>

[0016] Hereinafter, the embodiments of the present disclosure will be described with reference to the drawings. Each drawing is for explaining the embodiments. However, each embodiment is not limited to the description in each drawing. The same reference numerals are attached to similar configurations in each drawing. In addition, repeated explanations of similar configurations in each drawing may be omitted. In the drawings used in the following explanation, the description of parts not related to the explanation of the embodiment may be omitted and may not be illustrated.

[0017] [Configuration example of the first embodiment] The first embodiment will be described below with reference to the drawings.

[0018] In the first embodiment, the image processing device predicts the positions of an object group in a next input image (i.e., a new input image) based on the positions of the object group detected from a past input image. The image processing device also subjects the predicted area and its vicinity to object detection processing. This reduces the amount of data subject to object detection processing, and thus reduces the calculation load. This is expected to improve the throughput of the object detection processing. It is also expected to reduce the delay in the object detection processing.

[0019] [Configuration Description] The configuration of the first embodiment will be described with reference to the drawings.

[0020] Fig. 1 is a block diagram showing an example of the configuration of an image processing device 1 according to the present disclosure. The image processing device 1 includes an object position prediction unit 10, an object detection target region determination unit 20, a first object detection unit 30, a second object detection unit 40, an object position storage unit 50, and an image input unit 60. Note that the number and connection relationships of the components shown in Fig. 1 are merely an example. For example, the image processing device 1 may include a plurality of image input units 60.

[0021] The image processing device 1 may be configured using a computer device including a CPU (Central Processing Unit), a main memory, and a secondary storage device. In this case, the object position prediction unit 10, the object detection target area determination unit 20, the first object detection unit 30, the second object detection unit 40, the object position storage unit 50, and the image input unit 60 of the image processing device 1 shown in Fig. 1 are realized by the CPU executing processing according to a program stored in the secondary storage device. The hardware configuration of the image processing device 1 will be further described later.

[0022] In this embodiment, an example will be described in which the image processing device 1 uses an image as an input and detects a person in the image. For example, the image processing device 1 executes an object detection task of a person class. Note that the object detection task that can be executed by the image processing device 1 of the present disclosure is not limited to the person class.

[0023] In this embodiment, the object position prediction unit 10 predicts the position of a previously detected object (hereinafter also referred to as a known object) in a next input image. The first object detection unit 30 detects the known object in the image based on the prediction result by the object position prediction unit 10.

[0024] The second object detection unit 40 performs object detection for the purpose of preventing overlooking of an object whose information indicating its past position is not stored, or whose information indicating its past position has not been sufficiently accumulated to be used as an input for prediction by the object position prediction unit 10. Hereinafter, such objects are also referred to as new objects. Note that the operations of the object position prediction unit 10, the first object detection unit 30, and the second object detection unit 40 are not limited to these.

[0025] The object position prediction unit 10 uses position information indicating the past position of an object stored in the object position storage unit 50 as an input for prediction, and predicts the position of the object in the next image input from the image input unit 60. Hereinafter, the image input from the image input unit 60 is also referred to as an input image. The object position prediction unit 10 may use position information of an object detected by the second object detection unit 40 as an input for prediction. An example in which the object position prediction unit 10 uses position information of an object detected by the second object detection unit 40 as an input for prediction will be described later.

[0026] The object detection target area determination unit 20 determines an area in the input image in which the first object detection unit 30 should execute the object detection process, using the prediction result by the object position prediction unit 10, past position information of known objects stored in the object position storage unit 50, and information on new objects. Hereinafter, the area in the input image that is the target of the object detection process by the first object detection unit 30 is also referred to as the object detection target area. The object detection target area is, for example, a collection of partial areas that are part of the entire area of ​​the input image. Hereinafter, each partial area that constitutes the object detection target area is also referred to as an individual object detection target area. Details of the individual object detection target area will be described later.

[0027] The first object detection unit 30 executes an object detection task using an image as input, and outputs an object detection result. For example, the first object detection unit 30 uses only an area including the position of the object predicted by the object position prediction unit 10 (or an individual object detection target area determined by the object detection target area determination unit 20) of the input image input from the image input unit 60 as an input. The first object detection unit 30 performs object detection processing using, for example, deep learning. The first object detection unit 30 holds a trained model (specifically, the trained model is stored in a predetermined storage medium and can be referenced), and performs inference by applying the trained model to the input. In addition, the first object detection unit 30 outputs, for each object, a BB, a class, a score, and the like as an object detection result. However, the operation of the first object detection unit 30 is not limited to this. For example, the first object detection unit 30 may hold a plurality of trained models. The first object detection unit 30 may switch parameters used for inference (e.g., a trained model to be used or an input size) according to characteristics of an individual object detection target region (e.g., the height or width of the region). For example, the first object detection unit 30 may perform inference using a smaller input size for an individual object detection target region having a small height or width.

[0028] The second object detection unit 40 executes an object detection task on the input image input from the image input unit 60. The second object detection unit 40 outputs an object detection result obtained by executing the object detection task. The second object detection unit 40 performs an object detection process using deep learning, for example, similar to the first object detection unit 30. The second object detection unit 40 may use a trained model different from that of the first object detection unit 30. For example, the second object detection unit 40 may use a trained model that is lighter than the trained model used by the first object detection unit 30. However, the operation of the second object detection unit 40 is not limited to this. The second object detection unit 40 may intermittently execute an object detection task on the input image input from the image input unit 60. For example, the second object detection unit 40 may execute an object detection task at a predetermined interval (for example, once per second).

[0029] The object position storage unit 50 stores information about objects (known objects) detected by the object detection target area determination unit 20 and the first object detection unit 30. The object position storage unit 50 stores information about objects (new objects) detected by the second object detection unit 40. The object position storage unit 50 stores, as information about an object, all or part of, for example, a bounding box, a class, a score, a detection time, an identifier of an input image in which the object was detected, and an identifier of an image input unit 60 that generated the input image.

[0030] The image input unit 60 generates an image to be processed by the image processing device 1 (i.e., an input image). The image input unit 60 may be realized by, for example, a surveillance camera. The image input unit 60 may use, as the input image, for example, an image captured by a camera as it is, or may use an image after preprocessing such as image processing or clipping. The image input unit 60 may also input, as the input image, each image that is continuously output at a predetermined frame rate from a device external to the image processing device 1.

[0031] The method of predicting an object position by the object position predicting unit 10 will be described.

[0032] Methods for predicting the future location of an object from its past locations have been widely studied. For example, there is a research field called human trajectory prediction, which is related to predicting the location of a person. In recent years, human trajectory prediction methods using deep learning have been widely studied. In human trajectory prediction methods using deep learning, past movement trajectories are used as input, and inference is performed using a trained model to predict future movement trajectories.

[0033] The object position prediction unit 10 may perform the prediction process of the object position using any human trajectory prediction method. In many human trajectory prediction methods, in order to obtain high prediction accuracy, information summarizing the movement trajectories of the same person is used as an input. In addition, in many human trajectory prediction methods, in order to obtain information summarizing the movement trajectories of the same person, a tracking process is performed to determine the same person from multiple images captured at different times in the past. The execution of the tracking process involves a calculation load. Therefore, when a human trajectory prediction method requiring a tracking process is applied to the image processing device 1, there is a risk of causing a decrease in processing throughput.

[0034] The object position prediction unit 10 may use a prediction method that does not require tracking processing instead of a prediction method that requires tracking processing. An example of a prediction method that does not require tracking processing will be described below. The object position prediction unit 10 of this embodiment uses a prediction method that does not require tracking processing, which will be described below. This prediction method is also referred to as a first prediction method of this embodiment. However, the prediction method that the object position prediction unit 10 can use is not limited to the first prediction method of this embodiment.

[0035] When using position information of an object at multiple different times as input to prediction, a tracking process may be performed to use information that summarizes the movement trajectories of the same person. In this case, the input to prediction is not independent between the multiple times, and it can be said that there is a dependency. On the other hand, there is a case where a tracking process is not performed and only the position information of an object at each time is used as input to prediction. In this case, it can be said that the input to prediction is independent from each other between multiple times. The first prediction method of this embodiment is a prediction method that does not perform a tracking process. Therefore, the first prediction method of this embodiment uses information that is independent from each other between multiple different times as input to prediction.

[0036] The first prediction method of this embodiment uses position information of an object in an input image at a plurality of past times as an input. In this example, the number of past (i.e., observed) times is Nobs. The first prediction method of this embodiment predicts the position of an object in an input image at the next time. However, the prediction target of the first prediction method of this embodiment is not limited to this. For example, the first prediction method of this embodiment may predict the position of an object in an input image at a plurality of future times. Also, the number of objects included in the input image is an arbitrary number.

[0037] The first prediction method of this embodiment divides the input image and the output image (i.e., the prediction result) into a predetermined size, and manages the position of the object in units of the divided regions (grids). For example, when the original image is full HD (High Definition) [1920×1080] and the predetermined size is 32×32, the first prediction method of this embodiment manages the position of the object using 60×34 grids.

[0038] The first prediction method of the present embodiment uses position information of objects in Nobs past input images to create inference input data to be input to the trained model. The inference input data is, for example, a floating-point vector whose dimension is the number of grids. In this case, the inference input data includes, in an element corresponding to a grid, 1 when one or more objects are present in the grid, and 0 when no object is present at all. The output from the trained model (i.e., the inference output data) is, for example, a floating-point vector whose dimension is the number of grids. In this case, the output from the trained model (i.e., the inference output data) includes, in an element corresponding to a grid, a numerical value that is high when it is predicted that the probability of an object being present in the grid is high.

[0039] The first prediction method of this embodiment uses, for example, a predetermined threshold, and predicts that an object exists in a grid corresponding to each element of the inference output data if the value is equal to or greater than the threshold. The first prediction method of this embodiment predicts that an object does not exist in a grid corresponding to the element if the value is less than the threshold. The trained model used by the first prediction method of this embodiment is generated by training a model for prediction (i.e., a prediction model) using the inference input data and inference output data generated from a ground truth dataset including an image and position information related to a moving object.

[0040] The first prediction method of the present embodiment may use, for example, a deep learning model using a recurrent neural network (RNN) or a long short term memory (LSTM) network. The first prediction method of the present embodiment may also use an input in which past time series data is combined in the channel direction, and may use a deep learning model using a convolutional neural network (CNN) or a transformer.

[0041] The first prediction method of the present embodiment described above can predict the future positions of objects without identifying, separating, or tracking each object, even if multiple objects are captured in an image at a certain time. In other words, each of the Nobs floating-point vectors is independent of each other.

[0042] A method for determining an object detection target region by object-detection target region determining section 20 will be described.

[0043] The object detection target area determination unit 20 extracts grids in which an object exists from the prediction result by the object position prediction unit 10 (i.e., information indicating whether an object is located in each grid). Furthermore, the object detection target area determination unit 20 groups adjacent grids from among the extracted grids to create a set of grids. However, the operation of the object detection target area determination unit 20 is not limited to this. The number of grid sets is 0, 1, or multiple depending on the predicted position of the object (group).

[0044] Object detection target region determination unit 20 may determine an area corresponding to a set of the created grids as an individual object detection target region, and a set of the individual object detection target regions as the object detection target region. Alternatively, object detection target region determination unit 20 may determine, for each set of the created grids, a rectangle of a minimum size that encompasses the area corresponding to that set, and determine the determined rectangle as the individual object detection target region.

[0045] The object detection target area determination unit 20 may adjust the individual object detection target area. For example, the object detection target area determination unit 20 may multiply the height or width of the individual object detection target area by a constant or add a constant value. By enlarging the individual object detection target area, it is possible to avoid oversights or defects when the prediction is incorrect.

[0046] For example, the object detection target area determination unit 20 checks whether or not the areas overlap for each pair of two individual object detection target areas. If there is an overlap, the object detection target area determination unit 20 may integrate the two individual object detection target areas (i.e., eliminate overlap) to form one individual object detection target area. The object detection target area determination unit 20 may, for example, determine the smallest rectangle that includes the two individual object detection target areas as the individual object detection target area after integration. By integrating the individual object detection target areas (eliminating overlap), it is possible to avoid double detection in object detection for each individual object detection target area. The object detection target area determination unit 20 may adjust the individual object detection target areas described above after eliminating overlap.

[0047] When the object detection result by the second object detection unit 40 is available, the object detection target region determination unit 20 may update the existing object detection target region by using position information of the object detected by the second object detection unit 40. Furthermore, the object detection target region determination unit 20 may determine a new object detection target region by using the position information of the object detected by the second object detection unit 40.

[0048] For example, the object detection target area determination unit 20 checks whether or not the bounding box of each detected object included in the object detection result by the second object detection unit 40 is included in the object detection target area. If the bounding box of the detected object is not included in the object detection target area, the object detection target area determination unit 20 may add the bounding box of the detected object as an individual object detection target area and update the object detection target area. At that time, the object detection target area determination unit 20 may adjust or integrate the individual object detection target area described above. The object detection target area determination unit 20 may treat an object that is not included in the object detection target area among the detected objects included in the object detection result by the second object detection unit 40 as a new object. In this case, the object detection target area determination unit 20 may store information about the new object in the object position storage unit 50. The information about the new object may include, for example, position information such as a bounding box, as well as information that can identify the object as a new object.

[0049] When information about a new object is stored in object position storage unit 50, object detection target region determination unit 20 may update an existing object detection target region using the information about the new object. Also, object detection target region determination unit 20 may determine a new object detection target region using the information about the new object.

[0050] The object detection target area determination unit 20 may update the object detection target area for each new object, for example, by adding an area corresponding to the bounding box of the new object as an individual object detection target area. However, there is a possibility that the new object may be moving. Therefore, the object detection target area determination unit 20 may adjust the bounding box of the new object before setting it as the individual object detection target area. The object detection target area determination unit 20 may, for example, expand the bounding box up, down, left, and right by a predetermined value (i.e., add or subtract from the coordinate values). The object detection target area determination unit 20 may adjust or integrate the individual object detection target areas described above.

[0051] The object detection target region determination unit 20 may update the information on the new object stored in the object position storage unit 50, using the object detection result for the input image input from the image input unit 60. The object detection result is expected to include the detection result of the new object at its latest position.

[0052] For example, the object detection target area determination unit 20 searches for a detection result corresponding to the new object from among the object detection results for each new object. Then, the object detection target area determination unit 20 may update information about the new object stored in the object position storage unit 50 by using the corresponding detection result. The object detection target area determination unit 20 may use any method as a method for searching for a detection result corresponding to the new object from among the object detection results. For example, the object detection target area determination unit 20 may calculate an IoU (Intersection over Union) between the new object and each object included in the object detection result, and determine that the object with the maximum IoU corresponds to the new object. When the object detection target area determination unit 20 is unable to search for a detection result corresponding to the new object from among the object detection results (for example, when the maximum IoU is less than a predetermined threshold), it may delete information about the new object. Furthermore, the object detection target area determination unit 20 may execute a tracking process for the new object and search for a detection result corresponding to the new object from among the object detection results.

[0053] The object detection target area determination unit 20 may delete information about a new object stored in the object position storage unit 50 at any timing. For example, the object detection target area determination unit 20 may delete information about a new object when the information about the new object has been accumulated sufficiently to be used as an input for prediction by the object position prediction unit 10. The object detection target area determination unit 20 may determine that the information about the new object has been accumulated sufficiently to be used as an input for prediction by the object position prediction unit 10 when the object detection process for the new object by the first object detection unit 30 has been executed Nobs times or more.

[0054] [Explanation of operation] Next, an example of the operation of the image processing device 1 according to the first embodiment will be described with reference to the drawings.

[0055] (A) Operation of object detection process based on prediction Fig. 2 is a flow diagram showing an example of the operation of object detection processing in the image processing device 1 according to the present disclosure. The operation shown in Fig. 2 includes an operation in which the object detection target area determination unit 20 determines the object detection target area using object position information predicted based on position information of objects (known objects) detected in the past and stored in the object position storage unit 50. The operation shown in Fig. 2 also includes an operation in which the first object detection unit 30 performs object detection processing for each individual object detection target area constituting the object detection target area.

[0056] The image processing device 1 starts this operation, for example, every time an input image is input from the image input unit 60. However, when object detection is performed by the second object detection unit 40 (for example, when object detection is performed at predetermined time intervals), the image processing device 1 executes an object detection processing operation (B) that also uses the second object detection unit 40, which will be described later, instead of this operation.

[0057] The object position prediction unit 10 acquires position information of objects (known objects) included in the most recent Nobs input images from the object position storage unit 50. Based on the acquired position information, the object position prediction unit 10 creates inference input data to be input to the predictive learned model of the first prediction method of this embodiment (step S100).

[0058] Next, the object position prediction unit 10 performs inference using the learned model by using the inference input data created in step S100 as an input, and predicts the object position (step S101).

[0059] Next, object-detection target region determining section 20 determines an object-detection target region based on the prediction result obtained in step S101 (that is, information indicating whether or not an object is located in each grid unit) (step S102).

[0060] Next, for each individual object detection target region in the object detection target region obtained in step S102 (step S103), the first object detection unit 30 performs object detection processing on the region in the input image input from the image input unit 60 (step S104). For example, the first object detection unit 30 executes object detection processing using only the image corresponding to the region as input. The process of step S104 is repeated until it is executed for all individual object detection target regions. The first object detection unit 30 converts the bounding box information, which is the position information of the object obtained as a result of the object detection processing, from the coordinate values in the individual object detection target region to the coordinate values in the input image input from the image input unit 60.

[0061] Next, the first object detection unit 30 stores the position information of the object obtained in the processes of steps S103 to S104 in the object position storage unit 50 (step S105). The first object detection unit 30 may also store, in the object position storage unit 50 at the same time, information for identifying the input image, such as the time when the input image was generated, the identifier of the input image, and the identifier of the image input unit 60, together with the position information of the object. Further, the first object detection unit 30 may store the input image itself in the object position storage unit 50 at the same time.

[0062] (B) Operation of Object Detection Processing Using the Second Object Detection Unit 40 FIGS. 3 and 4 are flowcharts showing an example of the operation of object detection processing in the image processing apparatus 1 according to the present disclosure. FIGS. 3 and 4 show an example of the operation of object detection processing using the second object detection unit 40 in combination. This operation includes an operation in which the object detection target region determination unit 20 determines the object detection target region using the object detection result by the second object detection unit 40 in addition to the position information predicted based on the position information of the objects (known objects) detected in the past stored in the object position storage unit 50. Further, this operation includes an operation in which the first object detection unit 30 performs object detection processing for each individual object detection target region constituting the object detection target region.

[0063] For example, when an input image is input from the image input unit 60 and object detection is performed by the second object detection unit 40, the image processing device 1 executes this operation instead of the operation (A) of the object detection processing based on prediction. Note that in Figs. 3 and 4, explanations of steps similar to the operation (A) of the object detection processing based on prediction will be omitted.

[0064] The image processing device 1 performs the same processes as steps S100 to S102 shown in FIG.

[0065] Next, the second object detection unit 40 performs an object detection process using the input image input from the image input unit 60 as an input (step S110).

[0066] Next, the object-detection target region determining unit 20 updates the object-detection target region determined in step S102 based on the object detection result obtained in step S110 (step S111). The details of step S111 will be described later.

[0067] Next, for each individual object detection target region of the object detection target region obtained in step S111 (step S103), first object detection unit 30 performs object detection processing on that region in the input image input from image input unit 60 (step S104). For example, first object detection unit 30 performs object detection processing using only the image corresponding to that region as input. The processing of step S104 is repeated until it is performed on all individual object detection target regions.

[0068] Next, the first object detection unit 30 stores the object position information obtained in steps S103 and S104 in the object position storage unit 50 (step S105).

[0069] Next, the object detection target area determination unit 20 stores information (e.g., bounding box information) about the new object that was processed in step S111C (details of which will be described later) included in step S111, in the object position storage unit 50 (step S112).

[0070] Next, step S111 shown in FIG. 3 will be described in detail with reference to FIG.

[0071] Object-detection target region determining section 20 performs the processes of steps S111B to S111C described below for each detected object obtained in step S110 (step S111A).

[0072] Object detection target region determination unit 20 checks whether the bounding box of the detected object is completely contained in the object detection target region determined in step S102 (step S111B). If it is completely contained (Yes in step S111B), object detection target region determination unit 20 ends the process for the detected object.

[0073] If the detected object is not completely included (No in step S111B), object-detection target region determination unit 20 identifies the detected object as a new object. Furthermore, object-detection target region determination unit 20 updates the object-detection target region so that the bounding box of the detected object identified as a new object is completely included (step S111C).

[0074] When there are overlapping individual object detection target regions among the object detection target regions updated in steps S111A to S111C, object-detection target region determining section 20 merges them (step S111D).

[0075] (C) Object detection process using object position information of new objects Fig. 5 is a flow diagram showing an example of the operation of the object detection processing in the image processing device 1 according to the present disclosure. Fig. 5 shows an example of the operation of the object detection processing using object position information of a new object in combination. This operation includes an operation in which the object detection target area determination unit 20 determines the object detection target area using the detection result of a new object obtained by the operation (B) of the object detection processing using the second object detection unit 40 in combination, in addition to position information predicted based on position information of previously detected objects (known objects) stored in the object position storage unit 50. This operation also includes an operation in which the first object detection unit 30 performs object detection processing for each individual object detection target area constituting the object detection target area.

[0076] This operation is executed in place of the operation (A) of the object detection processing based on prediction until the next input image of the input image subjected to the operation (B) of the object detection processing using the second object detection unit 40 and Nobs-1 input images thereafter are processed. After a total of Nobs input images including the input image subjected to the operation (B) of the object detection processing using the second object detection unit 40 are processed, the position information of the new object is sufficiently accumulated to be used as an input to the prediction by the first prediction method of this embodiment. Therefore, the new object is treated as a known object. In addition, the information on the new object is deleted from the object position storage unit 50.

[0077] In FIG. 5, the description of steps similar to those in the operation (A) of the object detection processing based on prediction will be omitted.

[0078] The image processing device 1 performs the same processes as steps S100 to S102 shown in FIG.

[0079] Next, if information about a new object is stored in object position storage unit 50, object-detection target region determining unit 20 updates the object-detection target region determined in step S102 using the information about the new object (step S120).

[0080] Next, for each individual object detection target region of the object detection target region obtained in step S120 (step S103), first object detection unit 30 performs object detection processing on that region in the input image input from image input unit 60 (step S104). For example, first object detection unit 30 performs object detection processing using only the image corresponding to that region as input. The processing of step S104 is repeated until it is performed on all individual object detection target regions.

[0081] Next, the first object detection unit 30 stores the object position information obtained in steps S103 and S104 in the object position storage unit 50 (step S105).

[0082] Next, the object detection target region determination unit 20 updates the information on the new object stored in the object position storage unit 50 using the object position information (object detection result) obtained in steps S103 to S104 (step S121). For example, the object detection target region determination unit 20 searches for the new object stored in the object position storage unit 50 and the object with the maximum IoU from the object detection result. After that, the object detection target region determination unit 20 updates the information on the new object using the position information of the searched object. In addition, the object detection target region determination unit 20 may delete the information on the new object from the object position storage unit 50 when the next input image after the input image subjected to the operation (B) of the object detection process using the second object detection unit 40 and Nobs-1 input images thereafter are processed.

[0083] Based on the above operations, the image processing device 1 performs object detection processing on the input image input from the image input unit 60.

[0084] [Effect description] Next, the effects of the first embodiment will be described.

[0085] The image processing device 1 according to the first embodiment can improve the throughput of the object detection process. Also, the image processing device 1 according to the first embodiment can reduce the delay in the object detection process. The reason is as follows.

[0086] The object position prediction unit 10 of the image processing device 1 according to the first embodiment predicts the position of an object included in an input image. The object detection target area determination unit 20 determines an individual object detection target area, which is a partial area of ​​the input image, based on the predicted object position. The first object detection unit 30 performs object detection processing for each individual object detection target area. The size of the individual object detection target area is expected to be smaller than the input image. That is, in the image processing device 1 according to the first embodiment, the amount of data to be subjected to the object detection processing is reduced, and the calculation load is reduced accordingly. Therefore, the throughput of the object detection processing is expected to be improved. In addition, the delay of the object detection processing is expected to be shortened. That is, in the image processing device 1 according to the first embodiment, the period from when the target object is captured in the image to when it is detected is shortened (that is, the delay is shortened), making it possible to quickly respond to the target object.

[0087] [Variations] In the above description, an example has been shown in which the time interval applied to the input and output to the prediction by the object position prediction unit 10 is the same as the generation interval of the input image by the image input unit 60. However, the present disclosure is not limited to this. Each unit constituting the image processing device 1, including the object position prediction unit 10 and the image input unit 60, may be configured to operate at different time periods. For example, each unit constituting the image processing device 1 may operate using the latest information available at the timing when each unit operates.

[0088] The first object detection unit 30 and the second object detection unit 40 may output a score (or confidence or accuracy) for each detected object as an object detection result. When using the object detection result, each unit constituting the image processing device 1 may filter the detection result using a predetermined threshold and score. The predetermined threshold may be different for each application.

[0089] In the above description, an example was shown in which the trained model used in the first prediction method of the present embodiment is used for a correct answer data set including an image and position information related to a moving object, the inference input data and the inference output data are generated from the correct answer data set, and a model for prediction (prediction model) is trained and generated. However, the present disclosure is not limited to this. When training the model, conversion, processing, and augmentation may be performed on the correct answer position information included in the correct answer data set. For example, similar to the adjustment of the object detection target area, the object position may be enlarged up, down, left, right, and right. In such a case, the prediction result may be larger than the actual object position, but it is expected to have the effect of reducing oversight of object detection. In addition, within the Nobs frame used for one prediction, each object may be translated, flipped left and right or up, rotated, or scaled. In addition, the object may be scaled in the time direction. That is, the moving speed of the object may be slowed down or accelerated. For example, the Nobs frames may be extracted by thinning out one frame at a time from 2·Nobs frames of the correct answer data, and these may be used as the input for prediction. In this case, the objects in that frame will move twice as fast.

[0090] In the above description, an example has been shown in which the object position prediction unit 10 applies the first prediction method of the present embodiment to all known objects. However, the present disclosure is not limited to this. The object position prediction unit 10 may apply a different prediction method to some objects, different from the first prediction method of the present embodiment. The object position prediction unit 10 may apply, for example, a prediction method involving a tracking process. In general, a prediction method involving a tracking process increases the processing load, but on the other hand, improves the prediction accuracy. Therefore, it is expected that the detection accuracy will be improved by using a prediction method involving a tracking process (for example, it is expected that oversight due to prediction failure will be reduced). The object position prediction unit 10 may apply a different position prediction method to an object whose score specified as an object detection result is lower than a predetermined threshold value, for example. In addition, the object position prediction unit 10 may divide the input image into a plurality of regions and apply a different prediction method to each region, or may switch parameters related to prediction for each region. Configuration information of the regions and a switching pattern may be given in advance.

[0091] In the above description, an example has been shown in which both the second object detection unit 40 and the first object detection unit 30 perform object detection processing in the operation (B) of the object detection processing in which the second object detection unit 40 is also used. However, the present disclosure is not limited to this. In the operation (B) of the object detection processing in which the second object detection unit 40 is also used, the processing of step S111 and the object detection processing by the first object detection unit 30 (steps S103 and S104) may be omitted. In that case, the object detection result by the second object detection unit 40 (step S110) may be treated as the object detection result by the first object detection unit 30. Also, in step S110, instead of the object detection processing by the second object detection unit 40, the object detection processing by the first object detection unit 30 may be performed on the entire input image.

[0092] In the above description, an example has been shown in which an input image used as an input to prediction and an input image to be subjected to object detection (or an input image to be subjected to output as a prediction result) are input from the same image input unit. However, the present disclosure is not limited to this. For example, the image processing device 1 may be configured to include a plurality of image input units 60 (for example, the image input unit 60A and the image input unit 60B). In this case, the object position prediction unit 10 may predict the future position of the object in the input image input from the image input unit 60B by using the position of the object in the input image input from the image input unit 60A as an input to prediction. The trained model used by the object position prediction unit 10 for inference may be trained assuming such a configuration. The difference in the shooting range (or the angle of view) between the image input unit 60A and the image input unit 60B may be, for example, fixed. Also, information indicating the difference in the shooting range (or the angle of view) between the image input unit 60A and the image input unit 60B may be used during learning.

[0093] In the above description, an example in which the image processing device 1 executes an object detection task has been described. However, the present disclosure is not limited to this. The image processing device 1 may execute other tasks, for example, pose estimation or region recognition (segmentation). For example, the first object detection unit 30 may execute other tasks in addition to or instead of the object detection task. Also, the second object detection unit 40 may execute other tasks in addition to or instead of the object detection task. If the task does not directly generate position information of an object, the image processing device 1 may generate position information of an object, or information required for a prediction operation by the object position prediction unit 10, which is an alternative, based on the output from the task. For example, if the task is a pose estimation task, information (type, position, etc.) on the joint points of a person in the input image is obtained as the output. The image processing device 1 may generate a person rectangle from the obtained group of joint points and use it as position information of the object.

[0094] Also, for example, the first object detection unit 30 may execute an image classification task instead of the object detection task. In general, the image classification task has a smaller processing load than the object detection task, and is expected to have the effect of improving throughput. Also, the first object detection unit 30 may switch whether or not to use the image classification task depending on the characteristics of the individual object detection target area. For example, the first object detection unit 30 may select the image classification task in the following cases. When the size of the individual object detection target area is less than a predetermined threshold When the number of objects present in the individual object detection target area is expected to be 1 or less - When integration (reduction of overlap) processing is not applied to the individual object detection target area

[0095] In the above description, an example has been shown in which the image processing device 1 executes the operation of (A) prediction-based object detection processing every time an input image is input from the image input unit 60. However, the present disclosure is not limited to this. For example, when it can be determined in advance that no known object exists and the object detection target area obtained as a result of the prediction is empty, the image processing device 1 may omit execution of the operation. This is expected to reduce the load related to the prediction.

[0096] In the above description, an example in which the image input unit 60 generates an input image has been shown. However, the present disclosure is not limited to this. The image input unit 60 may receive compressed image data from an external device and decode the image data to generate an input image. The image input unit 60 may perform decoding of a compression method such as JPEG (Joint Photographic Experts Group) or MPEG (Moving Picture Experts Group). In addition, the image input unit 60 may switch the generation method using a prediction result for a past input image.

[0097] For example, when it can be determined in advance that no known object exists and the object detection target area obtained as a result of the prediction is empty, the image input unit 60 may not perform the decoding process, or may perform the decoding process using a low-load, low-quality decoding method. The image input unit 60 may also perform the decoding process only on a portion of the object detection target area. For example, the image input unit 60 may perform the decoding process for each individual object detection target area, or may perform the decoding process only on a minimum rectangular area including all individual object detection target areas. The image input unit 60 may fill in the area where the decoding process has not been performed with a dummy image (e.g., a blackened image). The image input unit 60 may also output area information for the area where the decoding process has not been performed to a component that uses the input image, and cause the component to use the input image by referring to the area information. The image input unit 60 may also generate an input image as usual, or may generate an input image using a low-load, low-quality decoding method, when it is the timing to execute the operation of the object detection process using the second object detection unit 40 (B).

[0098] [Hardware configuration] In the above description, an example has been used in which the object position prediction unit 10, the object detection target region determination unit 20, the first object detection unit 30, the second object detection unit 40, the object position storage unit 50, and the image input unit 60 are included in the same device (image processing device 1). However, the first embodiment is not limited to this.

[0099] For example, the image processing device 1 may be configured by connecting devices having functions corresponding to each component via a predetermined network.

[0100] Each component of the image processing device 1 may be configured as a hardware circuit, or multiple components in the image processing device 1 may be configured as a single piece of hardware.

[0101] Alternatively, the image processing device 1 may be realized as a computer device including a CPU, a read only memory (ROM), and a random access memory (RAM). The image processing device 1 may be realized as a computer device including an input / output connection circuit (IOC: Input and Output Circuit) in addition to the above configuration. The image processing device 1 may be realized as a computer device including a network interface circuit (NIC: Network Interface Circuit) in addition to the above configuration.

[0102] Alternatively, the image processing device 1 may be realized as a computer device further including an arithmetic unit that performs calculations for part or all of the processing related to tracking, such as calculation of feature amounts and inference.

[0103] FIG. 6 is a block diagram showing a configuration of an information processing device 600 which is an example of the hardware configuration of the image processing device 1. As shown in FIG.

[0104] The information processing device 600 includes a CPU 610, an arithmetic unit 611, a ROM 620, a RAM 630, an internal storage device 640, an IOC 650, and a NIC 680. The information processing device 600 constitutes a computer device.

[0105] The CPU 610 loads a program from the ROM 620 and / or the internal storage device 640. Based on the loaded program, the CPU 610 controls the RAM 630, the internal storage device 640, the arithmetic unit 611, the IOC 650, and the NIC 680. The computer device including the CPU 610 controls these configurations and realizes the functions of the object position prediction unit 10, the object detection target area determination unit 20, the first object detection unit 30, the second object detection unit 40, and the object position storage unit 50.

[0106] When implementing each function, the CPU 610 may use the RAM 630 or the internal storage device 640 as a temporary storage medium for a program.

[0107] Furthermore, the CPU 610 may read, using a storage medium reading device (not shown), a program contained in a storage medium 690 that stores a computer-readable program. Alternatively, the CPU 610 may receive a program from an external device (not shown) via the NIC 680, store the program in the RAM 630 or the internal storage device 640, and operate based on the stored program.

[0108] The arithmetic unit 611 may be, for example, any one of a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), and an AI (Artificial Intelligence) chip. The arithmetic unit 611 may perform calculations for part or all of processes such as object detection and prediction inference under the control of a program executed by the CPU 610. Data, programs, circuit information, etc. required for the execution of calculations by the arithmetic unit 611 may be stored in, for example, the ROM 620, the RAM 630, the internal storage device 640, etc.

[0109] The ROM 620 stores fixed data and programs executed by the CPU 610. The ROM 620 is, for example, a P-ROM (Programmable ROM) or a flash ROM.

[0110] The RAM 630 temporarily stores programs and data executed by the CPU 610. The RAM 630 is, for example, a dynamic RAM (D-RAM).

[0111] The internal storage device 640 stores data and programs that are to be stored long-term by the information processing device 600. The internal storage device 640 may operate as the object position storage unit 50. The internal storage device 640 may also operate as a temporary storage device for the CPU 610. The internal storage device 640 is, for example, a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), or a disk array device.

[0112] The ROM 620 and the internal storage device 640 are non-transitory recording media. On the other hand, the RAM 630 is a transitory recording media. And the CPU 610 is operable based on a program stored in the ROM 620, the internal storage device 640, or the RAM 630. That is, the CPU 610 is operable using a non-transitory recording media or a transitory recording media.

[0113] The IOC 650 mediates data between the CPU 610, the input device 660, and the display device 670. The IOC 650 is, for example, an IO interface card or a USB (Universal Serial Bus) card. Further, the IOC 650 is not limited to wired such as USB, and may be connectable wirelessly.

[0114] The input device 660 is a device that receives an instruction from an operator of the information processing apparatus 600. For example, the input device 660 receives parameters. The input device 660 is, for example, a keyboard, a mouse, or a touch panel. Also, the input device 660 may be an input device that functions as the image input unit 60. The image input unit 60 may be, for example, a camera device.

[0115] The display device 670 is a device capable of displaying information to an operator of the information processing apparatus 600. The display device 670 is, for example, a liquid crystal display, an organic electro-luminescence display, or an electronic paper.

[0116] The NIC 680 relays the exchange of data with an external device (not shown) via a network. The NIC 680 is, for example, a LAN (Local Area Network) card. The NIC 680 is not limited to wired, and may be connectable to an external device wirelessly.

[0117] The information processing device 600 configured in this manner can obtain the same effects as the image processing device 1. This is because the CPU 610 of the information processing device 600 can realize functions similar to those of the image processing device 1 based on a program. Also, the CPU 610 and the arithmetic unit 611 of the information processing device 600 can realize functions similar to those of the image processing device 1 based on a program.

[0118] <Second embodiment> Next, a second embodiment of the present disclosure will be described. An image processing device 1B according to the second embodiment generates an aggregated image by collecting image regions of an object detection target region, and performs object detection on the aggregated image.

[0119] The second embodiment will be described with reference to the drawings. In the drawings to be referred to in the description of the second embodiment, the same reference numerals are used for the configurations that perform the same operations as those in the first embodiment. Detailed description of these configurations will be omitted.

[0120] [Configuration Description] The configuration of an image processing device 1B according to the second embodiment will be described with reference to the drawings. Note that the image processing device 1B may be configured using a computer device as shown in FIG. 6, similarly to the first embodiment.

[0121] FIG. 7 is a block diagram showing an example of the configuration of an image processing device 1B according to the present disclosure.

[0122] The image processing device 1B illustrated in FIG. 7 includes an object position prediction unit 10, an object detection target area determination unit 20, a first object detection unit 30B, a second object detection unit 40, an object position memory unit 50, an image input unit 60, and an aggregate image generation unit 70.

[0123] The aggregated image generation unit 70 generates an aggregated image by collecting the image regions of the object detection target regions. The aggregated image generation unit 70 receives information indicating the object detection target region from the object detection target region determination unit 20, and generates an image (aggregated image) by collecting the image regions of the individual object detection target regions. Collecting the obtained image region group (that is, copying it) is called Packing. The aggregated image generation unit 70 may generate one or more aggregated images. When generating the aggregated image, the aggregated image generation unit 70 stores the individual object detection target region and information indicating the arrangement of the individual object detection target region on the aggregated image in association with each other. When performing Packing, the aggregated image generation unit 70 may provide a gap (interval) with a predetermined width between the respective image regions.

[0124] When performing Packing, the aggregated image generation unit 70 may change (that is, enlarge or reduce) the size of the individual object detection target region and copy it to the aggregated image. When reducing the size, the number of aggregated images decreases, so there is a possibility of shortening the inference processing time for object detection. Also, when enlarging the size, there is a possibility of improving the recognition accuracy. The aggregated image generation unit 70 may determine whether to change the size of the individual object detection target region and determine the size after the change based on a given threshold value or the like in advance. For example, the aggregated image generation unit 70 may determine whether to change the size of the individual object detection target region and determine the size after the change based on the area of the individual object detection target region. When changing the size of the individual object detection target region, the aggregated image generation unit 70 may perform image processing such as complement processing. Also, in addition to or instead of changing the size of the image region, the aggregated image generation unit 70 may perform arbitrary image processing. As the image processing, the aggregated image generation unit 70 may perform, for example, brightness adjustment, luminance adjustment, color adjustment, contrast adjustment, geometric correction, and the like.

[0125] The first object detection unit 30B has the same function as the first object detection unit 30 in the first embodiment. However, the first object detection unit 30B uses the aggregated image generated by the aggregated image generation unit 70 as an input instead of using the image of the individual object detection target region as an input.

[0126] [Explanation of operation] Next, an example of the operation of the image processing device 1B according to the second embodiment will be described with reference to the drawings. Among the operations (steps) of the image processing device 1B according to the second embodiment, the same operations (steps) as those of the image processing device 1 of the first embodiment will be given the same step numbers. Also, detailed descriptions of those operations (steps) will be omitted.

[0127] (A2) Operation of object detection processing based on prediction FIG. 8 is a flow diagram showing an example of the operation of the object detection process in the image processing device 1B according to the present disclosure.

[0128] The image processing device 1B performs the processes of steps S100 to S102.

[0129] Next, the aggregate image generating unit 70 generates an aggregate image based on the object detection target region obtained in step S102 (step S200).

[0130] Next, for each aggregate image generated in step S200 (step S201), first object detection unit 30B performs object detection processing on the aggregate image (step S202). For example, first object detection unit 30B executes object detection processing using the aggregate image as an input. First object detection unit 30B converts bounding box information, which is position information of an object obtained as a result of the object detection processing, from coordinate values ​​in the aggregate image to coordinate values ​​in the input image input from image input unit 60. The processing of step S202 is repeated until it is executed on all generated aggregate images.

[0131] The image processing device 1B in the second embodiment performs object detection processing for the following operations (B) and (C) in the first embodiment by replacing operations for each individual object detection target area with operations for each aggregated image, similar to the operation of the object detection processing based on the prediction (A2) described above. (B) Operation of object detection processing using the second object detection unit (C) Object detection process using object position information of new objects

[0132] [Effect description] Next, the effects of the second embodiment will be described.

[0133] The image processing device 1B according to the second embodiment can improve the throughput of the object detection process, as in the first embodiment, and can also reduce the delay in the object detection process.

[0134] The image processing device 1B generates an aggregate image based on the individual object detection target region. Next, the image processing device 1B performs object detection for each generated aggregate image. The size of the aggregate image is expected to be smaller than the input image. Therefore, in the image processing device 1B according to the second embodiment, the target data amount of the object detection process is reduced, and the calculation load is reduced accordingly. Therefore, it is expected that the throughput of the object detection process is improved. In addition, it is expected that the delay of the object detection process is shortened. That is, in the image processing device 1B according to the second embodiment, the period from when the target object is captured in the image to when it is detected is shortened (i.e., the delay is shortened), making it possible to quickly respond to the target object.

[0135] Next, an outline of the present disclosure will be described. FIG. 9 is a block diagram showing an outline of an image processing device according to the present disclosure. The image processing device 100 (in the embodiment, the image processing device 1 or the image processing device 1B) shown in FIG. 9 includes an object position prediction means 110 (in the embodiment, the object position prediction unit 10) that predicts the position of an object in a new input image based on the position of the object detected in a past input image, an object detection target area determination means 120 (in the embodiment, the object detection target area determination unit 20) that determines an object detection target area to be a target of object detection in the new input image based on the prediction result by the object position prediction means 110, and an object detection means 130 (in the embodiment, the first object detection unit 30 or the first object detection unit 30B) that performs object detection on the object detection target area determined by the object detection target area determination means 120. The size of the object detection target area is expected to be smaller than the input image. That is, in the image processing device 100, the amount of data to be subjected to the object detection process is reduced, and the calculation load is reduced accordingly. Therefore, it is expected that the throughput of the object detection process is improved. It is also expected that delays in object detection processing will be reduced.

[0136] A part or all of the above-described embodiments can be described as, but is not limited to, the following supplementary notes.

[0137] (Supplementary Note 1) An image processing device comprising: an object position prediction means for predicting a position of an object in a new input image based on the position of the object detected in a past input image; an object detection target area determination means for determining an object detection target area in the new input image to be subject to object detection based on a prediction result by the object position prediction means; and an object detection means for performing object detection within the object detection target area determined by the object detection target area determination means.

[0138] (Appendix 2) An image processing device as described in Appendix 1, further comprising a second object detection means for performing object detection on a new input image, wherein the object detection target area determination means updates the object detection target area based on the object detection result by the second object detection means.

[0139] (Supplementary Note 3) An image processing device as described in Supplementary Note 1 or Supplementary Note 2, wherein the object detection target area determination means identifies a new object among the objects detected by the second object detection means that is not included in the object detection target area determined based on the prediction result by the object position prediction means, and updates the object detection target area so as to include the identified new object.

[0140] (Supplementary Note 4) The image processing device according to Supplementary Note 3, wherein the object-detection target area determination means updates the object-detection target area so as to include the new object whose position information has been updated based on the object detection result by the object detection means.

[0141] (Supplementary Note 5) An image processing device according to any one of Supplementary Note 1 to Supplementary Note 4, wherein the object position prediction means uses information that is independent of each other at each time when using position information of objects detected in a plurality of input images generated at different times as input for prediction.

[0142] (Appendix 6) An image processing method comprising: predicting a position of an object in a new input image based on a position of the object detected in a past input image; determining an object detection target area in the new input image to be subject to object detection based on the prediction result; and performing object detection in the determined object detection target area.

[0143] (Supplementary Note 7) The image processing method according to Supplementary Note 6, further comprising: performing a second object detection on a new input image; and updating the object detection target area based on the object detection result from the second object detection.

[0144] (Supplementary Note 8) An image processing method as described in Supplementary Note 6 or Supplementary Note 7, which identifies a new object that is not included in the object detection target area determined based on the prediction result among the objects detected by the second object detection, and updates the object detection target area to include the identified new object.

[0145] (Supplementary Note 9) The image processing method according to Supplementary Note 8, further comprising updating the object detection target area so as to include the new object whose position information has been updated based on the object detection result by the object detection.

[0146] (Supplementary Note 10) An image processing method according to any one of Supplementary Note 6 to Supplementary Note 9, in which when position information of objects detected in a plurality of input images generated at different times is used as input for prediction, information that is independent of each other at each time is used.

[0147] (Supplementary Note 11) An image processing program for causing a computer to execute an object position prediction process that predicts a position of an object in a new input image based on the position of the object detected in a past input image, an object detection target area determination process that determines an object detection target area in the new input image that is to be subject to object detection based on a prediction result of the object position prediction process, and an object detection process that performs object detection in the object detection target area determined by the object detection target area determination process.

[0148] (Appendix 12) An image processing program as described in Appendix 11, which causes a computer to execute a second object detection process that performs object detection on a new input image, and in the object detection target area determination process, updates the object detection target area based on the object detection result by the second object detection process.

[0149] (Supplementary Note 13) An image processing program as described in Supplementary Note 11 or Supplementary Note 12, in which, in the object detection target area determination process, a new object is identified among the objects detected by the second object detection process that is not included in the object detection target area determined based on the prediction result by the object position prediction process, and the object detection target area is updated to include the identified new object.

[0150] (Supplementary Note 14) The image processing program according to Supplementary Note 13, wherein in the object detection target area determination process, the object detection target area is updated to include the new object whose position information has been updated based on the object detection result by the object detection process.

[0151] (Appendix 15) An image processing program according to any one of appendices 11 to 14, wherein, in the object position prediction process, when position information of objects detected in multiple input images generated at different times is used as input for prediction, information that is independent of each other at each time is used. [Industrial Applicability]

[0152] The present disclosure is preferably applied to image processing using inference based on machine learning. [Explanation of symbols]

[0153] 1, 1B, 100 Image processing device 10 Object position prediction unit 20 Object detection target area determination unit 30, 30B First object detection unit 40 Second object detection unit 50 Object position memory section 60 Image input section 70 Aggregated image generation unit 110 Object position prediction means 120 Object detection target area determination means 130 Object detection means 600 Information processing device 610 CPU 611 Calculation Unit 620 ROM 630 RAM 640 Internal storage 650 IOC 660 Input Devices 670 Display equipment 680NIC 690 Storage medium

Claims

1. an object position prediction means for predicting a position of an object in a new input image based on a position of the object detected in a past input image; an object detection target area determining means for determining an object detection target area in the new input image to be subject to object detection based on a prediction result by the object position predicting means; an object detection means for detecting an object in the object detection target area determined by the object detection target area determination means; 13. An image processing device comprising:

2. A second object detection means is provided for performing object detection on a new input image; The object detection target area determining means updates the object detection target area based on an object detection result by the second object detection means.

2. The image processing device according to claim 1.

3. The object detection target area determining means identifies a new object that is not included in the object detection target area determined based on a prediction result by the object position predicting means, among the objects detected by the second object detection means, and updates the object detection target area so as to include the identified new object.

3. The image processing device according to claim 1.

4. The object detection target area determination means updates the object detection target area so as to include the new object whose position information has been updated based on the object detection result by the object detection means.

4. The image processing device according to claim 3.

5. The object position prediction means uses, as an input for prediction, position information of objects detected in a plurality of input images generated at different times, and uses information that is independent of each other at each time.

3. The image processing device according to claim 1.

6. predicting a position of an object in a new input image based on a position of the object detected in a previous input image; determining an object detection target region in the new input image to be subjected to object detection based on the prediction result; Object detection is performed within the determined object detection target area.

13. An image processing method comprising:

7. On the computer, an object position prediction process for predicting a position of an object in a new input image based on a position of the object detected in a past input image; an object detection target area determination process for determining an object detection target area in the new input image to be subject to object detection based on a prediction result of the object position prediction process; an object detection process for detecting an object in the object detection target area determined in the object detection target area determination process; An image processing program for executing the above.

Citation Information

Patent Citations

  • Control program, control method, and information processing device

    JP2019036008A