Object detection system and object detection method thereof

By adding time information to pyramid images and combining it with neural network deep learning, the problem of insufficient utilization of time information in existing technologies is solved, thereby improving the detection accuracy and tracking stability of object detection systems.

CN112508839BActive Publication Date: 2026-01-06SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010650026.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-26
Filing Date
2020-07-08
Publication Date
2026-01-06
Estimated Expiration
2040-07-08

AI Technical Summary

Technical Problem

Existing object detection methods struggle to effectively utilize temporal information when processing time-series images, resulting in insufficient detection performance.

Method used

By generating pyramid images and adding time information to them, combined with spatial information, object detection and tracking are performed, and deep learning using neural networks is used to improve detection performance.

Benefits of technology

It improves the detection accuracy and tracking stability of the object detection system and enhances the ability to identify objects in time-series images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112508839B_ABST
    Figure CN112508839B_ABST
Patent Text Reader

Abstract

An object detection system for detecting an object by using a hierarchical pyramid structure, including a pyramid image generator configured to receive a plurality of input images respectively corresponding to a plurality of time points, and to generate a plurality of pyramid images corresponding to each of the plurality of input images; an object extractor configured to generate a plurality of pieces of object data by extracting at least one object from the plurality of pyramid images; and a buffer to store the plurality of pieces of object data on an object basis.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to Korean Patent Application No. 10-2019-0104574, filed on August 26, 2019, with the Korean Intellectual Property Office, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure relates to an object detection system, and more specifically to an object detection system and method for detecting objects by using a hierarchical pyramid. Background Technology

[0004] Object detection refers to data processing methods that detect objects of interest from images or videos and identify or classify them. Object detection plays an important role in various applications such as autonomous driving, driver assistance, unmanned aerial vehicles, and gesture-based interaction systems.

[0005] Along with the development of artificial intelligence technology, object detection, object classification and recognition methods using neural network-based deep learning technology and training have been developed and have been widely deployed. Summary of the Invention

[0006] Embodiments of this disclosure provide an object detection system and an object detection method used by the object detection system, which can add time information indicating the time of capturing the input image to at least one pyramid image generated using an input image, and detect objects from the input image using the added time information.

[0007] According to an aspect of this disclosure, an object detection system is provided, comprising: a pyramid image generator configured to receive a first input image captured at a first time and a second input image captured at a second time, and to generate a first pyramid image from the first input image and a second pyramid image from the second input image; an object extractor configured to detect objects in the first pyramid image and the second pyramid image, and to generate multiple object data representing the objects; and a buffer for storing the multiple object data representing the objects detected in the first pyramid image and the second pyramid image.

[0008] According to another aspect of this disclosure, an object detection method is provided, comprising: receiving a first input image captured at a first time and a second input image captured at a second time; generating a first pyramid image associated with the first time from the first input image and generating a second pyramid image associated with the second time from the second input image; and storing multiple object data in a buffer.

[0009] According to another aspect of this disclosure, a driving assistance system for driving a vehicle by detecting objects is provided, the driving assistance system comprising: a pyramid image generator configured to receive a first input image captured at a first time and a second input image captured at a second time, and to generate a first pyramid image from the first input image and a second pyramid image from the second input image; an object extractor configured to detect objects in the first pyramid image and the second pyramid image, and to generate multiple object data representing the objects by using deep learning based on a neural network; a buffer for storing the multiple object data representing the objects detected in the first pyramid image and the second pyramid image; and an object tracker configured to track the objects based on the multiple object data stored in the buffer. Attached Figure Description

[0010] Embodiments of this disclosure will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings, wherein:

[0011] Figure 1 This is a block diagram illustrating an electronic system according to an embodiment of the present disclosure;

[0012] Figure 2 This is a block diagram illustrating an electronic system according to an embodiment of the present disclosure;

[0013] Figure 3 This is a block diagram illustrating an object detection system according to an embodiment of the present disclosure;

[0014] Figure 4 This is a flowchart illustrating a method for operating an object detection system according to an embodiment of the present disclosure;

[0015] Figure 5 This is a diagram illustrating a neural network according to an embodiment of the present disclosure;

[0016] Figure 6 This is a diagram illustrating a method for detecting an object according to an embodiment of the present disclosure;

[0017] Figure 7 This is a block diagram illustrating an object detection system according to an embodiment of the present disclosure;

[0018] Figure 8 This is a flowchart illustrating a method for operating an object detection system according to an embodiment of the present disclosure;

[0019] Figure 9 This is a block diagram illustrating an object detection system according to an embodiment of the present disclosure;

[0020] Figure 10 This is a diagram illustrating object data according to an embodiment of the present disclosure;

[0021] Figure 11 This is a block diagram illustrating an object detection system according to an embodiment of the present disclosure;

[0022] Figure 12 This is a block diagram illustrating an object detection system according to an embodiment of the present disclosure;

[0023] Figure 13 This is a diagram illustrating a method for generating a pyramid image according to an embodiment of the present disclosure;

[0024] Figure 14 This is a block diagram illustrating an object detection system according to an embodiment of the present disclosure;

[0025] Figure 15 This is a block diagram illustrating an application processor according to an embodiment of the present disclosure; and

[0026] Figure 16 This is a block diagram illustrating a driving system according to an embodiment of the present disclosure. Detailed Implementation

[0027] In the following, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.

[0028] Figure 1 This is a block diagram illustrating an electronic system according to an embodiment of the present disclosure.

[0029] refer to Figure 1 Electronic system 10 can extract information by analyzing input data in real time, and based on the extracted information, it can determine a situation or control components of electronic devices in which electronic system 10 is located. In one example, electronic system 10 can detect objects from an input image and track the detected objects. As used herein, the term "object" can refer to at least one selected from buildings, articles, people, animals, and factories of interest to the user or electronic system 10. For example, electronic system 10 can be applied to unmanned aerial vehicles, advanced driver assistance systems (ADAS), robotic devices, smart TVs, smartphones, medical devices, mobile devices, image display devices, measuring devices, Internet of Things (IoT) devices, and so on, and can also be installed in one of a variety of electronic devices.

[0030] Electronic system 10 may include at least one intellectual property (IP) block (IP1, IP2, IP3, ... IPn) and object detection system 100. For example, electronic system 10 may include first IP block IP1 to third IP block IP3, but electronic system may include any number of IP blocks.

[0031] Electronic system 10 may include various IP blocks. For example, IP blocks may include processing units, multiple cores included in the processing units, various sensor modules, multi-format codecs (MFC), video modules (e.g., camera interfaces, Joint Picture Experts Group (JPEG) processors, video processors, mixers, etc.), three-dimensional (3D) graphics cores, audio systems, drivers, display drivers, volatile memory, non-volatile memory, memory controllers, input and output interface blocks, cache memory, etc. Each of the first IP block IP1 to the third IP block IP3 may include at least one of the various IP blocks described above.

[0032] IP blocks can be interconnected via at least one system bus 12. For example, the Advanced Microcontroller Bus Architecture (AMBA) protocol from Advanced RISC Machine (ARM) Ltd. can be used as a standard bus specification. The bus types of the AMBA protocol can include Advanced High Performance Bus (AHB), Advanced Peripheral Bus (APB), Advanced Extensible Interface (AXI), AXI4, AXI Coherence Extensions (ACE), and so on.

[0033] Object detection system 100 can receive an input image, detect objects included in the input image, and track detected objects or extract background by excluding detected objects from the input image. As used herein, the term "object" can refer to at least one selected from buildings, articles, people, animals, and factories of interest to the user or electronic system 10. The term "background" can refer to the remainder of the image obtained by excluding objects from image frames. In one embodiment, object detection system 100 can detect objects included in input image data by using a neural network and can track the extracted objects. (See reference...) Figure 2 This will be described in detail. In one embodiment, the object detection system 100 can generate at least one pyramid image by downsampling an input image, and can extract objects based on the at least one pyramid image. Here, the structure for hierarchically extracting objects based on multiple pyramid images with different resolutions generated by downsampling a single image data point can be referred to as a pyramid structure. Reference will be made below. Figure 6 This will be described in detail. Furthermore, for ease of description, the terms "input image" and "pyramid image" used in this paper can refer to the data corresponding to the input image and the data corresponding to the pyramid image, respectively.

[0034] According to this disclosure, when generating at least one pyramid image, the object detection system 100 can add temporal information corresponding to the time of capturing the input image to the at least one pyramid image. When detecting and tracking objects, the object detection system 100 can use the temporal information in addition to the spatial information of the at least one pyramid. Therefore, the detection performance of the object detection system 100 can be improved.

[0035] Figure 2 This is a block diagram illustrating an electronic system according to an embodiment of the present disclosure. Specifically, Figure 2 The diagram shows Figure 1 An embodiment of the electronic system 10 is shown. Regarding... Figure 2 The electronic system 10 will omit the reference. Figure 1 The given description is repetitive.

[0036] refer to Figure 2 The electronic system 10 may include a central processing unit (CPU) 21, a neural network device 22, a random access memory (RAM) 24, a memory 25, and a sensor module 26. The electronic system 10 may also include input / output modules, security modules, power control devices, and various computing devices. In one embodiment, all or some of the components of the electronic system 10 (CPU 21, neural network device 22, RAM 24, memory 25, and sensor module 26) may be mounted on a single semiconductor chip. For example, the electronic system 10 may be implemented as a system-on-a-chip (SoC). The components of the electronic system 10 may communicate with each other via at least one system bus 27.

[0037] CPU 21 can control the overall operation of electronic system 10. CPU 21 may include a single processor core (i.e., single-core) or multiple processor cores (i.e., multi-core). CPU 21 can process or execute programs and / or data stored in memory 25 and loaded into RAM 24. In one embodiment, by executing a program stored in memory 25, CPU 21 can execute a reference... Figure 1 The described object detection system 100 operates and can control the function of the neural network device 22 for object detection. The neural network device 22 can generate a neural network, train the neural network (or make the neural network learn), or perform calculations based on received input data and generate information signals based on the results of the calculations, or retrain the neural network. In one example, the neural network device 22 can receive an input image and generate at least one object data point by extracting objects included in the input image via calculations included in the neural network. The neural network device 22 may be referred to as a computing unit, computing module, etc.

[0038] Neural network models can include a variety of models, such as convolutional neural networks (CNNs) including GoogleNet, AlexNet, VGG networks, etc., region-based neural networks (R-CNN), region scheme networks (RPN), recurrent neural networks (RNN), stack-based deep neural networks (S-DNN), state-space dynamic neural networks (S-SDNN), deconvolutional networks, deep trust networks (DBN), restricted Boltzmann machines (RBM), fully convolutional networks, long short-term memory networks (LSTM), classification networks, and so on, but are not limited to these.

[0039] The neural network device 22 may include one or more processors for performing computations based on a model of the neural network. Additionally, the neural network device 22 may include a separate memory for storing programs corresponding to the model of the neural network. The neural network device 22 may be referred to as a neural network processor, neural network processing device, neural network integrated circuit, neural network processing unit (NPU), etc.

[0040] The neural network device 22 and CPU 21 can be included in the reference. Figure 1 The object detection system 100 described herein can receive data corresponding to an input image from a specific IP (e.g., RAM 24 or sensor module 26) and can detect objects included in the input image. In one embodiment, a CPU 21 included in the object detection system 100 can generate at least one pyramid image with a pyramid structure using the input image, and the generated pyramid image can include temporal information corresponding to the time the input image was captured. Additionally, a neural network device 22 included in the object detection system 100 can extract objects included in the input image through deep learning trained based on the network, spatial information, and temporal information of the pyramid image, and can track the extracted objects.

[0041] RAM 24 can store programs, data, or instructions. For example, programs and / or data stored in memory 25 can be loaded into RAM 24 under the control of CPU 21 or according to boot code. RAM 24 can be implemented using memory such as dynamic RAM (DRAM) or static RAM (SRAM). Memory 25 is a storage location for storing data and can, for example, store an operating system (OS), various programs, and individual data entries. Memory 25 may include at least one selected from volatile memory and non-volatile memory. Sensor module 26 can collect information about the surroundings of electronic system 10. Sensor module 26 can sense or receive image signals from outside electronic system 10 and can convert the sensed or received image signals into image data, i.e., image frames. For this purpose, sensor module 26 may include sensing devices, such as at least one of various sensing devices such as image pickup devices, image sensors, light detection and ranging (LIDAR) sensors, ultrasonic sensors, and infrared sensors, or may receive sensing signals from sensing devices. In one embodiment, sensor module 26 may provide image data including image frames to CPU 21 or neural network device 22. For example, sensor module 26 may include an image sensor that can generate a video stream by capturing images of the environment outside electronic system 10 and can sequentially provide continuous image frames of the video stream to CPU 21 or neural network device 22.

[0042] The electronic system 10 according to embodiments of the present disclosure can add temporal information corresponding to the image capture time of the image data to at least one pyramid image when generating at least one pyramid image, and can use the temporal information together with spatial information based on at least one pyramid image when detecting and tracking objects using a neural network. Therefore, the object detection performance of the electronic system 10 can be improved. As used herein, the term "spatial information" can refer to pixel data of the input image.

[0043] Figure 3 This is a block diagram illustrating an object detection system according to an embodiment of the present disclosure.

[0044] refer to Figure 3 The object detection system 100 may include a pyramid image generator 110, a feature extractor 120, a buffer 130, and an object tracker 140. The pyramid image generator 110 may receive multiple input images IM captured at multiple time points, and may generate multiple pyramid images PI from the received multiple input images IM.

[0045] The pyramid image generator 110 can generate multiple pyramid images based on an input image corresponding to a point in time, and each of the multiple pyramid images can include temporal information about the time when the input image was captured. In one example, the pyramid image generator 110 can generate a first pyramid image with a first resolution corresponding to a first input image at a first point in time, generate a second pyramid image with a second resolution by downsampling the first pyramid image, generate a third pyramid image with a third resolution by downsampling the second pyramid image, and add data corresponding to the first point in time when the first input image was captured to the first, second, and third pyramid images. The multiple pyramid images generated by downsampling and having different resolutions can be adapted based on the number and / or type and category of objects in the input image IM.

[0046] The pyramid image generator 110 can generate multiple pyramid images for each input image corresponding to each of multiple time points. In one example, the pyramid image generator 110 can generate a fourth pyramid image with a first resolution corresponding to a second input image at a second time point, generate a fifth pyramid image with a second resolution by downsampling the fourth pyramid image, generate a sixth pyramid image with a third resolution by downsampling the fifth pyramid image, and add data corresponding to the second time point of capturing the second input image to the fourth, fifth, and sixth pyramid images. In other words, a single time point of capturing the input image can be added to all pyramid images generated based on that input image. In one example, the pyramid image generator 110 can obtain information about the image capture time point from the meta-region of the input image (e.g., IM), or from an external device (e.g., in...). Figure 2 The sensor module 26 in the middle obtains the image capture time point.

[0047] In one embodiment, the pyramid image generator 110 may add time information about the image capture time of the input image to the header region of each of the generated plurality of pyramid images. This disclosure is not limited thereto, and the regions for which the pyramid image generator 110 adds time information to each of the plurality of pyramid images may be determined differently.

[0048] In one embodiment, the pyramid image generator 110 may add time information only to at least some of the plurality of pyramid images generated from the input image. In one example, the pyramid image generator 110 may add time information to a first pyramid image having a first resolution and to a second pyramid image having a second resolution, but may not add time information to a third pyramid image having a third resolution. In other words, a single time point capturing the input image may be added only to some or a subset of the pyramid images generated from the input image.

[0049] In one embodiment, the pyramid image generator 110 may differently determine the number of pyramid images generated from multiple input images corresponding to multiple time points, each having a resolution. In one example, based on the multiple input images corresponding to multiple time points, the pyramid image generator 110 may generate a first number of first pyramid images with a first resolution and a second number of second pyramid images with a second resolution. That is, a first number of first pyramid images can be generated from a first number of input images captured at different time points, and a second number of second pyramid images can be generated from a second number of input images captured at different time points.

[0050] Feature extractor 120 can receive multiple pyramid images PI from pyramid image generator 110 and can extract multiple object data ODs from the multiple pyramid images PI. In one embodiment, feature extractor 120 can extract multiple object data ODs from the multiple pyramid images PI using deep learning trained on a neural network. In one example, it can be done by... Figure 2 The feature extractor 120 is implemented using a neural network device 22.

[0051] According to this disclosure, feature extractor 120 can extract object data corresponding to the same object from multiple pyramid images corresponding to multiple time points. Feature extractor 120 can receive multiple pyramid images corresponding to multiple time points and a resolution from pyramid image generator 110, and can detect and extract an object based on time information included in the multiple pyramid images, thereby generating object data. In one example, feature extractor 120 can extract a first object from at least one first pyramid image with a first resolution, extract a second object from at least one second pyramid image with a second resolution, and extract a third object from at least one third pyramid image with a third resolution. In one embodiment, the first to third objects may be located at different distances from the image capture location, and reference... Figure 6This will be described in more detail. Feature extractor 120 can store at least one extracted object data OD in buffer 130. In one embodiment, feature extractor 120 can store at least one object data OD in different buffers or different areas of a buffer, depending on the type of object.

[0052] In one example, feature extractor 120 can store object data ODs corresponding to a first object and each corresponding to multiple time points in a first region of buffer 130, object data ODs corresponding to a second object and each corresponding to multiple time points in a second region of buffer 130, and object data ODs corresponding to a third object and each corresponding to multiple time points in a third region of buffer 130. In one example, feature extractor 120 can store multiple object data in buffer 130 on an object-by-object basis based on cascading operations.

[0053] Buffer 130 can store object data OD. For this purpose, buffer 130 may include at least one selected from volatile memory and non-volatile memory. According to one embodiment of this disclosure, buffer 130 may store object data OD in different areas therein on an object-by-object basis. In another embodiment, buffer 130 may include multiple storage devices and may store object data OD in different storage devices on an object-by-object basis.

[0054] Object tracker 140 can receive object data OD and can track objects based on the object data OD. In one embodiment of this disclosure, when tracking an object, object tracker 140 can use object data OD corresponding to multiple time points respectively. In one example, object tracker 140 can track a first object by using multiple object data corresponding to a first resolution and can track a second object by using multiple object data corresponding to a second resolution.

[0055] According to one embodiment of this disclosure, object tracker 140 can extract an object using object data ODs corresponding to multiple time points, respectively. In one example, the object may have a larger amount of data change over time compared to the background, and object tracker 140 can effectively track the object by comparing multiple object data ODs corresponding to multiple time points with each other.

[0056] Figure 4 This is a flowchart illustrating a method for operating an object detection system according to an embodiment of the present disclosure.

[0057] refer to Figure 3 and Figure 4The object detection system 100 can receive multiple input images corresponding to multiple time points (S110), and can add time information about the image capture time of each input image to the multiple input images (S120). The object detection system 100 can generate multiple pyramid images corresponding to one of the multiple input images by adding time information to the multiple input images. In one embodiment, the object detection system 100 can generate multiple pyramid images by repeatedly downsampling each input image to which time information is added.

[0058] The object detection system 100 can generate object data corresponding to multiple time points by extracting objects from each of multiple pyramid images (S140). In one embodiment, the object detection system 100 can generate multiple object data from multiple pyramid images by using a deep learning model trained on a neural network. In one example, the object detection system 100 can generate multiple time-point-by-time-point object data corresponding to a single object.

[0059] The object detection system 100 can store multiple object data entries on an object-by-object basis (S150). In one embodiment, the object detection system 100 can store multiple point-in-time object data entries in different areas of a buffer on an object-by-object basis, and can also store multiple point-in-time specific object data entries in the buffer using a cascading operation. The object detection system 100 can track the position and / or movement of objects by using both the multiple object data entries stored on an object-by-object basis and time information (S160).

[0060] Figure 5 This is a diagram illustrating a neural network according to an embodiment of the present disclosure. Specifically, Figure 5 The diagram illustrates the structure of a convolutional neural network as an example of a neural network structure. Figure 5 The diagram shows the... Figure 3 An example of the neural network used by feature extractor 120.

[0061] refer to Figure 5 A neural network NN may include multiple layers L1, L2, ... through Ln. Each of the multiple layers L1, L2, ... through Ln may be a linear or non-linear layer. In one embodiment, at least one linear layer and at least one non-linear layer may be coupled to each other and are therefore referred to as a layer. For example, a linear layer may include convolutional layers and fully connected layers, and a non-linear layer may include pooling and activation layers.

[0062] For example, the first layer L1 can be a convolutional layer, the second layer L2 can be a pooling layer, and the nth layer Ln can be a fully connected layer that serves as the output layer. A neural network NN can also include activation layers and layers that perform operations other than those discussed above.

[0063] Each of the multiple layers L1 to Ln can receive input data (e.g., image frames) or feature maps generated from previous layers as input feature maps, and can compute the input feature maps to generate an output feature map or recognition signal REC. Here, a feature map refers to data in which various features of the input data are represented. Each feature map FM1 to FMn can be in the form of, for example, a 2D matrix or a 3D matrix (or tensor). Each feature map FM1 to FMn can have a width W (or columns), a height H (or rows), and a depth D, which can correspond to the x-axis, y-axis, and z-axis in a coordinate system, respectively. Here, the depth D can be referred to as the number of channels.

[0064] The first layer L1 generates the second feature map FM2 through convolution of the first feature map FM1 and the weight map WM. The weight map WM can filter the first feature map FM1 and can also be referred to as a filter or kernel. The depth of the weight map WM, i.e., the number of channels, can be equal to the depth of the first feature map FM1, i.e., the number of channels, and convolution can be performed between the same channels of both the weight map WM and the first feature map FM1. The first feature map FM1 can be used as a sliding window, shifting the weight map WM by traversing it. The amount of shift can be referred to as the term "stride length" or "step". During each shift, the weight values ​​included in the weight map WM can be multiplied by all pixel data in the region overlapping with the first feature map. The results can then be summed. The data in the region of the first feature map FM1 overlapping with each weight value included in the weight map WM can be referred to as extracted data. When convolution is performed between the first feature map FM1 and the weight map WM, one channel of the second feature map FM2 can be generated. Figure 3 The diagram shows a weight map WM, but multiple weight maps can be convolved with the first feature map FM1, thereby generating multiple channels of the second feature map FM2. In other words, the number of channels in the second feature map FM2 can correspond to the number of weight maps.

[0065] The second layer L2 can generate a third feature map FM3 by modifying the spatial size of the second feature map FM2 via pooling. The term "pooling" can also be referred to as "sampling" or "downsampling." The pooling window PW can be shifted over the second feature map FM2 in units of the size of a 2D pooling window PW, and the maximum value (or average value) of the pixel data in the region overlapping with the pooling window PW can be selected. Thus, a third feature map FM3 with a spatial size different from that of the second feature map FM2 can be generated. The number of channels in the third feature map FM3 is equal to the number of channels in the second feature map FM2. In one embodiment of this disclosure, various operations for object extraction can be performed while the pooling window PW is shifted over feature maps FM2 and FM3.

[0066] The nth layer Ln can classify the category (i.e., CL) of the input data by combining the features of the nth feature map FMn. Additionally, the nth layer Ln can generate a recognition signal REC corresponding to the category. In one embodiment, the input data can correspond to a pyramid image generated using the input image, and the nth layer Ln can identify objects by extracting the category corresponding to the objects included in the image represented by the frame data based on the nth feature map FMn provided by the previous layer. Therefore, a recognition signal REC corresponding to the identified object can be output. In one embodiment, the feature extractor 120 ( Figure 3 The identification signal REC can be stored in buffer 130. Figure 3 The object data OD can be stored in buffer 130, or the object data OD generated using the recognition signal REC can be stored in buffer 130. Figure 3 )middle.

[0067] Figure 6 This is a diagram illustrating a method for detecting an object according to an embodiment of the present disclosure.

[0068] refer to Figure 3 and Figure 6 The pyramid image generator 110 can generate a first pyramid image PI1_1 with a first resolution based on the input image captured at a first time point t1. The pyramid image generator 110 can generate a second pyramid image PI1_2 with a second resolution by downsampling the first pyramid image PI1_1. The pyramid image generator 110 can generate a third pyramid image PI1_3 with a third resolution by downsampling the second pyramid image PI1_2. The pyramid image generator 110 can perform downsampling based on a preset integer ratio, and in one example, the pyramid image generator 110 can perform downsampling by multiplying the resolution of the existing image by... or To perform downsampling.

[0069] The pyramid image generator 110 can generate a fourth pyramid image PI2_1 with a first resolution based on the input image captured at a second time point t2. The pyramid image generator 110 can generate a fifth pyramid image PI2_2 with a second resolution by downsampling the fourth pyramid image PI2_1. The pyramid image generator 110 can generate a sixth pyramid image PI2_3 with a third resolution by downsampling the fifth pyramid image PI2_2.

[0070] According to one embodiment of the present disclosure, the pyramid image generator 110 can add time information corresponding to a first time point t1 to the first pyramid image PI1_1 to the third pyramid image PI1_3, and can add time information corresponding to a second time point t2 to the fourth pyramid image PI2_1 to the sixth pyramid image PI2_3.

[0071] Feature extractor 120 can extract multiple objects from different pyramid images. In one example, feature extractor 120 can extract the first object O1, which is closest to the image capture device that generated the input image, by using a third pyramid image PI1_3, which both have a third resolution as the lowest resolution, and a sixth pyramid image PI2_3. Similarly, feature extractor 120 can extract the second object O2, which is second closest to the image capture device that generated the input image, by using a second pyramid image PI1_2, which both have a second resolution as the second lowest resolution, and a fifth pyramid image PI2_2. Furthermore, feature extractor 120 can extract the third object O3, which is second closest to the image capture device that generated the input image, by using a first pyramid image PI1_1, which both have a first resolution as the highest resolution, and a fourth pyramid image PI2_1.

[0072] Object tracker 140 can track objects based on multiple object data generated by feature extractor 120. According to one embodiment of this disclosure, to track an object, object tracker 140 can use object data corresponding to multiple time points respectively using time information. In one example, to track a third object O3, in addition to using object data generated from the first pyramid image PI1_1, object tracker 140 can also use object data generated from the fourth pyramid image PI2_1 and the time difference between the first time point t1 and the second time point t2.

[0073] although Figure 6The illustration shows an example of extracting three objects using pyramid images with three resolutions, but this is merely an example; the number of pyramid images used for object extraction can be determined differently, and the number of objects extracted using these pyramid images can also be determined differently. Furthermore, it should be understood that this disclosure can also be applied to embodiments in which two or more objects are extracted using a single pyramid image.

[0074] Figure 7 This is a block diagram illustrating an object detection system according to an embodiment of the present disclosure. Specifically, Figure 7 This is a block diagram illustrating an embodiment of extracting an object based on input images captured at two time points. References will be omitted. Figure 3 Redundant description.

[0075] refer to Figure 6 and Figure 7 The object detection system 100 may include a pyramid image generator 110, a first feature extractor 121, a second feature extractor 122, a buffer 130, and an object tracker 140. The pyramid image generator 110 may receive a first input image IM1 captured at a first time point t1 and a second input image IM2 captured at a second time point t2. The pyramid image generator 110 may include a data manager 111 and a downsampling unit 112. The data manager 111 can generate a first pyramid image PI1_1 by adding time information corresponding to the first time point t1 to the first input image IM1, and can generate a fourth pyramid image PI2_1 by adding time information corresponding to the second time point t2 to the second input image IM2.

[0076] Downsampler 112 can generate second pyramid image PI1_2 and third pyramid image PI1_3 by downsampling the first pyramid image PI1_1. Additionally, downsampler 112 can generate fifth pyramid image PI2_2 and sixth pyramid image PI2_3 by downsampling the fourth pyramid image PI2_1.

[0077] The pyramid image generator 110 can output the generated first pyramid image PI1_1 to the third pyramid image PI1_3 to the first feature extractor 121, and can output the generated fourth pyramid image PI2_1 to the sixth pyramid image PI2_3 to the second feature extractor 122. The first feature extractor 121 can receive the first pyramid image PI1_1 to the third pyramid image PI1_3 corresponding to the first time point t1, and can generate first to third object data OD1_1, OD1_2, and OD1_3 by extracting objects from the received first pyramid image PI1_1 to the third pyramid image PI1_3 respectively. Figure 6In the example, the first feature extractor 121 can generate first object data OD1_1 by extracting a first object from the first pyramid image PI1_1, generate second object data OD1_2 by extracting a second object from the second pyramid image PI1_2, and generate third object data OD1_3 by extracting a third object from the third pyramid image PI1_3. The second feature extractor 122 can receive fourth to sixth pyramid images PI2_1, PI2_2, and PI2_3 corresponding to the second time point t2, and can generate fourth to sixth object data OD2_1 to OD2_3 by extracting objects from the received fourth to sixth pyramid images PI2_1 to PI2_3, respectively.

[0078] The first feature extractor 121 can store the generated first object data OD1_1 in the first region Ar1 of the buffer 130, the generated second object data OD1_2 in the second region Ar2 of the buffer 130, and the generated third object data OD1_3 in the third region Ar3 of the buffer 130. The second feature extractor 122 can store the generated fourth object data OD2_1 in the first region Ar1 of the buffer 130, the generated fifth object data OD2_2 in the second region Ar2 of the buffer 130, and the generated sixth object data OD2_3 in the third region Ar3 of the buffer 130.

[0079] In one embodiment, the first feature extractor 121 and the second feature extractor 122 can store the generated first object data to sixth object data OD1_1, OD1_2, OD1_3, OD2_1, OD2_2, and OD2_3 in buffer 130 using a cascading operation. Furthermore, although... Figure 7 The illustration shows an embodiment in which object data (e.g., OD1_1 to OD2_3) is stored on an object-by-object basis in different regions (e.g., Ar1 to Ar3) of a buffer 130, but this disclosure can also be applied to embodiments in which object data (e.g., OD1_1 to OD2_3) is stored on an object-by-object basis in different buffers, as described above.

[0080] Object tracker 140 can track objects by using object-based first object data to sixth object data OD1_1 to OD2_3. In one example, object tracker 140 can read the first object data OD1_1 and the fourth object data OD2_1 stored in the first region Ar1 of buffer 130, and can track the first object by using the first object data OD1_1 and the fourth object data OD2_1. Although Figure 7The figure illustrates an embodiment in which an object is extracted based on input images corresponding to two time points respectively, but this is merely an example, and it should be understood that this disclosure can also be applied to embodiments in which an object is extracted based on input images corresponding to more than two time points respectively.

[0081] Figure 8 This is a flowchart illustrating a method for operating an object detection system according to an embodiment of the present disclosure. Specifically, Figure 8 The figure illustrates a method for an operational object detection system that detects objects by using a different number of pyramid images for each resolution.

[0082] refer to Figure 3 and Figure 8 The object detection system 100 can generate a first pyramid image set with a first resolution by using multiple input images corresponding to multiple time points (S210). The object detection system 100 can generate a second pyramid image set with a second resolution by downsampling at least some of the multiple pyramid images included in the first pyramid image set (S220).

[0083] The object detection system 100 can generate N first object data points (where N is a natural number) corresponding to N time points respectively by extracting first objects from a first pyramid image set (S230). The object detection system 100 can generate M second object data points (where M is a natural number different from N) corresponding to M time points respectively by extracting second objects from a second pyramid image set (S240). The object detection system 100 can store the N first object data points in a first region of the buffer 130 (S250) and can store the M second object data points in a second region of the buffer 130 (S260). In one embodiment, the number N of first object data points can be greater than the number M of second object data points. According to one embodiment of this disclosure, the object detection system 100 can generate object data by extracting objects using a different number of pyramid images for each resolution. In one example, when the first object has insufficient spatial information compared to when the second object has relatively more spatial information, the object detection system 100 can generate object data by using more pyramid images. In other words, objects located far from the location where the image is captured can appear smaller in the image. Therefore, the object can be represented using a relatively small amount of information and / or pixels. Thus, for objects with insufficient spatial information in an image, the object can be extracted by using an increased number of pyramid images. Therefore, additional spatial and pixel information about the object can be obtained from the additional pyramid images, thereby improving object extraction performance.

[0084] Figure 9 This is a block diagram illustrating an object detection system according to an embodiment of the present disclosure. Specifically, Figure 9 The diagram illustrates an object detection system that detects objects by using a different number of pyramid images for each resolution. (References omitted) Figure 7 Redundant description.

[0085] refer to Figure 9 The object detection system 100 may include a pyramid image generator 110, a first feature extractor 121, a second feature extractor 122, a third feature extractor 123, a buffer 130, and an object tracker 140. The pyramid image generator 110 may receive a first input image IM1 captured at a first time point t1, a second input image IM2 captured at a second time point t2, and a third input image IM3 captured at a third time point t3.

[0086] The pyramid image generator 110 can generate a first pyramid image PI1_1 by adding time information corresponding to a first time point t1 to the first input image IM1, generate a second pyramid image PI1_2 by downsampling the first pyramid image PI1_1, and generate a third pyramid image PI1_3 by downsampling the second pyramid image PI1_2. The pyramid image generator 110 can output the first pyramid image PI1_1 to the third pyramid image PI1_3 to the first feature extractor 121 as a first pyramid image set PS1.

[0087] The pyramid image generator 110 can generate a fourth pyramid image PI2_1 by adding time information corresponding to the second time point t2 to the second input image IM2, and can generate a fifth pyramid image PI2_2 by downsampling the fourth pyramid image PI2_1. The pyramid image generator 110 can output the fourth pyramid image PI2_1 and the fifth pyramid image PI2_2 to the second feature extractor 122 as the second pyramid image set PS2. The pyramid image generator 110 can generate a sixth pyramid image PI3_1 by adding time information corresponding to the third time point t3 to the third input image IM3, and can output the sixth pyramid image PI3_1 to the third feature extractor 123 as the third pyramid image set PS3.

[0088] The first feature extractor 121 can receive first pyramid images PI1_1 to third pyramid images PI1_3 corresponding to the first time point t1, and can generate first object data to third object data OD1_1, OD1_2, and OD1_3 by extracting objects from the received first pyramid images PI1_1 to third pyramid images PI1_3 respectively. Figure 9In the example, the first feature extractor 121 can generate first object data OD1_1 by extracting a first object from the first pyramid image PI1_1, generate second object data OD1_2 by extracting a second object from the second pyramid image PI1_2, and generate third object data OD1_3 by extracting a third object from the third pyramid image PI1_3. The first feature extractor 121 can store the generated first object data OD1_1 in the first region Ar1 of the buffer 130, store the generated second object data OD1_2 in the second region Ar2 of the buffer 130, and store the generated third object data OD1_3 in the third region Ar3 of the buffer 130.

[0089] Feature extractor 122 can receive the fourth pyramid image PI2_1 and the fifth pyramid image PI2_2 corresponding to the second time point t2, and can generate fourth object data OD2_1 and fifth object data OD2_2 by extracting objects from the received fourth pyramid image PI2_1 and fifth pyramid image PI2_2, respectively. The second feature extractor 122 can store the generated fourth object data OD2_1 in the first region Ar1 of buffer 130, and can store the generated fifth object data OD2_2 in the second region Ar2 of buffer 130.

[0090] The third feature extractor 123 can receive the sixth pyramid image PI3_1 corresponding to the third time point t3, and can generate sixth object data OD3_1 by extracting the third object from the received sixth pyramid image PI3_1. The third feature extractor 123 can store the generated sixth object data OD3_1 in the first region Ar1 of the buffer 130.

[0091] Object tracker 140 can track objects by using first object data OD1_1 to sixth object data OD3_1 stored on an object-based basis. In one example, object tracker 140 can track a first object by using first object data OD1_1, fourth object data OD2_1, and sixth object data OD3_1 stored in a first region Ar1 of buffer 130.

[0092] According to one embodiment of this disclosure, the object detection system 100 can detect objects by using a different number of pyramid images for each object. In one example, the object detection system 100 can detect a third object by using three pyramid images (e.g., PI1_1, PI2_1, and PI3_1), a second object by using two pyramid images (e.g., PI1_2 and PI2_2), and a first object by using one pyramid image (e.g., PI1_3). In one embodiment, when the object is further away from the image capture location of the captured image, the object detection system 100 can detect the object by using more pyramid images.

[0093] Figure 10 This is a diagram illustrating object data according to an embodiment of the present disclosure. Specifically, Figure 10 The figure illustrates an embodiment in which the object detection system generates a different number of object data for each object.

[0094] refer to Figure 9 and Figure 10 The object detection system 100 can store the third object data OD1_3 corresponding to the first object O1 in the third region Ar3 of the buffer 130, the second object data OD1_2 and the fifth object data OD2_2 corresponding to the second object O2 in the second region Ar2 of the buffer 130, and the first object data OD1_1, the fourth object data OD2_1 and the sixth object data OD3_1 corresponding to the third object O3 in the first region Ar1 of the buffer 130.

[0095] The first object O1 can be an object relatively close to the image capture device, and relatively a large amount of spatial information can exist about the first object O1. In other words, an object located near the location where the image is captured can appear relatively large in the image. Therefore, the object can be represented by a correspondingly large amount of information and / or pixels. Therefore, the object detection system 100 can detect the first object O1 by using only the third object data OD1_3 corresponding to a first time point t1. On the other hand, the third object O3 can be an object relatively far from the image capture device, and relatively little spatial information can exist about the third object O3. Therefore, by using the first object data OD1_1, the fourth object data OD2_1, and the sixth object data OD3_1 corresponding to multiple time points (e.g., from the first time point t1 to the third time point t3), the object detection system 100 can supplement the relatively small amount of spatial information with object data corresponding to multiple time points respectively, and thus, can perform effective object detection.

[0096] Figure 11This is a block diagram illustrating an object detection system according to an embodiment of the present disclosure. Specifically, Figure 11 The figure illustrates an embodiment in which the object tracker 140 selectively determines the amount of object data required for object tracking. References will be omitted. Figure 7 Redundant description.

[0097] refer to Figure 11 The object detection system 100 may include a pyramid image generator 110, a first feature extractor 121, a second feature extractor 122, a third feature extractor 123, a buffer 130, and an object tracker 140. The pyramid image generator 110 may receive a first input image IM1 captured at a first time point t1, a second input image IM2 captured at a second time point t2, and a third input image IM3 captured at a third time point t3.

[0098] The pyramid image generator 110 can output first pyramid images to third pyramid images PI1_1, PI1_2, and PI1_3 to the first feature extractor 121 as a first pyramid image set PS1. The first pyramid images PI1_1 to the third pyramid images PI1_3 are generated by the method described above. Similarly, the pyramid image generator 110 can output fourth pyramid images to sixth pyramid images PI2_1, PI2_2, and PI2_3 to the second feature extractor 122 as a second pyramid image set PS2, and can output seventh pyramid images to ninth pyramid images PI3_1, PI3_2, and PI3_3 to the third feature extractor 123 as a third pyramid image set PS3.

[0099] The first feature extractor 121 can generate first object data OD1_1 to third object data OD1_3 by extracting objects from the first pyramid image PI1_1 to the third pyramid image PI1_3 corresponding to the first time point t1, respectively. The first feature extractor 121 can store the generated first object data OD1_1 in the first region Ar1 of the buffer 130, store the generated second object data OD1_2 in the second region Ar2 of the buffer 130, and store the generated third object data OD1_3 in the third region Ar3 of the buffer 130.

[0100] Similarly, the second feature extractor 122 can generate fourth object data OD2_1 to sixth object data OD2_3 by extracting objects from the fourth pyramid image PI2_1 to the sixth pyramid image PI2_3 corresponding to the second time point t2, respectively. The second feature extractor 122 can store the fourth object data OD2_1 in the first region Ar1 of the buffer 130, store the fifth object data OD2_2 in the second region Ar2 of the buffer 130, and store the sixth object data OD2_3 in the third region Ar3 of the buffer 130.

[0101] The third feature extractor 123 can generate seventh object data OD3_1 to ninth object data OD3_3 by extracting objects from the seventh pyramid image PI3_1 to the ninth pyramid image PI3_3 corresponding to the third time point t3, respectively. The third feature extractor 123 can store the seventh object data OD3_1 in the first region Ar1 of the buffer 130, store the eighth object data OD3_2 in the second region Ar2 of the buffer 130, and store the ninth object data OD3_3 in the third region Ar3 of the buffer 130.

[0102] Object tracker 140 can track objects by reading at least some of the object data stored on an object-based basis (e.g., OD1_1 to OD3_3). According to one embodiment of this disclosure, object tracker 140 can track objects by using only some of the object data stored on an object-based basis (e.g., OD1_1 to OD3_3). In one example, object tracker 140 can track a first object by using only some of the first object data OD1_1, the fourth object data OD2_1, and the seventh object data OD3_1 corresponding to the first object.

[0103] Figure 12 This is a block diagram illustrating an object detection system according to an embodiment of the present disclosure. Specifically, Figure 12 The diagram illustrates an object detection system that detects objects based on regions of interest (ROIs). (References omitted.) Figure 3 Redundant description.

[0104] refer to Figure 12The object detection device 100a may include a pyramid image generator 110, a feature extractor 120, a buffer 130, an object tracker 140, and a Region of Interest (ROI) manager 150. The ROI manager 150 can identify regions within an input image IM as Regions of Interest (ROIs) based on the input image IM, and can output ROI information RI that includes data indicating the ROIs. For example, when the object detection device 100a is included in a driver assistance system, the ROI manager 150 can analyze the input image IM to identify regions containing information necessary for vehicle driving as ROIs. For example, an ROI could include a road, another vehicle, traffic lights, a pedestrian crossing, etc.

[0105] The ROI manager 150 may include a depth generator 151. The depth generator 151 can generate a depth map that includes depth data about objects and background included in the input image IM. In one example, the input image IM may include left-eye and right-eye images, and the depth generator 151 can calculate the depth using the left-eye and right-eye images and obtain a depth map based on the calculated depth. In another example, the depth generator 151 can obtain a depth map about objects and background included in the input image IM using 3D information obtained from a distance sensor.

[0106] The ROI manager 150 can generate ROI information (RI) using a depth map generated by the depth generator 151. In one example, the ROI manager 150 can set an area within a specific distance as an ROI based on the depth map.

[0107] The ROI manager 150 can output the generated ROI information RI to the pyramid image generator 110, and the pyramid image generator 110 can generate a pyramid image PI based on the ROI information RI. In one embodiment, the pyramid image generator 110 can mask input images IM that are not part of the ROI based on the ROI information RI, and can generate the pyramid image PI by using only the unmasked portions. In other words, the pyramid image generator 110 can disregard regions of the input image IM outside the region of interest (ROI) indicated by the ROI information. This improves the efficiency of the pyramid image generator 110 in generating the input image IM.

[0108] Figure 13 This is a diagram illustrating a method for generating a pyramid image according to an embodiment of the present disclosure.

[0109] refer to Figure 12 and Figure 13The pyramid image generator 110 can receive a first input image IM1 corresponding to a first time point t1 and can add time information corresponding to the first time point t1 to the first input image IM1. Additionally, the pyramid image generator 110 can generate a first pyramid image PI1_1 by masking the region outside the ROI based on ROI information RI. The ROI can include all first objects O1 to third objects O3.

[0110] The pyramid image generator 110 can generate a second pyramid image PI1_2 by downsampling the masked first pyramid image PI1_1, and can generate a third pyramid image PI1_3 by downsampling the second pyramid image PI1_2. The object detection system 100 can detect a third object O3 using the masked first pyramid image PI1_1, a second object O2 using the masked second pyramid image PI1_2, and a third object O3 using the masked third pyramid image PI1_3. According to one embodiment of this disclosure, by detecting objects after masking the input image, the masked region outside the ROI can be disregarded, and detection performance can be improved.

[0111] Figure 14 This is a block diagram illustrating an object detection system according to an embodiment of the present disclosure. Specifically, Figure 14 The diagram illustrates an object detection system that uses object data to detect the background. (References omitted) Figure 3 Redundant description.

[0112] refer to Figure 14 The object detection system 100b may include a pyramid image generator 110, a feature extractor 120, a buffer 130, an object tracker 140, and a background extractor 160. The background extractor 160 may receive multiple object data ODs stored on an object-based basis from the buffer, and may extract the background of an input image IM based on the multiple object data ODs. In one example, the background extractor 160 may extract the background by removing at least one object from the input image IM based on the object data ODs. According to one embodiment of this disclosure, the background extractor 160 may remove objects from the background based on object data ODs corresponding to multiple time points.

[0113] Figure 15 This is a block diagram illustrating an application processor according to an embodiment of the present disclosure. Figure 15 The application processor 1000 shown can be a semiconductor chip and can be implemented via a system-on-a-chip (SoC).

[0114] Application processor 1000 may include processor 1010 and operating memory 1020. Additionally, application processor 1000 may include one or more IP modules connected to a system bus. Operating memory 1020 may store software such as various programs and instructions related to the operation of a system using application processor 1000. For example, operating memory 1020 may include operating system 1021, neural network module 1022, and object detection module 1023. Processor 1010 may execute object detection module 1023 loaded into operating memory 1020, and according to the embodiments described above, may perform the function of detecting objects from an input image based on time information.

[0115] One or more pieces of hardware may include a processor 1010 and may perform neural network operations by executing a neural network module 1022, and as an example, one or more pieces of hardware may generate object data from a pyramid image according to the embodiments described above.

[0116] Figure 16 This is a block diagram illustrating a driving system according to an embodiment of the present disclosure.

[0117] refer to Figure 16 The driving assistance system 2000 may include a processor 2010, a sensor unit 2040, a communication module 2050, a driving control unit 2060, an autonomous driving unit 2070, and a user interface 2080. The processor 2010 can control the overall operation of the driving assistance system 2000, and according to the embodiments described above, can detect objects from input images received from the sensor unit 2040 with reference to time information.

[0118] Sensor unit 2040 can collect information about objects sensed by driving assistance system 2000. In one example, sensor unit 2040 can be an image sensor unit and may include at least one image sensor. Sensor unit 2040 can sense or receive image signals from outside driving assistance system 2000 and can convert image signals into image data, i.e., image frames.

[0119] In another example, sensor unit 2040 may be a distance sensor unit and may include at least one distance sensor. The distance sensor may include, for example, at least one of a variety of sensing devices such as a light detection and ranging (LIDAR) sensor, a radio detection and ranging (RADAR) sensor, a time-of-flight (ToF) sensor, an ultrasonic sensor, an infrared sensor, and so on. Each of the LIDAR and RADAR sensors may be classified based on the effective measurement distance. For example, LIDAR sensors may be classified as long LIDAR sensors and short LIDAR sensors, and RADAR sensors may be classified as long RADAR sensors and short RADAR sensors. This disclosure is not limited thereto, and sensor unit 2040 may include at least one selected from, but not limited to, a geomagnetic sensor, a position sensor (e.g., a Global Positioning System (GPS)), an accelerometer, a barometric pressure sensor, a temperature / humidity sensor, a proximity sensor, and a gyroscope.

[0120] The communication module 2050 can transmit and receive data from the driving assistance system 2000. In one example, the communication module 2050 can perform communication in a vehicle-to-everything (V2X) manner. For example, the communication module 2050 can perform communication in a vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), vehicle-to-pedestrian (V2P), and vehicle-to-mobile device (V2N) manner. However, this disclosure is not limited thereto, and the communication module 2050 can transmit and receive data in a variety of communication methods known to the public. For example, the communication module 2050 can perform communication through communication methods such as 3G, LTE, Wi-Fi, Bluetooth, Bluetooth Low Energy (BLE), Zigbee, Near Field Communication (NFC), or ultrasonic communication, and may include both short-range and long-range communication.

[0121] Sensor unit 2040 can generate an input image by capturing images of the external environment or surroundings of the driver assistance system 2000 and can transmit the input image to processor 2010. Processor 2010 can detect objects (e.g., another vehicle) based on the input image and the time it was captured, and can control driving control unit 2060 and autonomous driving unit 2070. Although an example is provided in which processor 2010 detects objects based on an input image, in another example, processor 2010 can detect objects based on depth information output by a distance sensor.

[0122] The driving control unit 2060 may include: a vehicle steering device configured to control the direction of the vehicle; a throttle device configured to control acceleration and / or deceleration by controlling the vehicle's motor or engine; a braking system configured to control the vehicle's brakes; external lighting devices; and so on. The autonomous driving unit 2070 may include a computing device configured to implement autonomous control of the driving control unit 2060. For example, the autonomous driving unit 2070 may include at least one component of the driving assistance system 2000. The autonomous driving unit 2070 may include a memory storing multiple program instructions and one or more processors executing the program instructions. The autonomous driving unit 2070 may be configured to control the driving control unit 2060 based on sensing signals output from the sensor unit 2040. The user interface 2080 may include various electronic and mechanical devices, such as a display in the driver's seat showing the vehicle's dashboard, passenger seats, and so on.

[0123] The processor 2010 can use various sensing data, such as input images, depth information, etc., when detecting objects. In this case, the processor 2010 can use artificial neural networks for efficient operational processing and can execute any of the object detection methods described in this disclosure.

[0124] Although this disclosure has been specifically shown and described with reference to its embodiments, it will be understood that various changes in form and detail may be made therein without departing from the spirit and scope of the appended claims.

Claims

1. A subject detection system, comprising: a pyramid image generator configured to receive a first input image captured at a first time and a second input image captured at a second time, and to generate a first pyramid image from the first input image and a second pyramid image from the second input image; a subject extractor configured to detect subjects in the first pyramid image and the second pyramid image and to generate a plurality of pieces of subject data representing the subjects; and a buffer storing the plurality of pieces of subject data representing the subjects detected in the first input image and the second input image, wherein the subjects include a first subject and a second subject, the second subject being relatively close to the subject detection system compared to the first subject, wherein the pyramid image generator is further configured to: generate the first pyramid image having a first resolution by adding first time information corresponding to the first time to the first input image; generate a third pyramid image having a second resolution lower than the first resolution by downsampling the first pyramid image; and generate the second pyramid image having the first resolution by adding second time information corresponding to the second time to the second input image, and wherein the subject extractor is further configured to: generate first subject data of the first subject among the plurality of pieces of subject data by extracting the first subject from the first pyramid image using a deep learning model trained based on a neural network; generate second subject data of the first subject among the plurality of pieces of subject data by extracting the first subject from the second pyramid image using the deep learning model; and generate third subject data of the second subject among the plurality of pieces of subject data by extracting the second subject from the third pyramid image using the deep learning model.

2. The object detection system of claim 1, wherein, the pyramid image generator is further configured to: generate a fourth pyramid image having a third resolution lower than the second resolution by downsampling the third pyramid image; and generate a fifth pyramid image having the second resolution by downsampling the second pyramid image.

3. The object detection system of claim 2, wherein, the subjects further include a third subject, the third subject being relatively close to the subject detection system compared to the second subject, and wherein the subject extractor is further configured to: generate fourth subject data of the third subject among the plurality of pieces of subject data by extracting the third subject from the fourth pyramid image using the deep learning model trained based on the neural network; and generate fifth subject data of the second subject among the plurality of pieces of subject data by extracting the second subject from the fifth pyramid image using the deep learning model.

4. The object detection system of claim 1, wherein, the subject extractor is further configured to: store the first subject data of the first subject and the second subject data of the first subject in a first area of the buffer; and store the third subject data of the second subject in a second area of the buffer. 5.The subject detection system according to claim 4, further comprising: ​ The object tracker is configured to track the first object by using at least one object data among a plurality of object data selected from the first object data and the second object data stored in the first area and the first time information, and to track the second object by using at least one object data among a plurality of object data selected from the third object data stored in the second area and the second time information.

6. The object detection system of claim 1, wherein, The pyramid image generator is further configured to: generate a first pyramid image set including the first pyramid image and the second pyramid image; and generate a second pyramid image set by downsampling at least one of the first pyramid image and the second pyramid image in the first pyramid image set, and wherein the object extractor is further configured to: generate N pieces of first object data respectively corresponding to N time points among the plurality of object data from the first pyramid image set, wherein N is a natural number; and generate M pieces of second object data respectively corresponding to M time points among the plurality of object data from the second pyramid image set, wherein M is a natural number.

7. The object detection system of claim 6, wherein, N is greater than M.

8. The object detection system of claim 6, wherein, The objects include the first object and the second object, and wherein the object detection system further includes an object tracker configured to: track the first object by using P pieces of first object data among the N pieces of first object data, wherein P is a natural number less than or equal to N, and track the second object by using Q pieces of second object data among the M pieces of second object data, wherein Q is a natural number less than or equal to M. 9.The object detection system of claim 1, further comprising: a region of interest manager configured to set a first region of interest for the first input image and a second region of interest for the second input image, wherein the object extractor is further configured to: extract objects from a first area of the first input image and a second area of the second input image, the first area and the second area corresponding to the first region of interest and the second region of interest. 10.A method of detecting objects, the method comprising: receiving a first input image captured at a first time and a second input image captured at a second time; generating a first pyramid image associated with the first time from the first input image and a second pyramid image associated with the second time from the second input image; generating a plurality of object data representing objects detected in the first input image and the second input image based on the first pyramid image and the second pyramid image; and storing the plurality of object data in a buffer, wherein the objects include a first object and a second object, the second object being relatively close to the object detection system compared to the first object, wherein the generating the first pyramid image and the second pyramid image comprises: generating the first pyramid image having a first resolution by adding first time information corresponding to the first time to the first input image; generating a third pyramid image having a second resolution lower than the first resolution by downsampling the first pyramid image; and generating a second pyramid image having a third resolution lower than the second resolution by downsampling the second pyramid image. generating a second pyramid image having a first resolution by adding second time information corresponding to a second time to the second input image, and wherein the generating the plurality of pieces of object data includes: generating first object data of a first object among the plurality of pieces of object data by extracting the first object from the first pyramid image using a deep learning model trained based on a neural network; generating second object data of the first object among the plurality of pieces of object data by extracting the first object from the second pyramid image using the deep learning model; and generating third object data of a second object among the plurality of pieces of object data by extracting the second object from the third pyramid image using the deep learning model.

11. The method of claim 10, wherein, The generating the first pyramid image and the second pyramid image further includes: generating a fourth pyramid image having a third resolution by down-sampling the third pyramid image, the third resolution being lower than the second resolution; and generating a fifth pyramid image having the second resolution by down-sampling the second pyramid image.

12. The method of claim 11, wherein, The object further includes a third object, the third object being relatively close to the object detection system compared to the second object, and wherein the generating the plurality of pieces of object data further includes: generating fourth object data of the third object among the plurality of pieces of object data by extracting the third object from the fourth pyramid image using the deep learning model trained based on the neural network; and generating fifth object data of the second object among the plurality of pieces of object data by extracting the second object from the fifth pyramid image using the deep learning model.

13. The method of claim 11, wherein, The storing includes: storing the first object data of the first object and the second object data of the first object in a first area of the buffer; and storing the third object data of the second object in a second area of the buffer.

14. The method of claim 13, further comprising: tracking the first object by using at least one piece of object data among the plurality of pieces of object data selected from the first object data and the second object data stored in the first area and the first time information and the second time information; and tracking the second object by using at least one piece of object data among the plurality of pieces of object data selected from the third object data stored in the second area. The first object is relatively far from an image capturing device that captures the first input image and the second input image, and 15. The method of claim 10, wherein, The second object is relatively close to the image capturing device compared to the first object. The generating the first pyramid image and the second pyramid image includes:

16. The method of claim 10, wherein, generating a first pyramid image set including the first pyramid image and the second pyramid image; and generating a second pyramid image set by down-sampling at least one of the first pyramid image and the second pyramid image in the first pyramid image set, and wherein the generating the plurality of pieces of object data includes: generating N pieces of first object data among the plurality of pieces of object data respectively corresponding to N time points from the first pyramid image set, where N is a natural number; and generating N pieces of second object data among the plurality of pieces of object data respectively corresponding to the N time points from the second pyramid image set. M pieces of second object data respectively corresponding to M time points from among a plurality of pieces of object data generated from a second pyramid image set, where M is a natural number.

17. The method of claim 16, wherein, The objects include a first object and a second object, and wherein the method further comprises: tracking the first object by using P pieces of first object data from among the N pieces of first object data, where P is a natural number less than or equal to N; and tracking the second object by using Q pieces of second object data from among the M pieces of second object data, where Q is a natural number less than or equal to M.

18. A driving assistance system for driving a vehicle by detecting an object, the driving assistance system comprising: a pyramid image generator configured to receive a first input image captured at a first time and a second input image captured at a second time, and to generate a first pyramid image from the first input image and a second pyramid image from the second input image; an object extractor configured to detect an object in the first pyramid image and the second pyramid image and to generate a plurality of pieces of object data representing the object by using deep learning based on a neural network; a buffer storing the plurality of pieces of object data representing the object detected in the first input image and the second input image; and an object tracker configured to track the object based on the plurality of pieces of object data stored in the buffer, wherein the object includes a first object and a second object, a position of the second object is relatively close to the object detection system compared to the first object, wherein the pyramid image generator is further configured to: generate the first pyramid image having a first resolution by adding first time information corresponding to the first time to the first input image; generate a third pyramid image having a second resolution lower than the first resolution by downsampling the first pyramid image; and generate the second pyramid image having the first resolution by adding second time information corresponding to the second time to the second input image, and wherein the object extractor is further configured to: generate first object data of the first object from among the plurality of pieces of object data by extracting the first object from the first pyramid image using a deep learning model trained based on a neural network; generate second object data of the first object from among the plurality of pieces of object data by extracting the first object from the second pyramid image using the deep learning model; and generate third object data of the second object from among the plurality of pieces of object data by extracting the second object from the third pyramid image using the deep learning model. the pyramid image generator is further configured to:

19. The driver assist system of claim 18, wherein, generate a first pyramid image set including the first pyramid image and the second pyramid image; and generate a second pyramid image set by downsampling at least one of the first pyramid image and the second pyramid image, and wherein the object extractor is further configured to: generate N pieces of first object data respectively corresponding to N time points from among the plurality of pieces of object data from the first pyramid image set, where N is a natural number; and generate M pieces of second object data respectively corresponding to M time points from among a plurality of pieces of object data generated from a second pyramid image set, where M is a natural number. M pieces of second object data respectively corresponding to M time points from among a plurality of pieces of object data generated from a second pyramid image set, where M is a natural number. The objects include a first object and a second object, and wherein the method further comprises: tracking the first object by using P pieces of first object data from among the N pieces of first object data, where P is a natural number less than or equal to N; and tracking the second object by using Q pieces of second object data from among the M pieces of second object data, where Q is a natural number less than or equal to M.

18. A driving assistance system for driving a vehicle by detecting an object, the driving assistance system comprising: a pyramid image generator configured to receive a first input image captured at a first time and a second input image captured at a second time, and to generate a first pyramid image from the first input image and a second pyramid image from the second input image; an object extractor configured to detect an object in the first pyramid image and the second pyramid image and to generate a plurality of pieces of object data representing the object by using deep learning based on a neural network; a buffer storing the plurality of pieces of object data representing the object detected in the first input image and the second input image; and an object tracker configured to track the object based on the plurality of pieces of object data stored in the buffer, wherein the object includes a first object and a second object, a position of the second object is relatively close to the object detection system compared to the first object, wherein the pyramid image generator is further configured to: generate the first pyramid image having a first resolution by adding first time information corresponding to the first time to the first input image; generate a third pyramid image having a second resolution lower than the first resolution by downsampling the first pyramid image; and generate the second pyramid image having the first resolution by adding second time information corresponding to the second time to the second input image, and wherein the object extractor is further configured to: generate first object data of the first object from among the plurality of pieces of object data by extracting the first object from the first pyramid image using a deep learning model trained based on a neural network; generate second object data of the first object from among the plurality of pieces of object data by extracting the first object from the second pyramid image using the deep learning model; and generate third object data of the second object from among the plurality of pieces of object data by extracting the second object from the third pyramid image using the deep learning model. the pyramid image generator is further configured to: generate a first pyramid image set including the first pyramid image and the second pyramid image; and generate a second pyramid image set by downsampling at least one of the first pyramid image and the second pyramid image, and wherein the object extractor is further configured to: generate N pieces of first object data respectively corresponding to N time points from among the plurality of pieces of object data from the first pyramid image set, where N is a natural number; and M pieces of second object data respectively corresponding to M time points are generated from the second pyramid image set, wherein M is a natural number.

Citation Information

Patent Citations

  • Boilers and boiler systems and methods of operating boilers

    KR1020190104574A

  • Subcategory-aware convolutional neural networks for object detection

    US20170124415A1

  • Dynamic method for recognizing objects and image processing system therefor

    US5063603A