Information processing apparatus, sensing device, moving body, and information processing method

By mapping the coordinates of the detection object to the virtual space and using the Kalman filter to track the particle position and size, the processing load problem of high-precision tracking of the detection object position and size in the vehicle is solved, and the tracking effect of high-precision and low-load is achieved.

CN114868150BActive Publication Date: 2025-07-22KYOCERA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080089795.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-23
Filing Date
2020-12-22
Publication Date
2025-07-22
Estimated Expiration
2040-12-22

AI Technical Summary

Technical Problem

In a vehicle, as the relative position of the approaching vehicle and pedestrian changes, it is difficult for the prior art to reduce the processing load while tracking its position with high accuracy, and it is easy to lead to reduced tracking errors and accuracy.

Method used

By transforming the coordinate mapping of the detection object into the virtual space, the Kalman filter is used to track the position and velocity of the particles, and estimating the current size in combination with the current and past size observations, reducing the processing load.

Benefits of technology

It realizes the high-precision tracking of the position and size of the detection object while reducing processing load, improving tracking accuracy and reducing errors, and providing an easy-to-use image display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114868150B_ABST
    Figure CN114868150B_ABST
Patent Text Reader

Abstract

The information processing device includes: an input interface, a processor, and an output interface. The input interface acquires observation data obtained from an observation space. The processor detects a detection target included in the observation data. The processor maps and transforms the coordinates of the detected detection target into the coordinates of the detection target in the virtual space, tracks the position and velocity of the particle representing the detection target in the virtual space, and maps and transforms the coordinates of the tracked particle in the virtual space into the coordinates in the display space. The processor sequentially observes the size of the detection target in the display space, and estimates the size of the detection target based on the observed value of the size of the detection target at the current time point and the estimated value of the size of the detection target in the past. The output interface outputs output information based on the coordinates of the particle mapped and transformed into the display space and the estimated size of the detection target.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the priority of Japanese Patent Application No. 2019 - 231662 filed in Japan on December 23, 2019, and incorporates the entire disclosure of the prior application herein by reference. Technical field

[0003] The present invention relates to an information processing apparatus, a sensing apparatus, a moving body, and an information processing method. Background art

[0004] Conventionally, an image processing apparatus has been disclosed that processes an image signal output from a camera provided on a vehicle to acquire an image around the vehicle, detects approaching vehicles, pedestrians, etc., and displays a rectangular frame around the approaching vehicles and pedestrians in the image (for example, refer to Patent Document 1).

[0005] Prior art documents

[0006] Patent documents

[0007] Patent Document 1: Japanese Unexamined Patent Application Publication No. 11 - 321494 Summary of the invention

[0008] The information processing apparatus of the present invention includes: an input interface, a processor, and an output interface. The input interface is configured to acquire observation data obtained from an observation space. The processor is configured to detect a detection target included in the observation data. The processor is configured to map - transform the coordinates of the detected detection target into the coordinates of the detection target in a virtual space, track the position and velocity of a mass point representing the detection target on the virtual space, and map - transform the coordinates of the tracked mass point on the virtual space into the coordinates in a display space. The processor is configured to sequentially observe the size of the detection target in the display space, and estimate the size of the detection target at the current time point based on the observed value of the size of the detection target at the current time point and the estimated value of the size of the detection target in the past. The output interface is configured to output output information based on the coordinates of the mass point mapped - transformed into the display space and the estimated size of the detection target.

[0009] The sensing device of the present invention includes: a sensor, a processor, and an output interface. The sensor is configured to sense an observation space and obtain observation data of a detection object. The processor is configured to detect the detection object included in the observation data. The processor is configured to map-transform the coordinates of the detected detection object into the coordinates of the detection object in a virtual space, track the position and velocity of a particle representing the detection object in the virtual space, and map-transform the coordinates of the tracked particle in the virtual space into the coordinates in the display space. The processor is configured to sequentially observe the size of the detection object in the display space, and estimate the size of the detection object at the current time point based on the observed value of the size of the detection object at the current time point and the estimated value of the size of the detection object in the past. The output interface is configured to output output information based on the coordinates of the particle map-transformed into the display space and the estimated size of the detection object.

[0010] The mobile body of the present invention includes a sensing device. The sensing device includes: a sensor, a processor, and an output interface. The sensor is configured to sense an observation space and obtain observation data of a detection object. The processor is configured to detect the detection object included in the observation data. The processor is configured to map-transform the coordinates of the detected detection object into the coordinates of the detection object in a virtual space, track the position and velocity of a particle representing the detection object in the virtual space, and map-transform the coordinates of the tracked particle in the virtual space into the coordinates in the display space. The processor is configured to sequentially observe the size of the detection object in the display space, and estimate the size of the detection object at the current time point based on the observed value of the size of the detection object at the current time point and the estimated value of the size of the detection object in the past. The output interface is configured to output output information based on the coordinates of the particle map-transformed into the display space and the estimated size of the detection object.

[0011] The information processing method of the present invention includes the steps of: acquiring observation data from an observation space and detecting a detection object included in the observation data. The information processing method includes the steps of: mapping and transforming the coordinates of the detected detection object into the coordinates of the detection object in a virtual space, tracking the position and velocity of a mass point representing the detection object on the virtual space, and mapping and transforming the coordinates of the tracked mass point on the virtual space into the coordinates on the display space. The information processing method includes the steps of: sequentially observing the size of the detection object on the display space, and estimating the size of the detection object at the current time point based on the observed value of the size of the detection object at the current time point and the estimated value of the size of the detection object in the past. The information processing method includes the steps of: outputting output information based on the coordinates of the mass point mapped and transformed onto the display space and the estimated size of the detection object. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a block diagram showing a schematic configuration of an image processing system including an image processing apparatus as an information processing apparatus according to an embodiment.

[0013] Figure 2 is a diagram showing a vehicle and a pedestrian of an image processing system equipped with Figure 1 of the image processing system.

[0014] Figure 3 is a flowchart showing an example of processing for tracking a subject image on a moving image.

[0015] Figure 4 is a diagram showing an example of a subject image on a moving image.

[0016] Figure 5 is a diagram for explaining the relationship between a subject in the actual space, a subject image in a moving image, and a mass point in a virtual space.

[0017] Figure 6 is a diagram showing an example of the movement of a mass point in a virtual space.

[0018] Figure 7 is a diagram for explaining a method of tracking the size of a subject image in a moving image.

[0019] Figure 8 is a diagram showing an example of estimating the size of a subject image.

[0020] Figure 9 is an example of an image for displaying an image element (bounding box) on a moving image.

[0021] Figure 10It is a block diagram showing a schematic structure of a photographing device as a sensing device representing an embodiment.

[0022] Figure 11 It is a block diagram showing an example of a schematic structure of a sensing device including a millimeter-wave radar.

[0023] Figure 12 It represents Figure 11 A flowchart showing an example of processing executed by the information processing unit of the sensing device.

[0024] Figure 13 It is a diagram showing an example of observation data mapped and transformed onto a virtual space.

[0025] Figure 14 It is to Figure 13 A diagram for clustering the observation data. Detailed Embodiment

[0026] In an information processing device mounted on a vehicle or the like, as the relative positions of approaching vehicles, pedestrians, etc. with respect to the own vehicle change, the positions and sizes of the images of the vehicles, pedestrians, etc. in the display space change moment by moment. Therefore, accurately tracking the positions of approaching vehicles, pedestrians, etc. while grasping the sizes of detection objects increases the processing load and may lead to a decrease in tracking error and / or accuracy.

[0027] Preferably, the information processing device can reduce the processing load while accurately tracking the detection object.

[0028] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The drawings used in the following description are schematic diagrams. The dimensional ratios, etc. on the drawings are not necessarily consistent with the actual situation.

[0029] An example of an information processing device according to an embodiment of the present invention, i.e., an image processing device 20, is included in an image processing system 1. The image processing system 1 includes: a photographing device 10, an image processing device 20, and a display 30. The photographing device 10 is an example of a sensor for sensing an observation space. As Figure 2 Illustrated, the image processing system 1 is mounted on a vehicle 100 as an example of a moving body.

[0030] As Figure 2As shown, in the present embodiment, in the coordinates of the actual space, the x-axis direction is the width direction of the vehicle 100 on which the imaging device 10 is provided. The actual space is the observation space, which is the object for obtaining the observation data. The y-axis direction is the backward direction of the vehicle 100. The x-axis direction and the y-axis direction are parallel to the road surface on which the vehicle 100 is located. The z-axis direction is the direction perpendicular to the road surface. The z-axis direction can be referred to as the vertical direction. The x-axis direction, the y-axis direction, and the z-axis direction are orthogonal to each other. The method of defining the x-axis direction, the y-axis direction, and the z-axis direction is not limited to this. The x-axis direction, the y-axis direction, and the z-axis direction can be replaced with each other.

[0031] (Imaging device)

[0032] The imaging device 10 is configured to include: an imaging optical system 11, an imaging element 12, and a processor 13.

[0033] The imaging device 10 can be provided at various positions of the vehicle 100. The imaging device 10 includes a front camera, a left camera, a right camera, a rear camera, etc., but is not limited to these. The front camera, the left camera, the right camera, and the rear camera are respectively provided on the vehicle 100 so as to be able to image the surrounding areas in front of, to the left of, to the right of, and behind the vehicle 100. In the embodiment described as an example below, as Figure 2 shown, the imaging device 10 is mounted on the vehicle 100 with the optical axis direction facing downward from the horizontal direction so as to image the rear of the vehicle 100.

[0034] As Figure 1 shown, the imaging device 10 includes: an imaging optical system 11, an imaging element 12, and a processor 13. The imaging optical system 11 is configured to include one or more lenses. The imaging element 12 includes: a CCD image sensor (Charge-Coupled Device Image Sensor) and a CMOS image sensor (Complementary Metal Oxide Semiconductor Image Sensor). The imaging element 12 converts the subject image formed on the imaging surface of the imaging element 12 by the imaging optical system 11 into an electrical signal. The subject image is the image of the subject to be detected. The imaging element 12 can image a moving image at a specified frame rate. The moving image is an example of the observation data. Each still image constituting the moving image is called a frame. The number of images that can be taken in one second is called the frame rate. The frame rate can be set to, for example, 60 fps (frames per second), 30 fps, etc.

[0035] The processor 13 controls the entire imaging device 10 and performs various image processes on the moving image output from the imaging element 12. The image processes performed by the processor 13 may include any process such as distortion correction, brightness adjustment, contrast adjustment, and gamma correction.

[0036] The processor 13 may be composed of one or more processors. The processor 13 includes, for example, one or more circuits or units configured to perform one or more data calculation steps or processes by executing instructions stored in an associated memory. The processor 13 includes one or more processors, microprocessors, microcontrollers, application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), or any combination of these devices or structures, or a combination of other known devices or structures.

[0037] (Image processing device)

[0038] The image processing device 20 can be installed at any position of the vehicle 100. The image processing device 20 is configured to include: an input interface 21, a storage unit 22, a processor 23, and an output interface 24.

[0039] The input interface 21 is configured to be able to communicate with the imaging device 10 through a wired or wireless communication unit. The input interface 21 acquires the moving image from the imaging device 10. The input interface 21 may correspond to the transmission method of the image signal transmitted by the imaging device 10. The input interface 21 can be referred to as an input unit or an acquisition unit. The imaging device 10 and the input interface 21 may be connected through a vehicle-mounted communication network such as a CAN (Controller Area Network).

[0040] The storage unit 22 is a storage device that stores data and programs required for the processor 23 to perform processing. For example, the storage unit 22 temporarily stores moving images acquired from the imaging device 10. For example, the storage unit 22 sequentially stores data generated by the processing performed by the processor 23. The storage unit 22 can be configured by using any one or more of, for example, a semiconductor memory, a magnetic memory, and an optical memory. The semiconductor memory can include a volatile memory and a non-volatile memory. The magnetic memory can include, for example, a hard disk and a magnetic tape. The optical memory can include, for example, a CD (Compact Disc), a DVD (Digital Versatile Disc), and a BD (Blu-ray (registered trademark) Disc).

[0041] The processor 23 controls the entire image processing device 20. The processor 23 identifies a subject image included in the moving image acquired via the input interface 21. The processor 23 maps and transforms the coordinates of the identified subject image into the coordinates of the subject 40 in the virtual space, and tracks the position and velocity of the mass point representing the subject 40 in the virtual space. A mass point is a point with mass and no size. The virtual space is a virtual space used in an arithmetic device such as the processor 23 to describe the motion of an object. In the present embodiment, the virtual space is a two-dimensional space in a coordinate system composed of the x-axis, y-axis, and z-axis of the actual space, where the value in the z-axis direction is set to a prescribed fixed value. The processor 23 maps and transforms the coordinates of the tracked mass point in the virtual space into the coordinates in the image space for displaying the moving image. The image space is an example of a display space. The display space is a space in which a detection object is two-dimensionally represented for visual recognition by a user or for use by another device. In addition, the processor 23 sequentially observes the size of the subject image in the image space, and estimates the size of the subject image 42 at the current time point based on the observed value of the size of the subject image at the current time point and the estimated value of the size of the subject image in the past. The processing performed by the processor 23 will be described in detail later. The processor 23, like the processor 13 of the imaging device 10, can include a plurality of processors. In addition, like the processor 13, the processor 23 can also be configured by combining multiple types of devices.

[0042] The output interface 24 is configured to output an output signal as output information from the image processing device 20. The output interface 24 can be referred to as an output unit. The output interface 24 can output an output signal based on the coordinates of the mass points mapped and transformed onto the image space of the displayed moving image and the size of the subject image within the estimated image space. For example, the output interface 24 can overlap an image element representing the size of the subject image on the image output from the imaging device 10 and output it to the display 30. The image element representing the size of the subject image is, for example, a bounding box. The bounding box is a rectangular frame line that encloses the subject image. The output interface 24 can directly output the coordinates of the mass points and the size of the subject image as an output signal.

[0043] The output interface 24 can be configured to include a physical connector and a wireless communication device. In one of the multiple embodiments, the output interface 24 is connected to the network of the vehicle 100 such as CAN. The output interface 24 can be connected to the display 30, the control device of the vehicle 100, the alarm device, etc. via a communication network such as CAN. The information output from the output interface 24 can be appropriately utilized in each of the display 30, the control device, and the alarm device.

[0044] The display 30 can display the moving image output from the image processing device 20. The display 30 can also have the following function: when receiving the coordinates of the mass points representing the position of the subject image and the information on the size of the subject image from the image processing device 20, generating an image element according to this information and overlapping it with the moving image. The display 30 can adopt various devices. For example, the display 30 can adopt a liquid crystal display (LCD: Liquid Crystal Display), an organic EL (Electro-Luminescence) display, an inorganic EL display, a plasma display panel (PDP: Plasma Display Panel), a field emission display (FED: Field Emission Display), an electrophoretic display, a rolling ball display, etc.

[0045] (Tracking process of the subject image)

[0046] Next, refer to Figure 3The flowchart details the image processing method executed by the image processing apparatus 20. The image processing method is an example of an information processing method. The image processing apparatus 20 can be configured to read a program recorded in a non-transitory computer-readable medium and execute the processing performed by the processor 23 described below. Non-transitory computer-readable media include, but are not limited to, magnetic storage media, optical storage media, magneto-optical storage media, and semiconductor storage media. Magnetic storage media include magnetic disks, hard disks, and magnetic tapes. Optical storage media include optical discs such as CDs (Compact Discs), DVDs, and Blu-ray (registered trademark) Discs. Semiconductor storage media include ROMs (Read Only Memories), EEPROMs (Electrically Erasable Programmable Read-Only Memories), and flash memories.

[0047] Figure 3 The flowchart is to obtain consecutive frames of a moving image and perform the processing executed by the processor 23. The processor 23 of the image processing apparatus 20 follows Figure 3 The flowchart to track the position and size of the subject image 42 whenever a frame of the moving image is obtained. In the following description, as Figure 2 shown, it is assumed that the imaging device 10 provided at the rear of the vehicle 100 captures a pedestrian as the subject 40. The subject 40 is not limited to a pedestrian and can include various objects such as vehicles traveling on the road and obstacles on the road.

[0048] The processor 23 obtains each frame of the moving image from the imaging device 10 via the input interface 21 (step S101). An example of one frame of the moving image is shown in Figure 4 . In the example of Figure 4 , a subject image 42 of a pedestrian, who is the subject 40 trying to cross behind the vehicle 100, is displayed in the two-dimensional image space 41 constituted by the uv coordinate system. The u coordinate is the horizontal coordinate of the image. The v coordinate is the vertical coordinate of the image. In Figure 4 , the origin of the uv coordinates is the point at the upper left end of the image space 41. Also, the direction from left to right is the positive direction of the u coordinate. The direction from top to bottom is the positive direction of the v coordinate.

[0049] The processor 23 identifies the subject image 42 from each frame of the moving image through image recognition (step S102). Thereby, the processor 23 detects the subject 40. The method for identifying the subject image 42 includes various known methods. For example, the method for identifying the subject image 42 includes: a method based on the shape recognition of objects such as vehicles and pedestrians, a method based on template matching, a method of calculating feature amounts from the image and using them for matching, etc. In the calculation of the feature amounts, a function approximator that can learn the relationship between input and output can be used. A function approximator that can learn the relationship between input and output can use a neural network.

[0050] The processor 23 maps and transforms the coordinates (u, v) of the subject image 42 in the image space 41 into the coordinates (x', y') of the subject 40 in the virtual space (step S103). Generally, the coordinates (u, v) of the two-dimensional image space 41 cannot be transformed into the coordinates (x, y, z) of the actual space. However, by determining the height in the actual space and fixing the z coordinate to a specified value, the coordinates (u, v) of the image space 41 can be mapped into the two-dimensional virtual space coordinates (x', y') corresponding to the coordinates (x, y, z0) (z0 is a fixed value) of the actual space. Hereinafter, with reference to Figure 4 and Figure 5 it will be described.

[0051] In Figure 4 a representative point 43 located at the center of the lowermost part of the subject image 42 is determined. For example, the representative point 43 can be the position at the lowest v coordinate of the area occupied by the subject image 42 in the image space 41 and the central position within the range of the u coordinate. It is assumed that this representative point 43 is the position where the subject 40 corresponding to the subject image 42 contacts the road surface or the ground.

[0052] In Figure 5In [the figure], the relationship between the subject 40 located in the three-dimensional real space and the subject image 42 in the two-dimensional image space 41 is shown. When the internal parameters of the imaging device 10 are known, based on the coordinates (u, v) in the image space 41, the direction from the center of the imaging optical system 11 of the imaging device 10 to the corresponding coordinates (x, y, z) in the real space can be calculated. The internal parameters of the imaging device 10 include information such as the focal length of the imaging optical system 11, distortion, and the pixel size of the imaging element 12. In the real space, the point where the straight line facing the direction corresponding to the representative point 43 in the image space 41 intersects the reference plane 44 with z = 0 is defined as the mass point 45 of the subject 40. The reference plane 44 corresponds to the road surface or ground where the vehicle 100 is located. The mass point 45 has three-dimensional coordinates (x, y, 0). Therefore, when the two-dimensional space with z = 0 is regarded as the virtual space, the coordinates of the mass point 45 can be represented by (x', y'). The coordinates (x', y') of the mass point 45 on the virtual space correspond to the coordinates (x, y) of a specific point of the subject 40 on the xy plane (z = 0) when observing the subject 40 from the direction along the z-axis in the real space. The specific point is the point corresponding to the mass point 45.

[0053] As Figure 6 shown, the processor 23 tracks the position (x', y') and velocity (v x' , v y' ) of the mass point 45 mapped and transformed from the representative point 43 of the subject image 42 onto the virtual space 46 (step S104). Since the mass point 45 has information on the position (x', y') and velocity (v x' , v y' ), the processor 23 can predict the range of the position (x', y') of the mass point 45 in consecutive frames. The processor 23 can identify the mass point 45 located within the range predicted in the next frame as the mass point 45 corresponding to the tracked subject image 42. Each time the processor 23 receives the input of a new frame, it sequentially updates the position (x', y') and velocity (v x' , v y' ) of the mass point 45.

[0054] For example, the estimation using a Kalman filter based on a state space model can be used for tracking the mass point 45. By performing prediction / estimation using the Kalman filter, the robustness against non-detection, misdetection, etc. of the subject 40 to be tracked is improved. Generally, it is difficult to describe the subject image 42 in the image space 41 by an appropriate model for describing motion. Therefore, it is difficult to simply perform highly accurate position estimation on the subject image 42 in the image space 41. In the image processing apparatus 20 of the present embodiment, by mapping and transforming the subject image 42 into the mass point 45 in the real space, a model for describing motion in the real space can be applied, and thus the accuracy of tracking the subject image 42 is improved. In addition, by treating the subject 40 as a mass point 45 having no size, simple and easy tracking can be achieved.

[0055] Each time the processor 23 estimates a new position of the mass point 45, it maps and transforms the coordinates on the virtual space 46 of the mass point 45 into the coordinates (u, v) on the image space 41 (step S105). The mass point 45 located at the coordinates (x', y') on the virtual space 46 can be mapped and transformed onto the image space 41 as a point located at the coordinates (x', y', 0) in the real space. The coordinates (x', y', 0) in the real space can be mapped into the coordinates (u, v) on the image space 41 of the imaging device 10 by a known method.

[0056] The processor 23 can execute the processes of step S106 and step S107 described below in parallel with the processes of step S103 to step S105. For either the processes of step S103 to step S105 or the processes of step S106 to step S107, the processor 23 can execute them before or after the other process.

[0057] The processor 23 observes the size of the subject image 42 on the image space of the moving image identified in step S102 (step S106). The size of the subject image 42 includes the width and height occupied by the subject image 42 in the image space. The size of the subject image 42 can be expressed in pixels, for example.

[0058] The processor 23 estimates the size of the subject image 42 at the current time point based on the observed value of the size of the subject image 42 at the current time point and the estimated value of the size of the past subject image 42 (step S107). Herein, the "subject image at the current time point" refers to the subject image based on the image of the frame most recently acquired from the imaging device 10. The "first previous subject image" refers to the subject image based on the image of the previous frame of the frame most recently acquired from the imaging device 10. The processor 23 continuously observes the size of the subject image 42 for each frame image acquired as a moving image. In the present application, the observation at the current time point may be referred to as the "current observation", and the observation made the first previous one from the current observation may be referred to as the "previous observation". In the present application, "the current time point" and "the current", "the first previous one" and "the previous" are used with almost the same meaning.

[0059] As Figure 7 shown, the processor 23 performs tracking processing based on the estimated value W(k - 1) of the previous width deduced from the result of the previous observation and the observed value Wmeans(k) of the current width obtained from the result of the current observation, and calculates the estimated value W(k) of the current width. The processor 23 performs tracking processing based on the estimated value H(k - 1) of the previous height deduced from the result of the previous observation and the observed value Hmeans(k) of the current height obtained from the result of the current observation, and calculates the estimated value H(k) of the current height. Herein, k corresponds to the serial number of the frame included in the moving image. The current observation targets the k-th frame. The estimation of the width and height of the subject image 42 can be performed based on the following mathematical formulas (1) and (2).

[0060] W(k) = W(k - 1) + α(Wmeans(k) - W(k - 1)) (1)

[0061] H(k) = H(k - 1) + α(Hmeans(k) - H(k - 1)) (2)

[0062] The parameter α is a parameter included in the range of 0 ≤ α ≤ 1. The parameter α is a parameter set according to the reliability of the observed values Wmeans(k) and Hmeans(k) for the width and height. When α = 0, the estimated values W(k) and H(k) of the width and height this time are the same as the estimated values W(k - 1) and H(k - 1) of the width and height last time respectively. When α = 0.5, the estimated values W(k) and H(k) of the width and height this time are the average values of the estimated values W(k - 1) and H(k - 1) of the width and height last time and the observed values Wmeans(k) and Hmeans(k) of the width and height this time respectively. When α = 1, the estimated values W(k) and H(k) of the width and height this time are the observed values Wmeans(k) and Hmeans(k) of the width and height this time respectively.

[0063] The processor 23 can dynamically adjust the parameter α during the tracking. For example, the processor 23 can estimate the recognition accuracy of the subject image 42 included in the moving image and dynamically adjust the parameter α based on the estimated recognition accuracy. For example, the processor 23 can calculate values such as the brightness and contrast of the image based on the moving image, and when the image is darker or the contrast is lower, it can be determined that the recognition accuracy of the subject image 42 is lower, and the parameter α is decreased. The processor 23 can adjust the parameter α according to the speed of movement of the subject image 42 within the moving image. For example, when the subject image 42 within the moving image moves faster, in order to follow the movement of the subject image 42, the processor 23 can set the parameter α to a larger value than when the subject image 42 moves slower.

[0064] The processor 23 sometimes cannot detect the observed value of the size of the subject image 42 at the current time point from the current frame in step S102. For example, when two subject images 42 overlap in the image space where the moving image is displayed, sometimes the sizes of their respective subject images 42 cannot be detected. In such a case, the processor 23 can estimate the size of the subject image 42 at the current time point only based on the estimated value of the size of the past subject image 42. For example, the processor 23 can use the estimated values W(k - 1) and H(k - 1) of the width and height of the first one before as the estimated values W(k) and H(k) of the width and height at the current time point.

[0065] When the processor 23 cannot detect the observed value of the size of the subject image 42 at the current time point, it can consider the estimated values W(k−j), H(k−j) (j≥2) of the width and height of the subject image 42 included in two or more previous frames. For example, the processor 23 can use a parameter β within the range of 0≤β≤1 to estimate the estimated values W(k) and H(k) of the width and height of the subject image 42 at the current time point through the following mathematical formulas (3) and (4).

[0066] W(k) = W(k−1)+β(W(k−1)−W(k−2)) (3)

[0067] H(k) = H(k−1)+β(H(k−1)−H(k−2)) (4)

[0068] Thus, when the size of the subject image 42 at the current time point cannot be obtained, the size of the subject image 42 at the current time point can be estimated by reflecting the estimated values of the size of the subject images 42 in the previous two frames.

[0069] Use Figure 8 to illustrate an example of the process of tracking the size of the subject image 42. As the initial values (k = 0) for estimating the subject image 42, the processor 23 sets the observed values Wmeans(0), Hmeans(0) of the initial width and height as the estimated values W(0), H(0) of the width and height. In subsequent frames, the processor 23 sets the estimated values W(k−1), H(k−1) of the width and height of the previous frame as the predicted values W(k−1), H(k−1) of the width and height in this frame. The processor 23 uses the predicted values W(k−1), H(k−1) of the width and height and the observed values Wmeans(k), Hmeans(k) of the width and height in this frame to estimate the estimated values W(k), H(k) of the width and height in this frame.

[0070] In Figure 8In this case, it is assumed that the observation value cannot be obtained in the (k + 1)-th frame. In such a case, the processor 23 sets α = 0 in the mathematical expressions (1) and (2), and uses the predicted values W(k) and H(k) of the width and height of the subject image 42 in the immediately previous frame as the estimated values W(k + 1) and H(k + 1) of the width and height of the subject image 42 in the (k + 1)-th frame. Alternatively, the processor 23 may use the mathematical expressions (3) and (4) to calculate the estimated values W(k + 1) and H(k + 1) of the width and height of the subject image 42 in the (k + 1)-th frame. In this case, the estimated values W(k - 1) and H(k - 1) of the width and height in the second previous frame and the estimated values W(k) and H(k) of the width and height in the immediately previous frame are reflected in the estimated values W(k + 1) and H(k + 1).

[0071] In this way, even when the processor 23 cannot detect the observation values of the width and height of the subject image 42 from the moving image in the image space 41, it can stably calculate the estimated values W(k) and H(k) of the width and height of the subject image 42.

[0072] In step S105, when the coordinates (u, v) after the mapping transformation of the mass point 45 at the current time point to the image space 41 are obtained, and the size of the subject image 42 is estimated in step S107, the processor 23 proceeds to the process of step S108. In step S108, as Figure 9 shown, the processor 23 generates an image based on the position of the mass point coordinates in which the image element 48 representing the estimated size of the object image is superimposed on the image space for displaying the moving image. The image element 48 is, for example, a bounding box. The bounding box is a rectangular frame line surrounding the subject image 42. The processor 23 causes the moving image including the subject image 42 with the image element 48 added thereto to be displayed on the display 30 via the output interface 24. Thus, the user of the image processing system 1 can visually confirm the subject image 42 recognized by the image processing device 20 in a state where the subject image 42 is emphasized by the image element 48.

[0073] According to the present embodiment, the image processing device 20 tracks the position of the subject image 42 as the mass point 45 in the virtual space 46, tracks the size of the subject image 42 in the image space 41, and synthesizes the results thereof and displays them on the display 30. Thus, the image processing device 20 can track the position and size of the subject image 42 with high accuracy and reduce the processing load.

[0074] According to this embodiment, the image processing apparatus 20 uses a Kalman filter to track the mass point 45 corresponding to the position of the subject image 42. Therefore, even when the recognition error of the position of the subject image 42 in the image processing apparatus 20 is large, the position of the subject image 42 can be estimated with high accuracy.

[0075] According to this embodiment, the image processing apparatus 20 continuously observes the size of the subject image 42, and estimates the size of the subject image 42 at the current time point based on the observed value of the size of the subject image 42 at the current time point and the estimated value of the size of the subject image 42 in the past. Thus, even when the error of the observed size of the subject image 42 is large, the image processing apparatus 20 can estimate the size of the subject image 42 with high accuracy. In addition, the image processing apparatus 20 uses parameters α and β to reflect the estimated value of the subject image 42 in the past, and calculates the estimated value of the subject image 42 at the current time point based on this. Therefore, even when the observed values at each time point deviate due to errors, the flickering of the image element 48 displayed can be suppressed. Thus, the image processing apparatus 20 can provide an image that is easy for the user to view.

[0076] (Imaging device with tracking function)

[0077] The functions of the image processing apparatus 20 of this embodiment described in the above embodiment can be mounted on an imaging device. Figure 10 FIG. schematically shows an imaging device 50 according to an embodiment of the present invention having the functions of the image processing apparatus 20. The imaging device 50 includes: an imaging optical system 51, an imaging element 52, a storage unit 53, a processor 54, and an output interface 55. The imaging optical system 51 and the imaging element 52 are structural elements similar to Figure 1 the imaging optical system 11 and the imaging element 12 of the imaging device 10. The storage unit 53 and the output interface 55 are structural elements similar to Figure 1 the storage unit 22 and the output interface 24 of the image processing apparatus 20. The processor 54 is a structural element that combines the functions of Figure 1 the processor 13 of the imaging device 10 and the processor 23 of the image processing apparatus 20.

[0078] In the imaging device 50, the imaging element 52 captures a moving image of the subject 40 imaged by the imaging optical system 51. With respect to the moving image output by the imaging element 52, the processor 54 performs the same processing as that described in the Figure 3 flowchart. Thus, the imaging device 50 can display an image in which the image element 48 as a bounding box is added to the subject image 42 on the display 30 as shown in Figure 9 .

[0079] In the above-described embodiment, the information processing apparatus is described as the image processing apparatus 20, and the sensor is described as the imaging device 10. The sensor is not limited to an imaging device that detects visible light, and also includes a far-infrared camera that acquires an image through far-infrared rays. In addition, the information processing apparatus of the present invention is not limited to an apparatus that acquires a moving image as observation data and detects a detection target through image recognition. For example, the sensor may be a sensor other than an imaging device that can sense an observation space as an observation target to detect the direction and size of the detection target. The sensor includes, for example, a sensor that uses electromagnetic waves or ultrasonic waves. The sensor that uses electromagnetic waves includes a millimeter-wave radar and LiDAR (Laser Imaging Detection and Ranging). Therefore, the detection target is not limited to a subject captured as an image. The information processing apparatus can acquire observation data including information such as the direction and size of the detection target output from the sensor, and thereby detect the detection target. In addition, the display space is not limited to an image space that displays a moving image, and can also be a space that can two-dimensionally display the detected detection target.

[0080] (Sensing device including millimeter-wave radar)

[0081] As an example, refer to Figure 11 A sensing device 60 of one embodiment will be described. The sensing device 60 includes a millimeter-wave radar 61 as an example of a sensor, an information processing unit 62, and an output unit 63. The sensing device 60 can be mounted at various positions of a vehicle similarly to the imaging device 10.

[0082] The millimeter-wave radar 61 can use electromagnetic waves in the millimeter-wave band to detect the distance, speed, direction, etc. of the detection target. The millimeter-wave radar 61 includes a transmission signal generation unit 64, a high-frequency circuit 65, a transmission antenna 66, a reception antenna 67, and a signal processing unit 68.

[0083] The transmission signal generation unit 64 generates a chirp signal after frequency modulation. The chirp signal is a signal whose frequency rises or falls at a certain time interval. The transmission signal generation unit 64 is installed in, for example, a DSP (Digital Signal Processor). The transmission signal generation unit 64 can be controlled by the information processing unit 62.

[0084] After the chirp signal is subjected to D / A conversion, it undergoes frequency conversion in the high-frequency circuit 65 to become a high-frequency signal. The high-frequency circuit 65 transmits the high-frequency signal as radio waves into the observation space through the transmitting antenna 66. The high-frequency circuit 65 can receive, as a received signal, the reflected wave obtained by reflecting the radio waves transmitted from the transmitting antenna 66 by the detection object through the receiving antenna 67. The millimeter-wave radar 61 may have a plurality of receiving antennas 67. The millimeter-wave radar 61 can estimate the direction of the detection object by detecting the phase difference between the receiving antennas in the signal processing unit 68. The method of azimuth detection in the millimeter-wave radar 61 is not limited to the method using the phase difference. The millimeter-wave radar 61 can also detect the azimuth of the detection object by scanning with a millimeter-wave band beam.

[0085] The high-frequency circuit 65 amplifies the received signal, mixes it with the transmitted signal, and converts it into a beat signal representing the frequency difference. The beat signal is converted into a digital signal and output to the signal processing unit 68. The signal processing unit 68 processes the received signal and performs estimation processing such as distance, speed, and direction. The methods for estimating distance, speed, and direction in the millimeter-wave radar 61 are well known, so the content of the processing performed by the signal processing unit 68 is omitted. The signal processing unit 68 is installed in a DSP, for example. The signal processing unit 68 may be installed in the same DSP as the transmission signal generation unit 64.

[0086] The signal processing unit 68 outputs the information on the estimated distance, speed, and direction as the observation data of the detection object to the information processing unit 62. The information processing unit 62 can map-transform the detection object onto the virtual space based on the observation data and perform various processes. The information processing unit 62 is composed of one or more processors similar to the processor 13 of the imaging device 10. The information processing unit 62 can control the entire sensing device 60. The processing performed by the information processing unit 62 will be further described later.

[0087] The output unit 63 is an output interface that outputs the result processed by the information processing unit 62 to an external display device of the sensing device 60 or an ECU in the vehicle. The output unit 63 may include a communication processing circuit connected to a vehicle network such as CAN and a communication connector, etc.

[0088] Hereinafter, Figure 12 a flowchart is referred to for explaining a part of the processing performed by the information processing unit 62.

[0089] The information processing unit 62 acquires the observation data from the signal processing unit 68 (step S201).

[0090] Next, the information processing unit 62 maps the observation data onto the virtual space (step S202). In Figure 13An example of the observation data mapped onto the virtual space is shown. The observation data of the millimeter-wave radar 61 is obtained as information of points, where each point includes information on distance, speed, and direction respectively. The information processing unit 62 maps each piece of observation data onto the horizontal plane. In Figure 13 it, the horizontal axis is an axis with the center at 0, representing the left-right direction, i.e., the x-axis direction, in meters. The vertical axis is an axis with the nearest position at 0, representing the distance in the y-axis direction, i.e., the depth direction, in meters.

[0091] Next, the information processing unit 62 clusters the set of each point in the virtual space to detect the detection object (step S203). Clustering means extracting a point group as a set of points from the data representing each point. As Figure 14 shown by the dashed ellipse in it, the information processing unit 62 can extract a point group as a set of each point representing the observation data. The information processing unit 62 can judge that there is actually a detection object in a part of the multiple observation data sets. In contrast, it can be judged that the observation data corresponding to each discrete point is generated by observation noise. The information processing unit 62 can set a threshold for the number or density, etc. of the points corresponding to the observation data to judge whether the set of observation data is a detection object. The information processing unit 62 can estimate the size of the detection object based on the size of the area occupied by the point group.

[0092] Next, the information processing unit 62 tracks the positions of the detected point groups in the virtual space (step S204). The information processing unit 62 can use the center of the area occupied by each point group or the average of the coordinates of the positions of the points included in the point group as the position of each point group. The information processing unit 62 grasps the movement of the detection object in a time series by tracking the movement of the point group.

[0093] After step S204 or in parallel with step S204, the information processing unit 62 estimates the category of the detection object corresponding to each point group (step S205). The categories of the detection objects include "vehicle", "pedestrian", and "two-wheeled vehicle", etc. The determination of the category of the detection object can be performed using any one or more of the speed, size, shape, position, density of the points of the observation data, intensity of the detected reflected wave, etc. of the detection object. For example, the information processing unit 62 can accumulate the Doppler speed of the detection object obtained from the signal processing unit 68 in time series and estimate the category of the detection object based on the distribution pattern of the Doppler speed. In addition, the information processing unit 62 can estimate the category of the detection object based on the information on the size of the detection object estimated in step S203. Further, the information processing unit 62 can obtain the intensity of the reflected wave corresponding to the observation data from the signal processing unit 68 to estimate the category of the detection object. For example, since a vehicle containing a large amount of metal has a large radar cross-section, the intensity of the reflected wave is stronger compared to a pedestrian with a small radar cross-section. The information processing unit 62 can estimate the category of the detection object and calculate the reliability indicating the accuracy of the estimation.

[0094] After step S205, the information processing unit 62 maps and transforms the detection object from the virtual space to the display space, that is, the display space (step S206). The display space can be a two-dimensional plane like an image space that represents the three-dimensional observation space observed from the viewpoint of the user. The display space can be a two-dimensional space for observing the object from the z-axis direction (vertical direction). The information processing unit 62 can also directly map the observation data obtained from the signal processing unit 68 in step S201 to the display space without going through steps S203 to S205.

[0095] The information processing unit 62 can perform further data processing based on the detection object mapped and transformed onto the display space and data such as the position, speed, size, and category of the detection object obtained through steps S203 to S206 (step S207). For example, the information processing unit 62 can continuously observe the size of the detection object in the display space and estimate the size of the detection object at the current time point based on the observed value of the size of the detection object at the current time point and the estimated value of the size of the detection object in the past. Therefore, the information processing unit 62 can use the observation data of the millimeter-wave radar 61 to perform processing similar to the Figure 3 processing shown. In addition, the information processing unit 62 can output each data from the output unit 63 for processing in other devices (step S207).

[0096] As described above, when using the millimeter-wave radar 61 as a sensor, the sensing device 60 can also perform processing similar to the case of using a photographing device as a sensor and obtain similar effects. Figure 11The sensing device 60 is internally provided with a millimeter-wave radar 61 and an information processing unit 62. However, the millimeter-wave radar and the information processing device having the function of the information processing unit 62 may also be provided independently.

[0097] For the embodiments of the present invention, descriptions have been made based on the respective drawings and embodiments. However, it should be noted that those skilled in the art can easily make various deformations or modifications based on the present invention. Therefore, it should be noted that these deformations or modifications are included within the scope of the present invention. For example, the functions and the like included in each component or each step can be reconfigured in a logically non-contradictory manner, and multiple components or steps can be combined into one or divided. The embodiments of the present invention have been described centering on the device, but the embodiments of the present invention can also be implemented as a method including the steps executed by each component of the device. The embodiments of the present invention can be implemented as a method, a program, or a storage medium recording the program executed by a processor included in the device. It should be understood that these are also included within the scope of the present invention.

[0098] The "moving body" in the present invention includes vehicles, ships, and aircraft. The "vehicle" in the present invention includes automobiles and industrial vehicles, but is not limited thereto, and may also include railway vehicles, living vehicles, and fixed-wing aircraft traveling on a runway. Automobiles include, but are not limited to, sedans, trucks, buses, two-wheel vehicles, and trolleybuses, etc., and may include other vehicles traveling on roads. Industrial vehicles include industrial vehicles for agriculture and construction. Industrial vehicles include, but are not limited to, forklifts and golf carts. Industrial vehicles for agriculture include tractors, cultivators, transplanters, binders, combine harvesters, and lawn mowers, but are not limited thereto. Industrial vehicles for construction include bulldozers, scrapers, excavators, cranes, dump trucks, and loaders, but are not limited thereto. Vehicles include vehicles powered by human power. In addition, the classification of vehicles is not limited to the above examples. For example, automobiles may include industrial vehicles capable of traveling on roads, and the same vehicle may be included in multiple classifications. The ships in the present invention include hydroplanes, vessels, and oil tankers. The aircraft in the present invention include fixed-wing aircraft and rotary-wing aircraft, etc.

[0099] Description of reference numerals:

[0100] 1: Image processing system

[0101] 10: Imaging device (sensing device)

[0102] 11: Imaging optical system

[0103] 12: Imaging element

[0104] 13: Processor

[0105] 20: Image processing device (information processing device)

[0106] 21: Input interface

[0107] 22: Storage unit

[0108] 23: Processor

[0109] 24: Output interface

[0110] 30: Display

[0111] 40: Subject (detection object)

[0112] 41: Image space (display space)

[0113] 42: Subject image

[0114] 43: Representative point

[0115] 44: Reference plane

[0116] 45: Mass point

[0117] 46: Virtual space

[0118] 48: Image element

[0119] 50: Imaging device (sensing device)

[0120] 51: Imaging optical system

[0121] 52: Imaging element

[0122] 53: Storage unit

[0123] 54: Processor

[0124] 55: Output interface

[0125] 60: Sensing device

[0126] 61: Millimeter-wave radar (sensor)

[0127] 62: Information processing unit

[0128] 63: Output unit

[0129] 64: Transmission signal generation unit

[0130] 65: High-frequency circuit

[0131] 66: Transmission antenna

[0132] 67: Reception antenna

[0133] 68: Signal processing unit

[0134] 100: Vehicle (mobile body)

Claims

1. An information processing apparatus, wherein, Comprising: An input interface configured to obtain observation data obtained from an observation space; A processor configured to detect a detection object included in the observation data, map-transform the coordinates of the detected detection object into the coordinates of the detection object in a virtual space, the virtual space being a two-dimensional space in which the value in the z-axis direction is set to a specified fixed value in a coordinate system of a three-dimensional real space composed of the x-axis, y-axis, and z-axis of the real space, track the position and velocity of a mass point representing the detection object on the virtual space, map-transform the coordinates of the mass point on the virtual space to the coordinates on a display space, and sequentially observe the size of the detection object on the display space, and estimate the size of the detection object at the current time point based on the observed value of the size of the detection object at the current time point and the estimated value of the size of the detection object in the past; And An output interface configured to output output information based on the coordinates of the mass point mapped to the display space and the estimated size of the detection object; Based on a representative point at the center of the lowermost end of the detection object in the display space, combined with the internal parameters of the imaging device, calculate the intersection point of the representative point with a reference plane in the three-dimensional real space when the value in the z-axis direction is the specified fixed value through ray projection, and define this intersection point as the mass point.

2. The information processing apparatus according to claim 1, wherein The processor is configured to use a Kalman filter to track the position and velocity of the mass point.

3. The information processing apparatus according to claim 1 or 2, wherein The size of the detection object includes the width and height of the detection object, The estimated values of the width and height of the detection object at the first previous time point of the current time point are respectively set as W(k - 1) and H(k - 1), The observed values of the width and height of the detection object at the current time point are respectively set as Wmeans(k) and Hmeans(k), The processor is further configured to use a parameter α in the range of 0 ≤ α ≤ 1 and estimate the estimated values W(k) and H(k) of the width and height of the detection object at the current time point through the following mathematical expressions: W(k) = W(k - 1) + α(Wmeans(k) - W(k - 1)) (1) H(k) = H(k - 1) + α(Hmeans(k) - H(k - 1)) (2).

4. The information processing apparatus according to claim 3, wherein The processor is configured to estimate the accuracy of the detection of the detection object included in the observation data and dynamically adjust the parameter α based on the detection accuracy.

5. The information processing apparatus according to claim 3, wherein The processor is configured to dynamically adjust the parameter α according to the velocity of the movement of the detection object in the observation space.

6. The information processing apparatus according to claim 1 or 2, wherein In a case where an observed value of the size of the detection object at the current time point cannot be obtained, the processor is configured to estimate the size of the detection object at the current time point based only on an estimated value of the size of the detection object in the past.

7. The information processing apparatus according to claim 6, wherein the size of the detection object includes the width and height of the detection object, an estimated value of the width and height of the detection object that is the first one before the current time point is set as W(k - 1) and H(k - 1) respectively, an estimated value of the width and height of the detection object that is the second one before the current time point is set as W(k - 2) and H(k - 2) respectively, the processor is configured to use a parameter β included in the range of 0 ≤ β ≤ 1 and estimate the estimated values W(k) and H(k) of the width and height of the detection object at the current time point by the following mathematical expressions: W(k) = W(k - 1) + β(W(k - 1) - W(k - 2)) (3) H(k) = H(k - 1) + β(H(k - 1) - H(k - 2)) (4).

8. A sensing device, wherein, Comprising: a sensor configured to sense an observation space and acquire observation data of a detection object; a processor configured to detect a detection object included in the observation data, map-transform coordinates of the detected detection object into coordinates of a detection object in a virtual space, the virtual space being a two-dimensional space in which a value in the z-axis direction is set to a specified fixed value in a coordinate system of a three-dimensional real space composed of the x-axis, y-axis, and z-axis of the real space, track positions and velocities of particles representing the detection object on the virtual space, map-transform the coordinates of the tracked particles on the virtual space into coordinates on a display space, and sequentially observe the size of the detection object on the display space, and estimate the size of the detection object at the current time point based on an observed value of the size of the detection object at the current time point and an estimated value of the size of the detection object in the past; and an output interface configured to output output information based on the coordinates of the particles mapped to the display space and the estimated size of the detection object, Based on a representative point at the center of the lowermost end of the detection object in the display space, combined with internal parameters of the imaging device, calculate an intersection point of the representative point corresponding to a reference plane in the three-dimensional real space when the value in the z-axis direction is the specified fixed value by light projection, and define this intersection point as the particle.

9. A moving body, wherein it includes the sensing device according to claim 8.

10. An information processing method, wherein, Including: acquiring observation data from an observation space; detecting a detection object included in the observation data; map-transforming coordinates of the detected detection object into coordinates of a detection object in a virtual space, the virtual space being a two-dimensional space in which a value in the z-axis direction is set to a specified fixed value in a coordinate system of a three-dimensional real space composed of the x-axis, y-axis, and z-axis of the real space; tracking positions and velocities of particles representing the detection object on the virtual space; Map and transform the coordinates of the tracked particle in the virtual space into coordinates in the display space; Successively observe the size of the detection object in the display space; Based on the observed value of the size of the detection object at the current time point and the estimated value of the size of the detection object in the past, estimate the size of the detection object at the current time point; Output output information based on the coordinates of the particle mapped and transformed onto the display space and the estimated size of the detection object; Based on the representative point at the center of the lowermost end of the detection object in the display space, combined with the internal parameters of the imaging device, calculate the intersection point of the reference plane in the three-dimensional real space corresponding to this representative point when the value in the z-axis direction is the specified fixed value through ray projection, and define this intersection point as the particle.

Citation Information

Patent Citations

  • Rear side watching device

    JP1999321494A

  • Minimal user input video analytics systems and methods

    US20170270689A1