Information processing device and information processing method

The information processing device addresses the high processing load and accuracy limitations in image recognition by converting images to match the imaging device's optical system and generating multiple learning models for improved accuracy and efficiency.

JP7672801B2Active Publication Date: 2025-05-08CANON KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2020155371
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-09-16
Publication Date
2025-05-08
Estimated Expiration
2040-09-16

AI Technical Summary

Technical Problem

Existing image recognition methods face high processing loads and limitations in generating highly accurate learning models, particularly due to the need for conversion processing in each image recognition and the constraints on image collection.

Method used

An information processing device and method that generate a learning model for image recognition by converting images to have distortion characteristics matching the optical system of the imaging device, allowing for the generation of multiple learning models based on different image regions and skipping unnecessary image recognition processes.

Benefits of technology

This approach reduces processing load and enhances the accuracy of the learning model by accounting for distortion characteristics and optimizing image recognition processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672801000001
    Figure 0007672801000001
  • Figure 0007672801000002
    Figure 0007672801000002
  • Figure 0007672801000003
    Figure 0007672801000003
Patent Text Reader

Abstract

To provide an information processor and an information processing method capable of reducing a processing load and generating a highly accurate learning model.SOLUTION: An information processor 20 that generates a learning model for performing image recognition on a first image acquired by an imaging device 10 includes: an image conversion unit 215 that converts a second image for learning to generate a third image with distortion characteristics based on first optical characteristics; and a learning model generation unit 216 that generates a learning model based on the third image.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an information processing device and an information processing method. [Background technology]

[0002] One of the image recognition methods uses a learning model generated by machine learning. Patent Documents 1 and 2 disclose techniques related to recognition of captured images using a learning model.

[0003] Patent Document 1 discloses a method of converting an image to be processed based on the imaging conditions of the image used to generate a learning model and the imaging conditions of the image to be processed, and then inputting the converted image to the learning model.

[0004] Patent Document 2 discloses a method for generating a learning model using images captured by an imaging device including a lens that generates distortions on a subject in an image, such as a fisheye lens. In this method, a distorted captured image is converted into an equal image, related information is added, and then the equal image is converted into a distorted image to generate data for learning. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] JP 2019-125116 A [Patent Document 2] JP 2019-117577 A Summary of the Invention [Problem to be solved by the invention]

[0006] In the method of Patent Document 1, a conversion process needs to be performed on a captured image every time image recognition is performed, so the processing load during image recognition is large. Therefore, depending on the application, it may be difficult to apply the method of Patent Document 1.

[0007] In the method of Patent Document 2, the conversion process is performed during learning, not during image recognition, so the processing load described above can be reduced. However, since the learning images are captured using an image capture device for image recognition, there is a limit to the amount of images that can be collected, and there are cases in which the accuracy of the learning model is not sufficient.

[0008] An object of the present invention is to provide an information processing device and an information processing method capable of reducing the processing load and generating a highly accurate learning model. [Means for solving the problem]

[0009] According to one aspect of the present invention, there is provided an information processing device that generates a learning model for performing image recognition on a first image acquired by a first imaging device equipped with an optical system having a first optical characteristic, the information processing device including: a conversion unit that converts a second image for learning to generate a third image having distortion characteristics based on the first optical characteristic; and a generation unit that generates the learning model based on the third image. an image recognition unit that performs image recognition on the first image; the second image is an image acquired by a second imaging device having an optical system having a second optical characteristic different from the first optical characteristic. Each of the third images has a distortion characteristic corresponding to a position in the first image, the generation unit generates the learning models corresponding to the positions in the first image based on the third images, and the image recognition unit skips at least one of the learning models to omit image recognition for a part of the area of ​​the first image. The present invention provides an information processing apparatus comprising:

[0010] According to another aspect of the present invention, there is provided an information processing method for generating a learning model for performing image recognition on a first image acquired by a first imaging device equipped with an optical system having a first optical characteristic, the method including: converting a second image for learning to generate a third image having distortion characteristics based on the first optical characteristic; and generating the learning model based on the third image. performing image recognition on the first image; the second image is an image acquired by a second imaging device having an optical system having a second optical characteristic different from the first optical characteristic. The step of generating the third image generates a plurality of the third images having different distortion characteristics, each of the plurality of the third images having a distortion characteristic corresponding to a position in the first image, the step of generating the learning model generates a plurality of the learning models corresponding to a position in the first image based on the plurality of the third images, and the step of performing the image recognition skips at least one of the plurality of the learning models and omits image recognition for a part of an area of ​​the first image. The present invention provides an information processing method comprising the steps of: Effect of the Invention

[0011] The present invention provides an information processing device and an information processing method that can reduce the processing load and generate a highly accurate learning model. [Brief description of the drawings]

[0012] [Figure 1] 1 is a block diagram showing an overall configuration of an image recognition system according to a first embodiment. [Diagram 2] 1 is a block diagram showing a schematic configuration of an imaging device according to a first embodiment. [Diagram 3] 1 is a block diagram showing a hardware configuration of an information processing device according to a first embodiment. [Figure 4] 1 is a functional block diagram of an information processing device according to a first embodiment. [Diagram 5] 4 is a flowchart showing an outline of a learning process in the information processing device according to the first embodiment. [Figure 6] FIG. 2 is a diagram conceptually showing a neural network that can be used for image recognition in the information processing device according to the first embodiment. [Figure 7] 4 is a flowchart showing an outline of image recognition processing in the information processing device according to the first embodiment. [Figure 8] 10 is a flowchart showing an outline of a learning process in an information processing device according to a second embodiment. [Figure 9] 13A to 13C are diagrams illustrating an example of image conversion in an information processing device according to a second embodiment. [Figure 10] 13 is a flowchart showing an outline of a learning process in an information processing device according to a third embodiment. [Figure 11] FIG. 13 is a diagram illustrating an example of the configuration of an image recognition system and a moving object according to a fourth embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. The same elements or corresponding elements in multiple drawings are denoted by the same reference numerals, and the description thereof may be omitted or simplified.

[0014] [First embodiment] The image recognition system of this embodiment is a system that performs image recognition on a captured image and outputs a recognition result. An example of the use of the image recognition system is an automatic monitoring system that automatically determines whether or not a monitoring target is present within an image capture range. Typically, the image recognition system can realize real-time image recognition by repeatedly performing recognition processing based on a moving image or continuous images.

[0015] The image recognition system of this embodiment also has a machine learning function for generating a learning model using training images. The trained learning model generated by this function is used for the above-mentioned image recognition. The training images are stored in a database provided in the image recognition system in advance.

[0016] FIG. 1 is a block diagram showing the overall configuration of an image recognition system according to this embodiment. The image recognition system includes an imaging device 10 and an information processing device 20 that are communicably connected to each other. The imaging device 10 is a device that captures an image of the surroundings of the location where the imaging device 10 is installed and acquires an image. The imaging device 10 may be, for example, a surveillance camera, a digital still camera, a digital camcorder, a smartphone, an in-vehicle camera, an observation satellite, or the like. The imaging device 10 desirably uses a lens with a wide angle of view, such as a fisheye lens, in order to perform image recognition over a wide range. The information processing device 20 is a computer such as a PC or a server, and performs arithmetic processing such as image processing using images acquired from the imaging device 10. The information processing device 20 may have a function of controlling the imaging device 10 to perform imaging.

[0017] The device configuration of the image recognition system is not limited to that shown in FIG. 1. For example, the image recognition system may be an integrated image recognition device having the functions of the imaging device 10 and the information processing device 20. At least one of the imaging device 10 and the information processing device 20 may be provided in multiple numbers. For example, by providing multiple imaging devices 10, multiple shooting ranges may be captured in parallel. For example, by providing multiple information processing devices 20, the image processing of this embodiment may be performed by multiple devices in cooperation. The information processing device 20 may be divided into a learning device that generates a learning model and an image recognition device that performs image recognition using the learning model. The image recognition system may further include devices other than the imaging device 10 and the information processing device 20. For example, the image recognition system may further include a control device that controls the entire system, and in this case, the imaging device 10 and the information processing device 20 may perform image recognition processing according to the control of the control device.

[0018] Fig. 2 is a block diagram showing a schematic configuration of an image pickup device 10 according to this embodiment. As shown in Fig. 2, the image pickup device 10 has a photoelectric conversion device 101, a lens 102, an aperture 103, a barrier 104, a signal processing unit 105, a timing generating unit 111, and an overall control / calculation unit 110. The image pickup device 10 also has a memory unit 106, a recording medium control I / F (Interface) unit 109, and an external I / F unit 107.

[0019] The photoelectric conversion device 101 is a solid-state imaging element such as a CMOS image sensor or a CCD image sensor. The photoelectric conversion device 101 is typically a two-dimensional area sensor having a plurality of pixels arranged in a plurality of rows and a plurality of columns, and each of the plurality of pixels outputs a signal according to incident light. The lens 102 is for forming an optical image of a subject in an imaging area of ​​the photoelectric conversion device 101. As described above, the lens 102 may be a lens with a wide angle of view such as a fisheye lens. The aperture 103 is for varying the amount of light passing through the lens 102. The barrier 104 is for protecting the lens 102.

[0020] The signal processing unit 105 performs desired processing, correction, data compression, and the like on the signal output from the photoelectric conversion device 101. The signal processing unit 105 includes a circuit such as a digital signal processor. The signal processing unit 105 may be mounted on the same board as the photoelectric conversion device 101, or may be mounted on a different board. Also, a part of the functions of the signal processing unit 105 may be mounted on the same board as the photoelectric conversion device 101, and another part of the functions of the signal processing unit 105 may be mounted on a different board. Also, the photoelectric conversion device 101 may output an analog signal before AD conversion, instead of a digital signal. In that case, the signal processing unit 105 may further include an AD converter.

[0021] The timing generating unit 111 is for outputting various timing signals to the photoelectric conversion device 101 and the signal processing unit 105. The overall control / calculation unit 110 is a control unit that manages the overall driving and calculation processing of the imaging device 10. Here, control signals such as timing signals may be input from outside the imaging device 10, and it is sufficient for the imaging device 10 to have at least the photoelectric conversion device 101 and the signal processing unit 105 that processes signals output from the photoelectric conversion device 101.

[0022] The memory unit 106 is a frame memory unit for temporarily storing image data. The recording medium control I / F unit 109 is an interface unit for recording to the recording medium 108 or reading from the recording medium 108. The external I / F unit 107 is an interface unit for communicating with an external information processing device 20 or the like. The recording medium 108 is a recording medium such as a semiconductor memory for recording or reading imaging data. The recording medium 108 may be built into the imaging device 10 or may be removable.

[0023] 3 is a block diagram showing a hardware configuration of the information processing device 20 according to this embodiment. The information processing device 20 has a CPU 201, a RAM 202, a ROM 203, a HDD (Hard Disk Drive) 204, a communication I / F 205, an input device 206, and an output device 207. These units are connected to each other via a bus or the like.

[0024] The CPU 201 is a processor that reads out programs stored in the ROM 203 and the HDD 204 into the RAM 202, executes the programs, and performs calculation processing and controls each unit of the information processing device 20. The processing performed by the CPU 201 may include generation of a learning model, image recognition, and the like.

[0025] The RAM 202 is a volatile storage medium and functions as a work memory when the CPU 201 executes a program. The ROM 203 is a non-volatile storage medium and stores firmware and the like necessary for the operation of the information processing device 20. The HDD 204 is a non-volatile storage medium and stores programs, image data, and the like used in the learning process, image recognition, and other processes of this embodiment.

[0026] The communication I / F 205 is a communication device based on standards such as Wi-Fi (registered trademark), Ethernet (registered trademark), Bluetooth (registered trademark), etc. The communication I / F 205 is used for communication with the imaging device 10, other computers, etc.

[0027] The input device 206 is a device for inputting information to the information processing device 20, and is typically a user interface for a user to operate the information processing device 20. Examples of the input device 206 include a keyboard, a button, a mouse, a touch panel, and the like.

[0028] The output device 207 is a device that outputs information from the information processing device 20 to the outside, and is typically a user interface for presenting information to a user. Examples of the output device 207 include a display and a speaker.

[0029] The above-mentioned configuration of the information processing device 20 is an example and can be changed as appropriate. For example, examples of processors that can be mounted on the information processing device 20 include GPU, ASIC, FPGA, etc., in addition to the above-mentioned CPU 201. A plurality of these processors may be provided, and the plurality of processors may perform processing in a distributed manner. The function of storing information such as image data in the HDD 204 may be provided in another data server, not in the information processing device 20. The HDD 204 may be a storage medium such as an optical disk, a magneto-optical disk, or an SSD (Solid State Drive).

[0030] 4 is a functional block diagram of the information processing device 20 according to this embodiment. The information processing device 20 includes a learning image storage unit 211, a first distortion information storage unit 212, a second distortion information storage unit 213, a conversion parameter calculation unit 214, an image conversion unit 215, a learning model generation unit 216, an image acquisition unit 221, an image recognition unit 222, and a learning model storage unit 223.

[0031] The CPU 201 executes a program to perform a predetermined calculation process. The CPU 201 also executes a program to control each unit in the information processing device 20. Through these processes, the CPU 201 realizes the functions of a conversion parameter calculation unit 214, an image conversion unit 215, a learning model generation unit 216, an image acquisition unit 221, and an image recognition unit 222.

[0032] The HDD 204 functions as a database that stores the learning images, the first distortion information, the second distortion information, and the learning model. Thus, the HDD 204 functions as a learning image storage unit 211, a first distortion information storage unit 212, a second distortion information storage unit 213, and a learning model storage unit 223.

[0033] Some of the functional blocks shown in FIG. 4 may be provided in an external device of the information processing device 20, and the functional blocks shown in FIG. 4 may be realized by cooperation of a plurality of devices. For example, the functional blocks shown in FIG. 4 may be realized by a learning device and an image recognition device. In this case, the learning device may have a learning image storage unit 211, a first distortion information storage unit 212, a second distortion information storage unit 213, a conversion parameter calculation unit 214, an image conversion unit 215, and a learning model generation unit 216. Furthermore, the image recognition device may have an image acquisition unit 221, an image recognition unit 222, and a learning model storage unit 223. Furthermore, as another modified example, some or all of the functions of the learning image storage unit 211, the first distortion information storage unit 212, the second distortion information storage unit 213, and the learning model storage unit 223 may be realized by a data server external to the information processing device 20.

[0034] 5 is a flowchart showing an outline of the learning process in the information processing device 20 according to the first embodiment. This learning process is a process for generating a learning model, and is performed in advance based on a start operation by a user, prior to image recognition processing using the learning model. Note that this learning process may be a process for performing additional learning on an existing learning model for which learning has been completed.

[0035] The learning images are assumed to be stored in advance in the learning image storage unit 211, but may be acquired by the information processing device 20 from a database outside the information processing device 20 during learning. The learning images are typically a big data group including a large number of images obtained by capturing images of the object to be identified in various situations. Therefore, the learning images are usually images captured by an imaging device other than the imaging device 10. When setting related information such as the type of object is necessary for learning, this related information is assumed to be associated with the learning images in advance and stored in the learning image storage unit 211. The related information may be, for example, the name of an object included in the image.

[0036] In step S11, the conversion parameter calculation unit 214 acquires the first distortion information stored in the first distortion information storage unit 212 and the second distortion information stored in the second distortion information storage unit 213. Then, the conversion parameter calculation unit 214 calculates conversion parameters based on the first distortion information and the second distortion information.

[0037] Here, the first distortion information is information on distortion characteristics occurring in an image (first image) captured by the imaging device 10 due to optical characteristics (first optical characteristics) of the optical system of the imaging device 10 (first imaging device). More specifically, the first distortion information can be information on distortion of the lens 102 used in the imaging device 10. The first distortion information is stored in advance in the first distortion information storage unit 212 according to the imaging device 10 to be used for image recognition.

[0038] Moreover, the second distortion information is information on distortion characteristics occurring in an image due to optical characteristics (second optical characteristics) of the optical system of an imaging device (second imaging device) used to capture the learning image (second image). More specifically, the second distortion information can be information on distortion of the lens of the imaging device used to capture the learning image. The second distortion information is stored in advance in the second distortion information storage unit 213 according to the learning image used for learning this process. Note that the second distortion information may be stored in association with the learning image, and in this case, the conversion parameter calculation unit 214 may acquire the second distortion information from the learning image storage unit 211.

[0039] The transformation parameters generated by the processing of this step indicate the mode and degree of transformation in the image transformation described below. The transformation parameter calculation unit 214 calculates the transformation parameters so as to bring the distortion of the learning image closer to the distortion of the image captured by the imaging device 10. The transformation parameters may also include parameters such as the rotation angle of the image, the degree of enlargement / reduction, and brightness change.

[0040] In step S12, the image conversion unit 215 (conversion unit) acquires the learning image stored in the learning image storage unit 211 and the conversion parameters calculated in step S11. Then, the image conversion unit 215 converts the learning image based on the conversion parameters. Note that the conversion in this process may be, for example, a geometric conversion that performs a projection method conversion between lenses with different projection methods. In addition, pixel value complementation processing may be performed during the conversion. Note that the number of learning images converted in step S12 may be the number necessary to generate a learning model described later, and is generally more than one.

[0041] In step S13, the learning model generation unit 216 (generation unit) generates a learning model based on the converted learning image (third image). The learning model that can be used in this embodiment can be, for example, a neural network as exemplified in Fig. 6. Fig. 6 is a conceptual diagram showing a neural network that can be used for image recognition in the information processing device 20 according to this embodiment.

[0042] The neural network shown in FIG. 6 has a plurality of nodes 30. The plurality of nodes 30 form an input layer, an intermediate layer, and an output layer. Image data is input to the input layer. Each node 30 performs a calculation using an activation function including input values ​​from a plurality of other nodes 30, weighting coefficients, and bias values, and outputs the calculation result to the node 30 in the next layer. The node 30 in the output layer outputs the calculation result of the neural network based on the input from the intermediate layer in the previous stage. This output may mean some kind of judgment result regarding the input image data. Specific examples of the judgment result may be, for example, whether or not an object is present in the input image, the position of the object, the type of the object, etc.

[0043] In the case of a neural network example, the learning model corresponds to the structure of the neural network and the weighting coefficients and bias values ​​included in the activation function, and learning corresponds to appropriately determining the weighting coefficients and bias values.

[0044] An example of a neural network learning method will be described. Learning images for which correct answer values ​​have been set in advance are prepared. Then, the weighting coefficients and bias values ​​of each node 30 are optimized so that the correct answer value set when the learning image is input to the input layer is output from the node 30 in the output layer. An example of this optimization method is the back error propagation method. By performing such learning processing using a large number of learning images, a learning model that can perform appropriate image recognition even for unknown images other than the learning images can be generated.

[0045] The learning model that can be used in this embodiment is not limited to the one using a neural network as in the above example, and may be, for example, a random forest, a support vector machine, etc. In addition, the structure of the neural network shown in Fig. 6 is simplified for the convenience of explanation, and the actual number of layers, number of nodes, etc. may be much larger than those shown in the figure.

[0046] In step S14, the learning model storage unit 223 stores the learning model generated in step S13. Next, the image recognition process using the trained learning model will be described.

[0047] 7 is a flowchart showing an outline of image recognition processing in the information processing device 20 according to the first embodiment. This image recognition processing is processing for recognizing an image captured by the imaging device 10 using a learning model. In the case where this image recognition system repeats imaging and recognition like an automatic surveillance system, this processing is repeatedly executed every time imaging is performed by the imaging device 10. Alternatively, this processing may be executed based on a start operation from a user.

[0048] In step S21, the image acquisition unit 221 acquires an image captured by the imaging device 10. In step S22, the image recognition unit 222 recognizes the image acquired in step S21 using a learning model stored in the learning model storage unit 223. For example, when the learning model used in the image recognition unit 222 is a neural network as shown in FIG. 6, this process may be a process in which image data is input to an input layer and a calculation result is output from an output layer. In step S23, the information processing device 20 outputs the calculation result in the image recognition unit 222 to the outside. Note that the image acquisition unit 221 may further acquire the output of the distance sensor 40 as shown in FIG. 4 and use it for processing.

[0049] The effect of the learning process of this embodiment will be described in more detail. In general, in order to improve accuracy in machine learning for image recognition, it is effective to prepare many learning images. To this end, since the number of images obtained is limited in the method of capturing many images exclusively for learning and preparing learning images, it is desirable to utilize existing images, also called big data. However, if the optical characteristics of the imaging device used for image recognition are different from the optical characteristics of the imaging device used to capture the learning images, the same object will be captured with different shapes, etc., and the accuracy of image recognition may decrease. In particular, while lenses with a wide angle of view and large distortion, such as fisheye lenses, are often used for imaging devices for automatic monitoring, etc., existing images are often captured by imaging devices using general lenses of the central projection method. In such cases, the influence of the difference in image distortion due to the difference in lenses may become significant.

[0050] Therefore, in this embodiment, learning is performed by converting a learning image and using an image having distortion characteristics based on the optical characteristics of the optical system of the imaging device 10 used for image recognition. This allows learning to be performed taking into account the distortion characteristics of the image captured by the imaging device 10. Therefore, according to this embodiment, an information processing device and an information processing method capable of generating a highly accurate learning model are provided.

[0051] In this embodiment, in addition to the optical characteristics of the optical system of the imaging device 10 used for image recognition, the optical characteristics of the optical system of the imaging device that captured the learning image are also taken into consideration during conversion. Therefore, learning is performed taking into account the difference in distortion characteristics between the learning image and the image acquired by the imaging device 10. This further improves the accuracy of the learning model.

[0052] [Second embodiment] The image recognition system of this embodiment differs from the first embodiment in that, in the learning process, a single learning image is subjected to a plurality of different transformations according to the region of the lens. That is, in this embodiment, a plurality of transformed learning images having different distortion characteristics are generated. In the following, the differences from the first embodiment will be described, and the description of the parts common to the first embodiment will be omitted or simplified.

[0053] 8 is a flowchart showing an outline of the learning process in the information processing device according to the second embodiment. The learning process in this embodiment includes a loop process in which the lens 102 is divided into a plurality of regions, and the learning image corresponding to each region of the lens 102 is converted. This loop process includes steps S31, S11, and S12. The learning image corresponding to one region is converted each time the loop process goes through one cycle. The loop counter variable for the loop process is k.

[0054] In step S31, the transformation parameter calculation unit 214 acquires first distortion information corresponding to the k-th lens region among the multiple regions of the lens 102 from the first distortion information storage unit 212. The first distortion information may differ for each lens region. Note that the first distortion information for each lens region is assumed to be stored in advance in the first distortion information storage unit 212, but when calculating the transformation parameters, the information may be calculated by calculation from a theoretical formula indicating the lens projection method and the coordinates of the lens region.

[0055] Steps S11 and S12 are the same as those in the first embodiment except that the first distortion information does not correspond to the entire lens 102 but corresponds to the k-th lens region which is a part of the lens 102, and therefore a description thereof will be omitted. When the conversion of all the lens regions is completed, the loop process ends. In this loop process, different conversion processes are performed on one learning image for the number of lens regions, and converted learning images are obtained for the number of lens regions. Thereafter, the processes of steps S13 and S14 are performed as in the first embodiment. Through these processes, conversion is performed using different first distortion information corresponding to each of the multiple regions of the lens 102, and a learning model is generated from a large number of learning images corresponding to each conversion.

[0056] The conversion process for each lens region in this embodiment will be described in more detail with a specific example. Fig. 9(a) to Fig. 9(f) are diagrams for explaining an example of image conversion in the information processing device 20 according to the second embodiment. In the explanation of this specific example, the lens 102 of the imaging device 10 is a fisheye lens with a field angle of 180° using a projection method such as an equisolid angle projection method. In addition, the learning image is taken by an imaging device using a lens using a central projection method such that the subject and the image have a similar shape.

[0057] FIG. 9(a) is a diagram showing an image captured by an imaging device 10 using a fisheye lens and a method of dividing the image. When the optical axis of the fisheye lens is oriented vertically, the imaging range is the entire surface of the celestial sphere. In this case, the coordinates of each pixel in the captured image are the coordinates of the entire surface of the celestial sphere projected onto a circle on a plane. Due to this projection method, the distortion of the subject in the captured image becomes larger the closer to the surface of the celestial sphere (the closer to the outer periphery of the circle in FIG. 9(a)), and becomes smaller the closer to the zenith (the closer to the center of the circle in FIG. 9(a)). In addition, the image of the subject rotates at different angles depending on the position in the captured image. Therefore, as shown in FIG. 9(a), the region is divided by multiple center lines and multiple concentric circles. The range of the lens corresponding to each region in the image shown in FIG. 9(a) corresponds to the lens region described in FIG. 8.

[0058] Next, the deformation of an image by the conversion process of a learning image will be described with reference to Fig. 9(b) to Fig. 9(f). Fig. 9(b) shows an example of a learning image before conversion. The image in Fig. 9(b) forms a rectangle ABCD including an image of a flying object as a recognition target. In the image in Fig. 9(b), the subject and the image are similar in shape, so it is assumed that the shape of the actual flying object is correctly reflected.

[0059] Fig. 9(c) is a diagram showing the deformation of the learning image corresponding to the region P1 in Fig. 9(a). Fig. 9(d) is a diagram showing the deformation of the image of the flying object accompanying the deformation of the learning image. When the rectangle ABCD in the region P1 is converted from the central projection method to the equal solid angle projection method, as shown in Fig. 9(c), the rectangle ABCD is rotated by about 158° and deformed so that each side is distorted. As a result, as shown in Fig. 9(d), the shape of the image of the flying object is also rotated and deformed in the same way.

[0060] FIG. 9(e) is a diagram showing the deformation of the learning image corresponding to the region P2 in FIG. 9(a). FIG. 9(f) is a diagram showing the deformation of the image of the flying object accompanying the deformation of the learning image. When the rectangle ABCD in the region P2 is converted from the central projection method to the equal solid angle projection method, the rectangle ABCD is rotated by about 23° and deformed from a rectangle to a sector shape as shown in FIG. 9(e). As a result, the shape of the image of the flying object is also rotated and deformed in the same way as shown in FIG. 9(f). As can be understood by comparing FIG. 9(d) and FIG. 9(f), the manner and degree of deformation differ depending on the region.

[0061] When capturing an image of an object such as a flying object using a fisheye lens, it should be considered that the manner and degree of deformation of the image changes depending on the position in the image due to the nature of the projection method such as the equisolid angle projection method. It is difficult to consider this effect in a deformation that gives the same distortion over the entire surface of the fisheye lens. Therefore, in the conversion of the learning image in steps S11 and S12, the conversion parameters are made different depending on the position in the image, i.e., the position of the lens area, thereby making it possible to convert the learning image in consideration of the difference in the manner and degree of deformation depending on the position in the lens. Therefore, according to this embodiment, it is possible to provide an information processing device and an information processing method that can generate a learning model with higher accuracy.

[0062] [Third embodiment] The image recognition system of this embodiment differs from the second embodiment in that, in the learning process, a plurality of learning models are generated using images obtained by a plurality of different transformations according to the lens area. In the following, the differences from the second embodiment will be described, and the description of the parts common to the second embodiment will be omitted or simplified.

[0063] 10 is a flowchart showing an outline of the learning process in the information processing device according to the second embodiment. The learning process in this embodiment includes a loop process in which the lens 102 is divided into a plurality of regions, a learning image corresponding to each region of the lens 102 is converted, and a learning model corresponding thereto is generated. This loop process includes steps S31, S11, S12, S13, and S14. Each time the loop process goes through one cycle, the learning image corresponding to one region is converted and a learning model is generated. The loop counter variable for the loop process is k.

[0064] The specific contents of steps S31, S11, S12, S13, and S14 are generally similar to those of the second embodiment, and therefore will not be described. Unlike the second embodiment, in this embodiment, steps S13 and S14 are also included in the loop processing. As a result, in this embodiment, multiple learning models according to the lens area are generated.

[0065] 7, image recognition is performed using a different learning model for each lens region, which is a difference from the first embodiment, but other points are the same. Note that this image recognition is performed multiple times since it is performed for each lens region, but these multiple times of processing may be serial processing, parallel processing, or a combination of serial processing and parallel processing.

[0066] As described above, in this embodiment, as in the second embodiment, the conversion parameters are varied according to the position of the lens region, and the conversion is performed taking into account the above-mentioned differences in the form and degree of deformation. Therefore, the same effect as in the second embodiment is obtained. Furthermore, in this embodiment, since a different learning model is generated for each region, when each of the multiple learning models is viewed individually, the learning image of another region is not used for learning. Therefore, the recognition accuracy of the learning model is improved. Therefore, according to this embodiment, it is possible to provide an information processing device and an information processing method that can generate a learning model with higher accuracy.

[0067] In this embodiment, since there are individual learning models subdivided for each region, it is possible to perform image recognition not over the entire region of the lens 102 but not over some regions. For example, if there is a region within the shooting range of the imaging device 10 where it is known that there is no object to be identified, a learning model may be prepared for only the region excluding that region, and image recognition may be performed by skipping that region. This method can omit image recognition for some regions, thereby reducing the processing load. The region for which image recognition is skipped may be selected in advance, or a region that is determined to have no difference by difference detection with a past image may be selected.

[0068] [Fourth embodiment] An image recognition system and a moving object according to a fourth embodiment of the present invention will be described with reference to Fig. 11. Fig. 11(a) and Fig. 11(b) are diagrams showing the configurations of an image recognition system 600 and a moving object according to this embodiment.

[0069] FIG. 11(a) is a block diagram showing an example of an image recognition system 600 related to an in-vehicle camera. The image recognition system 600 has the imaging device 10 described in the first to third embodiments. The image recognition system 600 also has an image recognition unit 612 that is equipped with a learning model generated by the information processing device 20 described in the first to third embodiments and performs image recognition on an image captured by the imaging device 10. Here, the output result of the image recognition performed by the image recognition unit 612 is information related to the possibility of collision with an object. For example, the image recognition system 600 has a distance sensor 40 separate from the imaging device 10, and inputs the results of the distance sensor and the results of the imaging device 10 to the image recognition unit 612. The image recognition unit 612 merges the results of both. Then, the image recognition system 600 can take optimal avoidance action.

[0070] The image recognition system 600 has a collision determination unit 618 that determines whether or not there is a possibility of collision based on the distance calculated by the image recognition unit 612. Here, the image recognition unit 612 is an example of a distance information acquisition means (or a distance information acquisition circuit) that acquires distance information to an object. That is, the distance information is information related to parallax, defocus amount, distance to the object, and the like.

[0071] The collision determination unit 618 may determine the possibility of a collision using any of these pieces of distance information. The distance information acquisition means may be realized by dedicated hardware, a software module, or a combination of these. It may also be realized by an FPGA, an ASIC, or the like. It may also be realized by a combination of these.

[0072] The image recognition system 600 is connected to a vehicle information acquisition device 620 and can acquire vehicle information such as vehicle speed, yaw rate, steering angle, etc. The image recognition system 600 is also connected to a control ECU 630, which is a control means (control circuit) that outputs a control signal for generating a braking force for the vehicle based on the determination result in the collision determination unit 618.

[0073] The image recognition system 600 is also connected to an alarm device 640 that issues an alarm to the driver based on the determination result of the collision determination unit 618. For example, when the collision determination unit 618 determines that there is a high possibility of a collision, the control ECU 630 performs vehicle control to avoid a collision and reduce damage by applying the brakes, releasing the accelerator, suppressing engine output, etc. The alarm device 640 warns the user by sounding an alarm, displaying alarm information on the screen of a car navigation system, etc., vibrating the seat belt or steering wheel, etc.

[0074] In this embodiment, the surroundings of the vehicle, for example the front or rear, are imaged by the imaging device 10 of the image recognition system 600. Fig. 11(b) shows the image recognition system 600 when imaging the area in front of the vehicle (imaging range 650). The vehicle information acquisition device 620 sends instructions to the image recognition system 600 or the imaging device 10 to perform a predetermined operation. With this configuration, the accuracy of distance measurement can be further improved. The vehicle may further include a control means for controlling the vehicle, which is a moving body, based on the distance information.

[0075] Although the above example describes control to prevent collision with other vehicles, the image recognition system 600 can also be applied to control to automatically drive following other vehicles, control to automatically drive without going out of the lane, etc. Furthermore, the image recognition system 600 can be applied not only to vehicles, but also to moving bodies (moving devices) such as ships, aircraft, and industrial robots. In addition, the image recognition system 600 can be applied not only to moving bodies, but also to a wide range of devices that use object recognition, such as intelligent transport systems (ITS).

[0076] According to this embodiment, by using the image recognition unit 612 equipped with a highly accurate learning model, it is possible to provide a higher performance image recognition system 600 and a moving object.

[0077] [Modified embodiment] The present invention is not limited to the above-described embodiments, and various modifications are possible. For example, an example in which a part of the configuration of any of the embodiments is added to another embodiment, or an example in which a part of the configuration of another embodiment is replaced with another embodiment is also an embodiment of the present invention.

[0078] In the above-described second embodiment, a fisheye lens of an equisolid angle projection type is exemplified as the lens 102 of the imaging device 10, but the projection type is not limited to this. The projection type may be, for example, an equidistant projection type, an orthogonal projection type, a stereoscopic projection type, or the like.

[0079] In the description of the above embodiment, it is assumed that the learning images are captured by an imaging device different from the imaging device 10, but some of the learning images may include images captured by the imaging device 10. For example, when the image recognition system of this embodiment is in operation, images captured by the imaging device 10 may be added to the learning images. In these cases, image conversion processing is not essential for this image.

[0080] The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

[0081] It should be noted that the above-mentioned embodiments are merely examples of the embodiment of the present invention, and the technical scope of the present invention should not be interpreted as being limited by these. In other words, the present invention can be implemented in various forms without departing from its technical idea or main features. For example, it should be understood that an embodiment in which a part of the configuration of any of the embodiments is added to another embodiment, or an embodiment in which a part of the configuration of any of the embodiments is replaced with a part of the configuration of another embodiment, is also an embodiment to which the present invention can be applied. [Explanation of symbols]

[0082] 10. Imaging device 20 Information processing device 211 Learning image memory unit 212 First distortion information storage section 213 Second distortion information storage section 214 Conversion parameter calculation unit 215 Image conversion unit 216 Learning Model Generation Unit

Claims

1. An information processing device that generates a learning model for performing image recognition on a first image acquired by a first imaging device having an optical system having a first optical characteristic, a conversion unit that converts a second image for learning to generate a third image having distortion characteristics based on the first optical characteristic; A generation unit that generates the learning model based on the third image; an image recognition unit that performs image recognition on the first image, the second image is an image acquired by a second imaging device including an optical system having a second optical characteristic different from the first optical characteristic, each of the third images has a distortion characteristic according to a position in the first image; The generation unit generates a plurality of the learning models according to positions in the first image based on a plurality of the third images, The information processing device is characterized in that the image recognition unit skips at least one of the multiple learning models and omits image recognition for a portion of the first image.

2. The conversion unit converts the distortion characteristic of the second image further based on the second optical characteristic.

2. The information processing apparatus according to claim 1, Place.

3. The optical system of the first imaging device includes a fisheye lens.

3. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.

4. The conversion process performed by the conversion unit includes a geometric conversion for converting a projection method of the second image. The information processing device according to claim 1 .

5. An image recognition device having the learning model generated by the information processing device according to any one of claims 1 to 4; The first imaging device; An image recognition system comprising:

6. A mobile object, An image recognition device having the learning model generated by the information processing device according to any one of claims 1 to 4; a control means for controlling the moving object based on a result of the image recognition by the image recognition device; A moving object comprising:

7. An information processing method for generating a learning model for performing image recognition on a first image acquired by a first imaging device including an optical system having a first optical characteristic, comprising: converting the second image for training to generate a third image having distortion characteristics based on the first optical characteristic; generating the learning model based on the third image; performing image recognition on the first image; the second image is an image acquired by a second imaging device including an optical system having a second optical characteristic different from the first optical characteristic, The step of generating the third image includes generating a plurality of the third images having different distortion characteristics from each other, each of the third images has a distortion characteristic according to a position in the first image; The step of generating the learning model includes generating a plurality of the learning models according to positions in the first image based on a plurality of the third images; The information processing method, characterized in that the step of performing image recognition includes skipping at least one of the multiple learning models and omitting image recognition for some areas of the first image.

8. A program for causing a computer to execute the information processing method according to claim 7.

Citation Information

Patent Citations

  • Information processing apparatus, control method thereof, and computer program

    JP2016038732A

  • Program, learning processing method, learning model, data structure, learning device and object recognition device

    JP2019117577A

  • Information processing device, information processing method, and program

    JP2019125116A

  • Image processing apparatus, image processing method, and imaging apparatus

    JP2019186918A