Image processing method, imaging control method, program, image processing apparatus, and imaging apparatus

The image processing method addresses the challenge of balancing detection time and accuracy in omnidirectional images by duplicating the image's edge area to connect ends and set a detection target area, resulting in efficient and accurate person state detection.

JP7697234B2Active Publication Date: 2025-06-24RICOH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021043464
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-17
Publication Date
2025-06-24
Estimated Expiration
2041-03-17

AI Technical Summary

Technical Problem

Conventional image processing techniques struggle to balance detection time and accuracy when identifying a person's state within an image, particularly in omnidirectional images where the person may occupy a small proportion of the screen.

Method used

An image processing method that acquires an image, identifies the person's area, sets a detection target area based on the person's position, and duplicates the image's edge area to connect the ends, allowing for accurate detection within a reduced pixel range.

Benefits of technology

This method enables both faster detection times and improved accuracy when identifying a person's state within an image, even when the person occupies a small portion of the screen, and specifically addresses the challenges of omnidirectional imaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007697234000002
    Figure 0007697234000002
  • Figure 0007697234000003
    Figure 0007697234000003
  • Figure 0007697234000004
    Figure 0007697234000004
Patent Text Reader

Abstract

To provide an image processing method.SOLUTION: An image processing method causes a computer to execute a step of obtaining an image (S101), a step of identifying the position of a person region from the obtaining image (S103), a step of setting a detection target region on the basis of the identified position of the person region (S106), and a step of detecting the state of a person on the basis of the detection target region in the image (S107).SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image processing technology, and more particularly to an image processing method, an imaging control method, a program, an image processing apparatus, and an imaging apparatus.

Background Art

[0002] Conventionally, techniques for detecting a person's posture and the like from an image using deep learning have been known. In addition, a technique for automatically taking a picture when detecting a predetermined state of a person such as the above posture, facial expression, or gesture has been known. For example, Japanese Patent No. 4227257 (Patent Document 1) discloses a configuration for recognizing a subject's face, posture, and movement and automatically taking a picture when the face, posture, and movement reach a predetermined state. Japanese Patent No. 6729043 (Patent Document 2) discloses a technique for accurately specifying the position of a person in an image. More specifically, when another object overlaps the detected person area, if the other object is a moving object, the person position is specified based on the person's posture, and if the other object is a non-moving object, a predetermined position in the person area is specified as the person position.

[0003] However, the above conventional techniques had room for improvement from the viewpoint of achieving both shortening of the detection time and detection accuracy when detecting the state of a person from within the screen.

Summary of the Invention

Problems to be Solved by the Invention

[0004] The present disclosure has been made in view of the above points, and an object thereof is to provide an image processing method capable of achieving both shortening of the detection time and detection accuracy when detecting the state of a person from within the screen.

Means for Solving the Problems

[0005] In the present disclosure, in order to solve the above problems, an image processing method having the following features is provided. This image processing method causes a computer to Circulate in at least one directionA step of acquiring an image, and the position of a person area from the acquired image And size A step of specifying, and the position of the specified person area And determining the position and size of the detection target area based on the size, a step of determining whether the detection target area is located at an edge of the image based on the position and size of the detection target area, and when it is determined that the detection target area is located at an edge of the image, at least in the detection target area, a step of duplicating the edge area so that one edge area of the image is connected to the other edge, and an image obtained by adding the duplication of one edge area of the image to the other edge of the image and A step of detecting the state of a person based on the detection target area is executed.

Effect of the Invention

[0006] With the above configuration, it becomes possible to achieve both shortening of the detection time and detection accuracy when detecting the state of a person from within the screen.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Mode for Carrying Out the Invention

[0008] Hereinafter, this embodiment will be described. However, the embodiment is not limited to the embodiment described below. In the following embodiment, as an example of an image processing apparatus and an imaging apparatus, an omnidirectional imaging device 10 including two fisheye lenses will be used for explanation.

[0009] Hereinafter, the overall configuration of the omnidirectional imaging device 10 according to this embodiment will be described with reference to FIGS. 1 and 2. FIG. 1 is a cross-sectional view of the omnidirectional imaging device 10 according to this embodiment. The omnidirectional imaging device 10 shown in FIG. 1 includes an imaging body 12, a housing 14 that holds the imaging body 12 and components such as a controller and a battery, and a shooting button 18 provided on the housing 14.

[0010] The imaging body 12 shown in Fig. 1 includes two lens optical systems 20A and 20B and two imaging elements 22A and 22B. The imaging elements 22A and 22B are, for example, CMOS (Complementary Metal Oxide Semiconductor) sensors or CCD (Charge Coupled Device) sensors. The lens optical system 20 is configured as, for example, a fisheye lens of 6 groups and 7 elements or 10 groups and 14 elements. In the embodiment shown in Fig. 1, the fisheye lens has an overall angle of view greater than 180 degrees (= 360 degrees / n; optical coefficient n = 2), and preferably has an angle of view of 190 degrees or more. In the embodiment to be described, although it is described that two fisheye lenses having an overall angle of view of 180 degrees or more are used, as long as a predetermined angle of view can be obtained as a whole, it may include three or more lens optical systems and imaging elements. Also, in the embodiment to be described, although it is described that a fisheye lens is used, as long as a predetermined angle of view can be obtained as a whole, it is not precluded from using other wide-angle lenses or ultra-wide-angle lenses instead of the fisheye lens.

[0011] The optical elements (lenses, prisms, filters, and aperture stops) of the two lens optical systems 20A and 20B are defined in a positional relationship with respect to the imaging elements 22A and 22B. The optical axes of the optical elements of the lens optical systems 20A and 20B are positioned so as to be orthogonal to the center of the light-receiving area of the corresponding imaging element 22, and the light-receiving area is positioned so as to be the imaging plane of the corresponding fisheye lens. In the embodiment to be described, in order to reduce parallax, a bending optical system that divides the light collected by the two lens optical systems 20A and 20B among the two imaging elements 22A and 22B by two 90-degree prisms is adopted, but it is not limited thereto. In order to further reduce parallax, a three-time refraction structure may be used, or a straight optical system may be used to reduce costs.

[0012] In the embodiment shown in FIG. 1, the lens optical systems 20A and 20B have the same specifications and are combined in opposite directions so that their respective optical axes coincide. The imaging elements 22A and 22B convert the received light distribution into an image signal and sequentially output the images to the image processing block on the controller. Although details will be described later, the images captured by the imaging elements 22A and 22B respectively are synthesized, whereby an image of a solid angle of 4π steradians (hereinafter referred to as the "omnidirectional image") is generated. The omnidirectional image is an image that captures all directions that can be seen from the shooting location. In the embodiment to be described, although it is described as generating an omnidirectional image, it may be a full circumference image that captures only 360 degrees of the horizontal plane, that is, a so-called 360-degree panoramic image, or an image that captures a part of the panorama of the entire sphere or 360 degrees of the horizontal plane (for example, a full circumference (dome) image that captures 360 degrees horizontally and 90 degrees vertically from the horizon). Also, the omnidirectional image can be acquired as a still image or as a moving image.

[0013] FIG. 2 shows the hardware configuration of the omnidirectional imaging device 10 according to the present embodiment. The omnidirectional imaging device 10 is composed of a digital still camera processor (hereinafter simply referred to as the processor) 100, a lens barrel unit 102, and various components connected to the processor 100. The lens barrel unit 102 has the two sets of lens optical systems 20A and 20B and the imaging elements 22A and 22B described above. The imaging element 22 is controlled by a control command from the CPU (Central Processing Unit) 130 in the processor 100. Details of the CPU 130 will be described later.

[0014] Processor 100 includes an ISP (Image Signal Processor) 108, a DMAC (Direct Memory Access Controller) 110, and an arbiter (ARBMEMC) 112 for mediating memory access. Further, processor 100 includes a MEMC (Memory Controller) 114 for controlling memory access, a distortion correction / image composition block 118, and a face detection block 119. ISP108A and 108B each perform automatic exposure (AE) control, white balance setting, and gamma setting on the images input through the signal processing of image sensors 22A and 22B. In FIG. 2, two ISPs 108A and 108B are provided corresponding to the two image sensors 22A and 22B, but it is not particularly limited, and one ISP may be provided for the two image sensors 22A and 22B.

[0015] An SDRAM (Synchronous Dynamic Random Access Memory) 116 is connected to MEMC 114. And data is temporarily stored in SDRAM 116 when being processed in ISP108A, 108B, and the distortion correction / image composition block 118. The distortion correction / image composition block 118 performs distortion correction and zenith correction, etc. using the information from the motion sensor 120 on the two captured images obtained from the two sets of the lens optical system 20 and the image sensor 22, and composes the corrected images. The motion sensor 120 may include a three-axis acceleration sensor, a three-axis angular velocity sensor, a geomagnetic sensor, and the like. The face detection block 119 performs face detection on the image and identifies the position of the person's face. Note that, instead of the face detection block 119, an object recognition block for recognizing other subjects such as a full body image of a person, a face of an animal such as a cat or a dog, a car or a flower may be provided.

[0016] Processor 100 further includes DMAC 122, image processing block 124, CPU 130, image data transfer section 126, SDRAMC (SDRAM Controller) 128, memory card control block 140, USB (Universal Serial Bus) block 146, peripheral block 150, audio unit 152, serial block 158, LCD driver 162, and bridge 168.

[0017] CPU 130 controls the operations of each part of the omnidirectional imaging device 10. Image processing block 124 performs various image processes on the image data. The processor 100 is provided with a resize block 132, which is a block for enlarging or reducing the size of the image data by interpolation processing. The processor 100 is also provided with a still image compression block 134, which is a codec block for performing still image compression and decompression such as JPEG (Joint Photographic Experts Group) and TIFF (Tagged Image File Format). The still image compression block 134 is used to generate still image data of the generated omnidirectional image. The processor 100 is further provided with a video compression block 136, which is a codec block for performing video compression and decompression such as MPEG (Moving Picture Experts Group)-4 AVC (Advanced Video Coding) / H.264. The video compression block 136 is used to generate video data of the generated omnidirectional image. Also, the processor 100 is provided with a power controller 137.

[0018] The image data transfer unit 126 transfers the image processed by the image processing block 124. The SDRAMC 128 controls the SDRAM 138 connected to the processor 100, and the image data is temporarily stored in the SDRAM 138 when various processes are performed on the image data within the processor 100. The memory card control block 140 controls the reading and writing of the memory card inserted into the memory card slot 142 and the flash ROM (Read Only Memory) 144. The memory card slot 142 is a slot for detachably mounting a memory card on the omnidirectional imaging device 10. The USB block 146 controls USB communication with external devices such as a personal computer connected via the USB connector 148. A power switch 166 is connected to the peripheral block 150.

[0019] The audio unit 152 is connected to a microphone 156 through which the user inputs an audio signal and a speaker 154 that outputs the recorded audio signal, and controls audio input and output. The serial block 158 controls serial communication with external devices such as a personal computer, and a wireless NIC (Network Interface Card) 160 is connected thereto. The LCD (Liquid Crystal Display) driver 162 is a drive circuit that drives the LCD monitor 164 and converts it into a signal for displaying various states on the LCD monitor 164. In addition to those shown in FIG. 2, a video interface such as HDMI (High-Definition Multimedia Interface, registered trademark) may be provided.

[0020] The flash ROM 144 stores a control program and various parameters described in a code that can be decoded by the CPU 130. When the power is turned on by operating the power switch 166, the above control program is loaded into the main memory, and the CPU 130 controls the operations of each part of the apparatus according to the program read into the main memory. At the same time, data necessary for control is temporarily stored in the SDRAM 138 and a local SRAM (Static Random Access Memory) (not shown). By using a rewritable flash ROM 144, it becomes possible to change the control program and parameters for control, and the function can be easily upgraded.

[0021] FIG. 3 is a diagram for explaining the overall flow of image processing in the omnidirectional imaging apparatus 10 according to the present embodiment, and shows main functional blocks. As shown in FIG. 3, images are captured by each of the image sensors 22A and 22B under predetermined exposure condition parameters. Subsequently, the images output from each of the image sensors 22A and 22B are subjected to the processing of the first image signal processing (processing 1) by the ISP 108A and 108B shown in FIG. 2. As the processing of the first image signal processing, an optical black (OB) correction process, a defective pixel correction process, a linear correction process, a shading correction process, and a region division average process are executed, and the results are stored in the memory.

[0022] When the processing of the first image signal processing (ISP1) is completed, subsequently, the second image signal processing (processing 2) is performed by the ISP 108A and 108B. As the second image signal processing, a white balance (WB (White Balance) gain) process 176, a gamma (γ) correction process, a Bayer interpolation process, a YUV conversion process, an edge enhancement (YCFLT) process, and a color correction process are executed, and the results are stored in the memory.

[0023] For the Bayer RAW image output from the imaging device 22A, the ISP 108A performs the first image signal processing, and the image is stored in the memory. Similarly, for the Bayer RAW image output from the imaging device 22B, the ISP 108B performs the first image signal processing, and the image is stored in the memory.

[0024] Note that each imaging device 22A and 22B may be set to proper exposure using the area integration value obtained by the region division averaging process so that the brightness of the image boundary portion of the images of both eyes matches (compound eye AE). Also, the imaging device 22 may have an independent simple AE processing function, and each of the imaging device 22A and the imaging device 22B may be individually set to proper exposure.

[0025] The data for which the second image signal processing has been completed is subjected to distortion correction and synthesis processing by the distortion correction / image synthesis block 118, and a full-sphere image is generated. During the process of the distortion correction and synthesis processing, information from the motion sensor 120 is appropriately obtained, and zenith correction and rotation correction are performed. When storing the captured image, if the image is a still image, it is appropriately JPEG-compressed by the still image compression block 134 shown in FIG. 2, stored in the memory, and file storage (tagging) is performed. If it is a moving image, the image is appropriately compressed into a moving image format such as MPEG-4 AVC / H.264 by the moving image compression block 136 shown in FIG. 2, stored in the memory, and file storage (tagging) is performed. Further, the data may be stored in a medium such as an SD card. When transferring to an information processing device 50 such as a smartphone, the transfer is performed using wireless LAN (Wi-Fi) or Bluetooth (registered trademark), etc.

[0026] Hereinafter, with reference to FIG. 4, the generation of the omnidirectional image and the generated omnidirectional image will be described. FIG. 4(A) illustrates the data structure of each image and the data flow of the images in the generation of the omnidirectional image. First, the images directly captured by each of the image sensors 22A and 22B are images that capture approximately a hemisphere of the omnidirectional sphere within the field of view. The light incident on the lens optical system 20 is imaged on the light-receiving area of the image sensor 22 according to a predetermined projection method. The image captured here is captured by a two-dimensional image sensor whose light-receiving area forms a planar area, and becomes image data expressed in a planar coordinate system. Also, typically, the obtained image is configured as a fisheye image including the entire image circle on which each shooting range is projected, as shown by "partial image A" and "partial image B" in FIG. 4(A).

[0027] The plurality of partial images captured by these plurality of image sensors 22A and 22B are subjected to distortion correction and compositing processing to form one omnidirectional image. In the compositing process, first, each image including complementary hemisphere portions is generated from each partial image configured as a planar image. Then, the two images including the hemisphere portions are aligned (stitched) based on the matching of the overlapping regions, and the images are composited to generate an omnidirectional image including the entire omnidirectional sphere. Although the images of the hemisphere portions include overlapping regions with other images, in image compositing, blending is performed on the overlapping regions so as to form a natural seam.

[0028] FIG. 4(B) is a diagram for explaining the data structure of the image data of the omnidirectional image used in the present embodiment represented in a plane. FIG. 4(C) is a diagram for explaining the data structure of the image data of the omnidirectional image represented on a spherical surface. As shown in FIG. 4(B), the image data of the omnidirectional image is expressed as an array of pixel values with the vertical angle φ made with respect to a predetermined axis and the horizontal angle θ corresponding to the rotation angle around the predetermined axis as coordinates. The vertical angle φ ranges from 0 degrees to 180 degrees (or -90 degrees to +90 degrees), and the horizontal angle θ ranges from 0 degrees to 360 degrees (or -180 degrees to +180 degrees).

[0029] Each coordinate value (θ, φ) in the all-sky format is associated with each point on the spherical surface representing the entire azimuth centered on the shooting location, as shown in Fig. 4(C), and the entire azimuth is associated with the all-sky image. The planar coordinates of the partial image captured by the fisheye lens and the coordinates on the spherical surface of the all-sky image are associated with each other by a predetermined conversion table. The conversion table is data that has been pre-created by the manufacturer or the like according to a predetermined projection model based on the design data of each lens optical system, etc., and is data for converting the partial image into the all-sky image.

[0030] As described above, a technique is known in which deep learning is used to detect a person's posture, gesture, or facial expression from an image, and automatic shooting is performed when a predetermined posture, gesture, or facial expression is detected.

[0031] The imaging control based on the above-described person's posture, gesture, and facial expression detects the state of the person who is the subject by image processing, and the person who is the subject (detection target) often occupies a certain proportion within the screen. However, in applications where the imaging device is remotely operated by a specific posture or posture change so as not to be captured in the screen as much as possible, the proportion occupied by the operator in the screen tends to be small. That is, the operator enters the screen for remote operation, and it can be said that the operator is not the main subject.

[0032] When detecting the above-described state of a person, in order to shorten the detection time, it is common to perform the detection process by reducing the number of pixels for the detection process by reducing the image. However, there is a trade-off relationship between the processing accuracy and the detection time, and the detection accuracy decreases when the detection operation is performed with the image reduced. Furthermore, when the proportion of the screen occupied by the operator who is the detection target is small, since the number of pixels occupied by the detection target is small, if the entire image is reduced, the number of pixels becomes even smaller, and the detection accuracy of the posture change deteriorates.

[0033] Furthermore, the omnidirectional imaging device 10 as described above also has difficulties specific to omnidirectional images. For example, in a fisheye image, the distortion of the imaged person is large, and when a person is at the boundary between multiple fisheye images (a state where the person straddles between multiple fisheye images), the pose detection accuracy decreases. By converting multiple fisheye images into an Equirectangular image at once, stitching them together to form an omnidirectional image, and performing detection on the omnidirectional image, the distortion near the equator can be reduced, and the person is not divided between multiple images, thus expecting an improvement in the person's pose detection accuracy. However, as will be described below, there are still difficulties involved.

[0034] More specifically, even if the distance from the imaging device to the operator is the same, the wider the lens in the imaging device, the smaller the proportion of the person occupying the screen. In an omnidirectional image with a horizontal field of view of 360 degrees, in particular, the proportion of the person occupying the screen is small. In an omnidirectional image, since two fisheye images are joined, there may be a case where it is possible to prevent the person from being divided to a certain extent when the person is at the boundary between multiple fisheye images. However, since the omnidirectional image circulates 360 degrees horizontally, it is cut off at the peripheral part or the end. If the person is divided at this image end, the pose detection accuracy of the person decreases. In particular, in a configuration where two imaging units are provided front and back as shown in FIG. 1 and the front of the lens is associated with the central part of the omnidirectional image, when the photographer himself / herself does not want to be captured in the screen, he / she often remotely operates from the side of the imaging device so as not to be located at the center of the image. In such a case, the person is divided at both ends of the omnidirectional image.

[0035] In view of the above points, the omnidirectional imaging device 10 according to the present embodiment acquires an image, first identifies the position of the person area from the acquired image, and sets a detection target area based on the identified position of the person area. Then, the state of the person is detected based on the set detection target area in the image. With the above configuration, instead of performing detection processing on the entire image, the area where a person is present within the screen is identified, and detection processing is performed within a limited range based on this identified area where the person is present, thereby reducing the number of pixels for which detection processing is performed and shortening the detection time. Since the image is not reduced, the number of pixels of the person within the screen remains unchanged, preventing deterioration of detection accuracy. As a result, it becomes possible to achieve both shortening of the detection time and detection accuracy when detecting the state of a person from within the screen.

[0036] In a more preferred embodiment, corresponding to the omnidirectional image circulating in at least one direction, in at least the detection target area, the end area of the omnidirectional image can be replicated so that one end area of the omnidirectional image is connected to the other end. With the configuration of the above preferred embodiment, even if the detection target is divided at both ends when the omnidirectional image is an equirectangular image, it becomes possible to accurately detect a person.

[0037] Hereinafter, with reference to FIGS. 5 to 8, the imaging control based on the posture detection of a person executed by the omnidirectional imaging device 10 according to the first embodiment will be described in more detail.

[0038] FIG. 5 is a functional block diagram for realizing imaging control based on the posture detection of a person according to the first embodiment. The functional block 200 shown in FIG. 5 includes an omnidirectional image generation unit 210, an image acquisition unit 220, a replication unit 230, a position identification unit 240, a region setting unit 250, a posture detection unit 260, and an imaging control unit 270.

[0039] The omnidirectional image generation unit 210 generates an omnidirectional image (equirectangular image) captured by the imaging device 22 and synthesized by the distortion correction / image synthesis block 118. Note that the imaging control based on the detection of the human posture may be the control before the actual shooting before pressing the shutter button. However, in the embodiment to be described, even at the stage before the actual shooting, the conversion from the fisheye image to the omnidirectional image is performed, and it should be noted that the omnidirectional image after this conversion is the processing target of the posture detection.

[0040] The image acquisition unit 220 acquires the image to be processed. In the omnidirectional imaging device 10, the acquired image is an image having a viewing angle of 360 degrees at least in the first direction. More specifically, it is an omnidirectional image of 360 degrees in the horizontal direction and 180 degrees in the vertical direction (including 360 degrees in the horizontal direction, so 360 degrees in the horizontal and 360 degrees in the vertical in total).

[0041] Although the omnidirectional image is an image that circulates in the horizontal direction as the shooting range, as image data, it is a single image with a predetermined horizontal position as the end. If a person is located at this end, the area including the person may be separated, which may affect the accuracy of the posture detection. Therefore, in order to address this discontinuity at the image end in the present embodiment, the replication unit 230 replicates this end region so that one end region of the omnidirectional image is connected to the other end, and adds this replication to the other end. The replication unit 230 performs the replication at a stage before the process of specifying the position of the person region by the position specifying unit 240 described later.

[0042] The position specifying unit 240 specifies the position of the person area from the acquired image. Any technique can be provided for the position of the person area, and known lightweight person detection, face detection, etc. can be applied. As described above, in the present embodiment, the position specifying unit 240 specifies the position of the person area based on the modified image obtained by adding a copy of one end area of the omnidirectional image to the other end of the omnidirectional image. Note that, in the embodiment to be described, the position of the person area is detected by performing person detection based on the acquired image or by performing face detection based on the acquired image. However, in other embodiments, it can also be detected by performing moving object detection based on the difference between images of a plurality of continuously acquired frames.

[0043] The area setting unit 250 sets a detection target area for the omnidirectional image based on the specified position of the person area. The setting of the detection target area may be set for a part of the omnidirectional image, and the detection process described later may be performed on this part. Alternatively, an image of a part corresponding to the detection target area in the omnidirectional image may be copied, and the detection process may be executed on this copied data.

[0044] The posture detection unit 260 uses the set detection target area as a processing target and detects the posture of the person based on the image features of the detection target area in the omnidirectional image. At this time, preferably, the detection target area is not reduced. Alternatively, even if trimming, white painting or black painting, or reduction is performed to match the input layer of the deep learning model used by the posture detection unit 260, since it is not the entire omnidirectional image that is reduced but only a limited part of the detection target area, a decrease in the number of pixels can be suppressed. The posture detection unit 260 preferably can include skeleton detection of a person based on the acquired image and posture detection based on the detected skeleton. A deep learning model can be used for skeleton detection and posture detection.

[0045] The imaging control unit 270 controls the imaging body 12 based on the posture of the person detected by the above-described processing. More specifically, it detects a specific posture and performs control related to the functions of the camera, such as taking a picture (shutter release), setting a timer, or changing the shooting parameters or modes, according to the detected posture.

[0046] FIG. 6 is a flowchart showing imaging control based on human posture detection according to the first embodiment.

[0047] The processing shown in FIG. 6 is executed for each frame in response to the start of generation of an image frame due to the activation of the omnidirectional imaging device 10 or the activation of the imaging control function based on posture detection. Note that FIG. 6 shows a series of processes from detecting a person to detecting a predetermined posture and releasing the shutter, and will be described as being performed for each frame of the equirectangular format omnidirectional image output from the image sensors 22A and 22B through the ISPs 108A and 108B to the SDRAM. However, it is not particularly limited, and in other embodiments, it may be performed at regular frame intervals.

[0048] In step S101, the processor acquires, by the image acquisition unit 220, one frame of the omnidirectional image generated by the omnidirectional image generation unit 210. In step S102, the processor generates, by the duplication unit 230, a modified image in which the end region of the omnidirectional image is duplicated so that one end region is connected to the other end, and the duplication is added to the other end of the image.

[0049] FIG. 7 illustrates a process of duplicating an end region of an omnidirectional image so as to connect it to the other end in the omnidirectional imaging apparatus according to the present embodiment. In step S102 shown in FIG. 6, as shown in FIGS. 7(A) and 7(B), a modified omnidirectional image is created by duplicating one end region T of the omnidirectional image to the other end S as T'. For simplicity, the size to be duplicated is fixed. When the resolution of the image is 3840x1920 (AxB), for example, for the original omnidirectional image, 384x1920 (CxB), which is 10% of the horizontal image size from the left end, is fixedly duplicated to the right end to generate a modified omnidirectional image of 4224x1920 (D×B).

[0050] In step S103, the processor specifies the position of the person region from the modified omnidirectional image by the position specifying unit 240. As a result, coordinates (px, py) and size (height H and width W) representing a rectangular area (detection frame) R including the detected person P as shown in FIG. 7(C) are output. Here, when multiple persons are detected, it is assumed that coordinates and sizes are output for the number of detected persons. Note that any person detection algorithm can be used for the process of specifying the position of the person region in step S103. For example, techniques such as SVM (Support Vector Machine) or AdaBoost using a statistical learning method can be used. These techniques are generally lighter than, for example, pose detection processing. Note that in the described embodiment, it is described that person detection is performed, but face detection may be performed instead.

[0051] Also, in some cases, the same subject may be detected in both the original image region on the left end and the duplicated image region on the right end in FIG. 7. In that case, it may be appropriate to use the detection result in the region including the duplicated image region, or alternatively, to include the results detected in both the original image region and the duplicated image region.

[0052] In step S104, the processor determines whether a person has been detected. If it is determined in step S104 that no person has been detected (NO), the process branches to step S112 and the processing for the frame ends. On the other hand, if it is determined in step S104 that at least one person has been detected (YES), the process proceeds to step S105. In step S105, the processor sets N to the initial value 0, sets the number of detected persons to NMAX, and in steps S106 to S110, repeats the processing for each person for the number of persons identified. Note that the order of processing the detected persons may prioritize those with a larger area of the rectangular area of the detected persons.

[0053] In step S106, the processor sets a detection target area based on the position of the identified person area by the area setting unit 250. Here, the detection frame (position (px, py), size (W, H)) in person detection or face detection may be set as the detection target area as it is, or it may be set as an area with a predetermined margin added thereto.

[0054] In step S107, the processor detects the state of the person, more specifically the posture of the person, based on the detection target area in the image by the posture detection unit 260. Posture detection is performed on the image obtained by cutting out the coordinate part of the set detection target area. For example, when the coordinates (px, py) = (3200, 460) and the size is W = 180 and H = 800, a rectangular range of 800x180 from (3200, 460) to (3380, 1260) is cut out. The posture detection unit 260 performs skeleton detection on the cut-out part. Note that the skeleton detection by the posture detection unit 260 may be performed only on the part of the set detection target area instead of on the cut-out image. The skeleton detection may be executed by a neural network learned by deep learning. With the image as the input, the coordinates of the body parts of the person are output. The skeleton detection may, for example, detect the body parts of the person as 18 parts numbered from 0 to 17 as shown in FIG. 8, and the positions (x coordinate, y coordinate) of each body part are output. For example, numbers 4 and 7 represent the positions of the wrists, and numbers 14 and 15 represent the positions of the eyes.

[0055] In step S108, the processor determines whether or not a predetermined posture has been detected by the posture detection unit 260. In step S108, based on the coordinates of the skeleton detection result, it is determined whether or not it is a predetermined posture. Table 1 exemplifies the relationship between the predetermined posture and camera control.

[0056]

Table 1

[0057] As exemplified in Table 1, on the condition that the Y coordinate of 4 (right wrist) or 7 (left wrist) of the detected skeleton is above the Y coordinate of 14 (right eye) or 15 (left eye), it is determined that the position of the wrist is above the eye, and it is determined as a predetermined posture, and an operation (shooting) of pressing the shutter can be performed at S111. When a plurality of postures are determined, priorities are set as exemplified in Table 1, and the one with the higher priority can be adopted as the determination result. In the embodiment to be described, although it is described as detecting a static posture, camera control may be performed by detecting a dynamic change of the posture including the time series of the posture.

[0058] In step S108, if it is determined that a predetermined posture has not been detected yet (NO), the process branches to step S109. In step S109, N is incremented, and in step S110, it is determined whether N has reached NMAX. In step S110, if it is determined that N has reached NMAX (YES), the process branches to step S112, and the process for the frame ends. On the other hand, if it is determined that N has not reached NMAX yet (NO), the process loops back to step S106 to continue the process for the remaining detected persons.

[0059] When returning to step S108 again, if it is determined in step S108 that a predetermined posture has been detected (YES), the process branches to step S111. In step S111, the processor performs camera control corresponding to the conditions exemplified in Table 1 by the imaging control unit 270, and the process for the frame ends in step S112. For example, as an operation of pressing the shutter, the captured image can be recorded as a file.

[0060] According to the embodiment described above, it is possible to achieve both shortening of the detection time and detection accuracy when detecting the posture of a person from within the screen. In particular, since a limited detection target range is set, it is possible to accurately and quickly detect the posture of a person far away.

[0061] Next, with reference to FIGS. 9 to 13, imaging control based on human posture detection performed by the omnidirectional imaging device 10 according to the second embodiment will be described. In the above-described first embodiment, in the stages of detecting the human region and specifying the position of the human region, a process of adding a copy of one end region of the omnidirectional image to the other end is performed, and based on this modified image, the position of the human region is specified and the posture of the human is detected. On the other hand, in the second embodiment described below, before the process of copying and adding the end region, the position of the human region is specified, a detection target region corresponding to the position and size of the human region is set, and only when necessary, a process of adding a copy of one end region of the omnidirectional image to the other end is performed to detect the posture of the human.

[0062] FIG. 9 is a functional block diagram for realizing imaging control based on human posture detection according to the second embodiment. The functional block 300 shown in FIG. 9 includes an omnidirectional image generation unit 310, an image acquisition unit 320, a position specification unit 330, a size determination unit 340, a necessity determination unit 350, a copying unit 360, a region setting unit 370, a posture detection unit 380, and an imaging control unit 390. Note that the functional blocks shown in FIG. 9 have the same or similar functions as the functional blocks with the same names shown in FIG. 5, and detailed descriptions thereof will be omitted unless otherwise specified.

[0063] The omnidirectional image generation unit 310 generates an omnidirectional image captured by the imaging device 22 and synthesized by the distortion correction / image synthesis block 118. Also in this embodiment, conversion from a fisheye image to an omnidirectional image is performed even at the stage before actual shooting. The image acquisition unit 320 acquires an image to be processed.

[0064] The position specifying unit 330 specifies the position of the person area from the acquired image. Any technique can be provided for specifying the position of the person area, and known lightweight person detection, face detection, moving object detection, etc. can be applied. In the second embodiment, the position specifying unit 330 specifies the position of the person area based on the original omnidirectional image. Therefore, in the second embodiment, it is preferable to perform moving object detection based on the difference between images of a plurality of frames continuously acquired to specify the position of the person area. This is because even when a person is located at the boundary, the person area can be preferably detected even from a part of the person.

[0065] The sizing unit 340 determines the position and size of the detection target area to be set based on the size of the specified person area. Here, it can be set as an area obtained by adding a predetermined margin to the detection frame (position (px, py), size (W, H)) in moving object detection.

[0066] The necessity determination unit 350 determines whether it is necessary to perform replication based on the position and size of the detection target area to be set. Depending on the position and size of the detection target area, the detection target area to be set may protrude from the range of the omnidirectional image. The necessity determination unit 350 determines whether the detection target area to be set protrudes from the range of the omnidirectional image based on the position and size of the detection target area to be set, and determines that replication is necessary if it protrudes.

[0067] Similar to the first embodiment, the replication unit 360 performs processing of replicating this end region so that one end region of the omnidirectional image is connected to the other end at least in the detection target area to be set in order to cope with this discontinuity at the image end, and adding this replication to the other end. The replication unit 230 according to the second embodiment performs replication at a stage after the process of specifying the position of the person area by the above-described position specifying unit 330, but performs replication only when it is determined by the necessity determination unit 350 that replication is necessary.

[0068] The area setting unit 370 sets a detection target area for the omnidirectional image based on the position of the specified person area. The setting of the detection target area may be set for a part of the omnidirectional image, and the detection process described later may be performed on this part. Alternatively, an image of a part corresponding to the detection target area in the omnidirectional image may be separately copied (cut out) for the detection process, and this copied data may be the target of the detection process. Also, when copying an image of a corresponding part for the detection process, the above-mentioned copying may be performed by adding a copy of one end area of the omnidirectional image to the other end as in the first embodiment, and copying an image of a part corresponding to the detection target area from the changed image. Or, after separately copying (cutting out) the part of the omnidirectional image included in the detection target area, a process of copying and adding from the other end area of the omnidirectional image only for the insufficient part of the detection target area may be performed.

[0069] The posture detection unit 260 uses the set detection target area as a processing target and detects the posture of a person based on the image features of the detection target area in the omnidirectional image. At this time, preferably, the detection target area is not reduced. Alternatively, even if trimming, white painting or black painting, or reduction is performed to conform to the input layer of the deep learning model used by the posture detection unit 260, the whole omnidirectional image is not reduced, but only a limited part of the detection target area is reduced. Therefore, a decrease in the number of pixels can be suppressed.

[0070] The imaging control unit 270 controls the imaging body 12 based on the posture of the person detected by the above processing.

[0071] FIG. 10 is a flowchart showing imaging control based on the posture detection of a person according to the second embodiment.

[0072] The process shown in FIG. 10 is executed for each frame by starting the omnidirectional imaging device 10 or starting the imaging control function based on the posture detection. Similar to the first embodiment, the process shown in FIG. 10 may be performed for each frame or at a certain frame interval.

[0073] In step S201, the processor acquires, via the image acquisition unit 320, a full-sphere image for one frame generated by the full-sphere image generation unit 210. In step S202, the processor specifies, via the position specification unit 330, the position and size of the person region from the full-sphere image by detecting a moving object.

[0074] FIG. 11 is a diagram for explaining a process of detecting a difference between frames and a process of determining whether replication is necessary in the full-sphere imaging device according to the second embodiment. FIGS. 11(A) and (B) schematically show two consecutive frames, and coordinates (px, py) representing a rectangular area and size (height H, width W) of the moving object part M as a person region including a person P are output based on the difference between the previous frame and the current frame.

[0075] Referring to FIG. 10 again, in step S203, the processor determines whether a person has been detected. If it is determined in step S203 that no person has been detected (NO), the process branches to step S214 and the process for the frame ends. On the other hand, if it is determined in step S203 that at least one person has been detected (YES), the process proceeds to step S204. In step S204, the processor sets N to an initial value of 0, sets NMAX to the number of detected persons (moving objects), and in steps S205 to S212, repeats the process for each person for the number of specified persons.

[0076] In steps S205 to S208, the moving object part is set as the person region, the detection frame is expanded according to the size of the person region, the detection target range is set within the range of the detection frame, and at that time, if necessary, one end region of the full-sphere image is replicated and added to the other end.

[0077] More specifically, in step S205, the processor determines, via the size determination unit 340, the position and size of the detection target region to be set based on the position and size of the specified person region.

[0078] FIG. 12 is a diagram illustrating how to expand the detection target area in the omnidirectional imaging device according to the second embodiment. In FIG. 12, assume that the result detected in step S202 is the detection frame dct_box2 with height H and width W. In that case, the detection target area can be set by expanding the width of the original output result by 50% (W / 2) on each side left and right, and also expanding the height by 50% (H / 2) on each side up and down with respect to the height H, within the range of dct_box2. Note that when reaching the upper or lower end when expanding in the height direction, it can be limited to the upper limit (Y = 0) or the lower limit (for example, Y = 1920).

[0079] In step S206, the processor determines whether replication needs to be performed based on the position and size of the detection target area (which may be the detection frame if the detection target area is fixedly determined with respect to the detection frame). In a specific embodiment, whether replication is required can be determined based on whether the detected target area exceeds the edge of the image, from the position and size of the detected target area determined according to the size of the specific person area located at the edge of the screen.

[0080] Specifically, first, it can be determined whether the person area is located at the edge of the screen. FIG. 11(C) illustrates the process of determining whether the person area is located at the edge of the screen. As shown in FIG. 11(C), boundaries for determining whether it is located at the edge of the screen are set on the left and right in the omnidirectional image. The left boundary is le_xrange, and the right boundary is re_xrange. When setting the boundaries at 10% of the horizontal size of 3840x1920, the widths le_xrange and re_xrange up to the left and right boundaries are 384. In the example described above, if px or px + w exists in the range of the x - coordinate from 0 to le_xrange or from 1920 - re_xrange to 1920, it is determined that the person area is located at the edge of the screen. When it is determined that the person area is located at the edge of the screen, further, whether replication is required is determined according to whether the detection target area exceeds the edge of the omnidirectional image when the size of the detection target area is set according to the size of the person area.

[0081] In step S206, if it is determined that duplication is necessary (YES), the process branches to step S207. In step S207, the processor causes the duplication unit 360 to duplicate the end region in the detection target region so that one end region of the omnidirectional image is connected to the other end, and generates a modified image with the duplication added to the other end of the image. In step S208, the processor sets the detection target region by means of the region setting unit 370.

[0082] As described above, when separately duplicating (cropping) a corresponding part of the image for detection processing, the above duplication may perform a process of adding an image duplicated from the other end region only to the insufficient part of the detection target region after duplicating the part of the omnidirectional image included in the detection target region. FIG. 13 illustrates the process of duplicating the part of the omnidirectional image included in the detection target region and then adding and duplicating only the insufficient part of the detection target region from the other end region.

[0083] When expanding the person region dct_box1 detected at the left end as shown in FIG. 13 by 50% in the vertical and horizontal directions, the left side will exceed the image end. Therefore, a region of horizontal W / 2 - px and vertical 2H, which is the excess from the right end, is duplicated and added to the trimming region.

[0084] In step S209, the processor detects the state of the person, more specifically the posture of the person, based on the detection target region in the image by means of the posture detection unit 380. In step S210, the processor determines whether a predetermined posture has been detected by means of the posture detection unit 260. In step S210, it is determined whether it is a predetermined posture based on the coordinates of the skeleton detection result.

[0085] In step S210, if it is determined that a predetermined posture has not been detected yet (NO), the process branches to step S211. In step S211, N is incremented, and in step S212, it is determined whether N has reached NMAX. If it is determined in step S212 that N has reached NMAX (YES), the process branches to step S214 and the processing for the frame ends. On the other hand, if it is determined in step S212 that N has not reached NMAX yet (NO), the process loops back to step S205 to continue the processing for the remaining detected persons.

[0086] When returning to step S210 again, if it is determined in step S210 that a predetermined posture has been detected (YES), the process branches to step S213. In step S213, the processor performs camera control corresponding to the conditions exemplified in Table 1, for example, by the imaging control unit 390, and the processing for the frame ends in step S214. For example, as an operation of taking a shutter, the captured image can be recorded as a file.

[0087] In the second embodiment, at the end of the posture detection process, the current frame is saved for use in moving object detection in the next frame.

[0088] According to the second embodiment, it is possible to achieve both shortening of the detection time and detection accuracy when detecting the posture of a person from within the screen. In particular, since a limited detection target range is set, it is possible to accurately and quickly detect the posture of a person far away. In the second embodiment, in particular, the image at the end is not copied for each frame, and the copy is performed only when the detection target area straddles the end of the image. Therefore, an effect of reducing the processing time is expected even when compared with the first embodiment.

[0089] According to the embodiments described above, it is possible to achieve both shortening the detection time and improving the detection accuracy when detecting the state of a person within a screen. In particular, even when the proportion of the screen occupied by the person is small, detection can be performed in a short detection time. Further, with the above configuration, since detection can be performed in a short detection time even when the proportion of the screen occupied by the person is small, it becomes possible to suitably apply remote operation of the imaging device based on a specific posture or posture change.

[0090] In addition, in each of the above-described embodiments, an equirectangular image is described as a specific example. Since the above-described embodiments also have parts specific to the omnidirectional image, they can be suitably used when the target is an equirectangular image, but are not particularly limited. In other embodiments, the image to be processed is not limited to an equirectangular image. Further, as the state of the person to be detected, a posture based on the detection of the person's skeleton is exemplified, but it is not limited thereto. It is not limited to the case of detecting the posture from the whole body of the person, and it may be possible to detect the expression of the person's face (movement of eyes and mouth) and the state of the body part of the person (sign using hands).

[0091] Each function of the embodiments described above can be realized by one or more processing circuits. Here, the "processing circuit" in the present embodiment refers to a processor programmed to execute each function by software, such as a processor implemented by an electronic circuit, an ASIC (Application Specific Integrated Circuit) designed to execute each function described above, a DSP (digital signal processor), an FPGA (field programmable gate array), an SOC (System on a chip), a GPU, and devices such as conventional circuit modules.

[0092] Furthermore, the above functions can be realized by a computer-executable program written in legacy programming languages such as assembler, C, C++, C#, Java (registered trademark), and object-oriented programming languages, and can be stored in a device-readable recording medium such as ROM, EEPROM, EPROM, flash memory, flexible disk, CD-ROM, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, Blu-ray disc, SD card, MO, etc., or distributed through an electric communication line.

[0093] So far, the image processing apparatus, image processing system, image processing method, and program according to an embodiment of the present invention have been described. However, the present invention is not limited to the above-described embodiment, and can be changed within the scope that those skilled in the art can conceive, such as addition, change, or deletion of other embodiments. As long as the effects of the present invention can be achieved in any aspect, it is included in the scope of the present invention.

Explanation of Signs

[0094] 10…Omnidirectional imaging device, 12…Imaging body, 14…Housing, 18…Shooting button, 20…Lens optical system, 22…Image sensor, 100…Processor, 102…Lens barrel unit, 108…ISP, 110, 122…DMAC, 112…Arbiter (ARBMEMC), 114…MEMC, 116, 138…SDRAM, 118…Distortion correction and image synthesis block, 119…Face detection block, 120…Motion sensor, 124…Image processing block, 126…Image data transfer section, 128…SDRAMC, 130…CPU, 132…Resize block, 134…Still image compression block, 136…Video compression block, 140…Memory card control block, 142…Memory card slot, 144…Flash ROM, 146…USB block, 148…USB connector, 150…Peripheral block, 152…Audio unit, 154…Speaker, 156…Microphone, 158…Serial block, 160…Wireless NIC, 162…LCD driver, 164…LCD monitor, 166…Power switch, 168…Bridge, 200, 300…Function block, 210, 310…Omnidirectional image generation section, 220, 320…Image acquisition section, 230, 360…Duplication section, 240, 330…Position determination section, 250, 370…Region setting section, 260, 380…Posture detection section 260, 270, 390…Imaging control section, 340…Size determination section, 350…Necessity determination section

Prior art documents

Patent documents

[0095]

Patent Document 1

Patent Document 2

Claims

1. An image processing method, comprising steps of: a computer acquiring an image that circulates in at least one direction; identifying the position and size of a human region from the acquired image; determining the position and size of a detection target region based on the identified position and size of the human region; determining whether the detection target region is located at an edge of the image based on the position and size of the detection target region; when it is determined that the detection target region is located at an edge of the image, replicating at least the edge region of the image such that one edge region of the image is connected to the other edge in at least the detection target region; detecting the state of a person based on the image obtained by adding the replication of one edge region of the image to the other edge of the image and the detection target region An image processing method that performs the above steps.

2. The image processing method according to claim 1, wherein the acquired image has a viewing angle of 360 degrees in at least a first direction.

3. The image processing method according to claim 1 or 2, wherein the detection includes skeleton detection of the person based on the acquired image and posture detection based on the detected skeleton.

4. The step of identifying the position of the human region is characterized by identifying the position of the human region by performing human detection based on the acquired image, performing face detection based on the acquired image, or performing moving object detection based on the difference between a plurality of consecutively acquired frame images. The image processing method according to any one of claims 1 to 3.

5. An imaging control method including the image processing method according to any one of claims 1 to 4, wherein the computer controls a device equipped with imaging means, and the computer performs the step of executing the image processing method; performs the step of controlling the imaging means based on the detected state of the person An imaging control method that performs the above steps.

6. A program for causing a computer to execute the method according to any one of claims 1 to 5.

7. An image acquisition unit that acquires an image that circulates in at least one direction, a position identification unit that identifies the position and size of a human region from the acquired image, a determination unit that determines the position and size of a detection target region based on the identified position and size of the human region, A determination unit that determines whether or not the detection target area is located at an edge of the image based on the position and size of the detection target area; A duplication unit that duplicates the end region so that at least in the detection target area, one end region of the image is connected to the other end when it is determined that the detection target area is located at an edge of the image; A detection unit that detects the state of a person based on the image obtained by adding the duplication of one end region of the image to the other end of the image and the detection target area An image processing apparatus comprising: **Claim 8** The image processing apparatus according to claim 7, wherein the acquired image has an angle of view of 360 degrees at least in a first direction. **Claim 9** The image processing apparatus according to claim 7 or 8, wherein the detection includes skeleton detection of the person based on the acquired image and posture detection based on the detected skeleton. **Claim 10** The image processing apparatus according to any one of claims 7 to 9, Imaging means, An imaging control unit that controls the imaging means based on the detected state of the person An imaging apparatus comprising:

Citation Information

Patent Citations

  • Method for minimizing dead zones in panoramic camera and system therefor

    JP2006191535A

  • Joint position estimation device and joint position estimation program

    JP2018057596A

  • camera

    JP4227257B2

  • Information processing device and information processing system

    JP6729043B2

  • System and method for detecting human gaze and gesture in unconstrained environments

    WO2019204118A1