Image processing device, image processing method, image processing program and image processing system

The image processing apparatus addresses the challenge of understanding wide-angle videos by detecting specific features and generating flattened videos with adjusted display specifications, resulting in improved intuitive understanding and feature detection accuracy.

WO2025115350A1PCT designated stage expired Publication Date: 2025-06-05KONICA MINOLTA INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/032687
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-09-12
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Wide-angle videos obtained by wide-angle lens cameras are difficult for people to intuitively understand, making it challenging to grasp movements or identify specific features within these videos.

Method used

An image processing apparatus that detects feature portions in wide-angle videos and generates flattened videos by adjusting display specifications, such as display position and occupancy rate, to highlight specific features and improve understanding.

Benefits of technology

The solution enables the generation of videos that are intuitively easier to understand, improving the detection accuracy of specific features and allowing users to easily grasp the current state of monitored areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024032687_05062025_PF_FP_ABST
    Figure JP2024032687_05062025_PF_FP_ABST
Patent Text Reader

Abstract

An image processing device (100) comprises: a detection unit that detects, from within a wide-angle image, a feature portion in which a predetermined feature appears; a generation unit that generates, on the basis of predetermined display designation information, a planar image (N) obtained by flattening a target region that includes the feature portion in the wide-angle image; and an output unit that outputs image data indicating the planar image (N).
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device, image processing method, image processing program, and image processing system

[0001] The present disclosure relates to image processing technology, and more particularly to technology for generating an image that is easy for people to intuitively understand from a wide-angle image.

[0002] There is known a technology for capturing images of a predetermined space with a camera and understanding the movement of people or objects based on the captured images. For example, Japanese Patent Laid-Open Publication No. 2007-158680 (Patent Document 1) discloses a system for capturing images of a lecture and extracting from the images portions that show the lecturer and lecture information (such as letters and / or diagrams written on a blackboard).

[0003] International Publication No. 2001 / 072034 (Patent Document 2) discloses a technology for controlling the directionality of a camera using image and audio information. Japanese Patent Application Laid-Open No. 2015-70503 (Patent Document 3) discloses a device for filming a sport such as soccer and determining events that occur during the sport based on the video of the sport and sounds obtained from a microphone.

[0004] Japanese Patent Application Laid-Open Publication No. 2019-106631 (Patent Document 4) discloses a device that photographs an airport arrival lobby, etc., detects pairs of suspicious individuals who have made highly suspicious transactions based on the photographed footage, and displays images of each person in the pair.

[0005] JP 2007-158680 A International Publication No. 2001 / 072034 JP 2015-70503 A JP 2019-106631 A

[0006] In Patent Document 1 or Patent Document 4, a characteristic portion of a captured image that shows a predetermined characteristic is cut out as a part of the whole image. Generally, wide-angle images captured by wide-angle lens cameras are images that are difficult for people to intuitively understand. Therefore, if the captured image is a wide-angle image, an image cut out from the wide-angle image is also an image that is difficult for people to intuitively understand.

[0007] An object of the present disclosure is to generate an image that is easy for people to intuitively understand from a wide-angle image.

[0008] An image processing device according to one aspect of the present disclosure includes a detection unit that detects a characteristic portion in a wide-angle image that shows a predetermined characteristic, a generation unit that generates a flattened image obtained by flattening a target area in the wide-angle image that includes the characteristic portion based on predetermined display specification information, and an output unit that outputs video data representing the flattened image.

[0009] Preferably, the feature is a particular identity of a person, a particular action of a person, a particular identity of an object, or a particular state of an object.

[0010] Preferably, the display specification information includes a display position of the object having the characteristic in the flattened image and an occupancy rate of the object in the flattened image.

[0011] Preferably, the object is moving, and the generation unit adjusts the display position of the object in the flattened image based on the moving direction and moving speed of the object.

[0012] Preferably, the generating unit generates the flattened image so that an object having a characteristic is highlighted.

[0013] Preferably, the image processing device further includes a warning unit that instructs the warning device to output an alert when the characteristic portion is detected.

[0014] Preferably, the detection unit generates a feature detection image used to detect a feature portion from the wide-angle image. The feature detection image is an image obtained by flattening the wide-angle image.

[0015] Preferably, the generation unit converts first coordinates indicating a first display position of the object having the feature in the feature detection image into second coordinates indicating a second display position of the object in the wide-angle image, and determines the target area based on the second coordinates and the display specification information.

[0016] An image processing method according to another aspect of the present disclosure includes detecting a characteristic portion in a wide-angle image that shows a predetermined feature, flattening a target area in the wide-angle image that includes the characteristic portion to generate a flattened image based on predetermined display specification information, and outputting image data representing the flattened image.

[0017] An image processing program according to another aspect of the present disclosure causes one or more computers to execute the image processing method described above.

[0018] An image processing system according to another aspect of the present disclosure includes the image processing device described above and a camera that captures wide-angle images.

[0019] According to the present disclosure, an image that is easy for people to intuitively understand can be generated from a wide-angle image.

[0020] 1 is a diagram illustrating a first example of an image processing system according to the present embodiment. FIG. 2 is a diagram illustrating a second example of an image processing system according to the present embodiment. FIG. 3 is a diagram illustrating an example of a hardware configuration of an image processing device 100. FIG. 4 is a diagram illustrating an overview of the processing of the system 11. FIG. 5 is a diagram illustrating an example of a management table 114. FIG. 6 is a diagram illustrating the processing of a detection unit 153. FIG. 7 is a diagram illustrating another example of a method for generating a feature detection image. FIG. 8 is a diagram illustrating an example of a detection operation. FIG. 9 is a diagram illustrating the processing of a generation unit 154. FIG. 10 is a diagram illustrating an example of a flattened image when display specification information d2 is specified as the display specification information. FIG. 11 is a diagram illustrating an example of a flattened image when display specification information d3 is specified as the display specification information. FIG. 12 is a diagram illustrating another example of coordinate transformation by the generation unit 154. FIG. 13 is a diagram illustrating an example of a method for calculating a horizontal rotation angle and a vertical rotation angle by the generation unit 154. FIG. 14 is a diagram illustrating an example of a method for calculating an FOV by the generation unit 154. FIG. 15 is a diagram illustrating a first example of an alert output by a warning device 500. FIG. 16 is a diagram illustrating a second example of an alert output by the warning device 500. FIG. 17 is a flowchart illustrating an example of a processing procedure of the image processing device 100. Fig. 10 is a diagram for explaining a case where display designation information d6 or d7 is designated as display designation information. Fig. 11 is a diagram for explaining a case where a plurality of people raise their hands at the same time. Fig. 12 is a diagram for explaining a case where a plurality of people raise their hands in turn. Fig. 13 is a diagram for explaining a case where a plurality of features are designated as detection targets and the plurality of features are detected.

[0021] Hereinafter, embodiments and modifications according to the present disclosure will be described with reference to the drawings. In the following description, the same parts and components are denoted by the same reference numerals. Their names and functions are also the same. Therefore, detailed descriptions thereof will not be repeated. Note that the embodiments and modifications described below may be selectively combined as appropriate.

[0022] <A. Example of a System to Which the Technology of the Present Disclosure Can Be Applied> The configuration of an image processing system to which the technology of the present disclosure can be applied will be described with reference to FIGS. 1 and 2 .

[0023] 1 is a diagram showing a first example of an image processing system according to the present embodiment. A system 11 is the first example of an image processing system according to the present embodiment. The system 11 is used by a user Q to monitor a monitoring area 120. The user Q is, for example, an operator who monitors the monitoring area 120. The monitoring area 120 may be any area, such as a factory, port, airport, construction site, research laboratory, warehouse, office, school, stadium, museum, art gallery, library, concert venue, etc.

[0024] The system 11 includes an image processing device 100 , an input device 200 , a display device 300 , a camera 400 , and a warning device 500 .

[0025] The camera 400 has a wide-angle lens. A wide-angle lens is a lens that can capture a wider range than what the human eye can see. The wide-angle lens is, for example, a fisheye lens. The camera 400 captures an image of the monitoring area 120 and acquires a wide-angle image of the monitoring area 120. The wide-angle image is a curved image (distorted image) that is not flattened. The wide-angle image is, for example, a fisheye image. People cannot intuitively understand wide-angle images. Furthermore, wide-angle images are not suitable for image recognition processing.

[0026] The camera 400 transmits wide-angle image data representing a wide-angle image to the image processing device 100. The camera 400 is connected to the image processing device 100 via a network 15. The network 15 may include a wireless network, a wired network, a local area network (LAN), a public network, or any other network. The camera 400 may be an omnidirectional camera formed by combining multiple cameras.

[0027] The image processing device 100 receives wide-angle image data from the camera 400. The wide-angle image data is stored in the storage 103 of the image processing device 100. The image processing device 100 detects a characteristic portion showing a predetermined characteristic from the wide-angle image represented by the wide-angle image data. The predetermined characteristic is a specific identity of a person, a specific behavior of a person, a specific identity of an object, or a specific state of an object. The specific identity of a person is, for example, height, the positional relationship of skeletal features, and / or the positional relationship of facial features. The specific behavior of a person is, for example, a gesture or a gait style. The specific identity of an object is, for example, shape, color, and / or size. The specific state of an object is, for example, position, movement speed, acceleration, and / or color. FIG. 1 illustrates an example in which the predetermined characteristic is a gesture of a person raising their hand.

[0028] The image processing device 100 generates a flattened image N obtained by flattening an object region including a characteristic portion in a wide-angle image based on predetermined display specification information. The display specification information includes a display mode of an object R having a predetermined characteristic in the flattened image N. As an example, the display specification information includes a display position of the object having the predetermined characteristic in the flattened image, and an occupancy rate of the object having the predetermined characteristic in the flattened image. As the occupancy rate, one of the following four ratios is used: - Ratio of the vertical size of the object to the vertical size of the flattened image - Ratio of the horizontal size of the object to the horizontal size of the flattened image - Ratio of the vertical size of the object to a size that is smaller than the vertical size of the flattened image by a first predetermined amount - Ratio of the horizontal size of the object to a size that is smaller than the horizontal size of the flattened image by a second predetermined amount

[0029] When the ratio of the vertical size of the object to a size that is smaller than the vertical size of the flattened image by a first predetermined amount is used as the occupancy rate, a portion of the object is prevented from being out of the frame of the flattened image more than when the ratio of the vertical size of the object to the vertical size of the flattened image is used as the occupancy rate. Furthermore, when the ratio of the horizontal size of the object to a size that is smaller than the horizontal size of the flattened image by a second predetermined amount is used as the occupancy rate, a portion of the object is prevented from being out of the frame of the flattened image more than when the ratio of the horizontal size of the object to the horizontal size of the flattened image is used as the occupancy rate. The first and second predetermined amounts may be determined in advance or may be specified by user Q. In the following description, the ratio of the vertical size of the object to a size that is smaller than the vertical size of the flattened image by a first predetermined amount is used as the occupancy rate. 1, the display designation information specifies that the display position of object R having a predetermined characteristic in flattened image N is to be the center, and that the occupancy rate of the entire body of object R in flattened image N is 100%. Therefore, in FIG. 1, the entire body of person P1 (i.e., object R) raising his hand is displayed in the center of flattened image N with an occupancy rate of 100%.

[0030] The image processing device 100 outputs flattened image data representing the flattened image N to the display device 300. The display device 300 displays the flattened image N based on the flattened image data. Because the flattened image N is a flattened image that is not curved, people can intuitively understand the flattened image N. Therefore, the user Q can easily grasp the current state of the monitored area 120 by checking the flattened image N displayed on the display device 300.

[0031] Furthermore, when a characteristic portion that captures a predetermined feature is detected in the wide-angle image, the image processing device 100 instructs the warning device 500 to output an alert. The warning device 500 includes, for example, at least one of a speaker 501 and a warning light 502. The warning device 500 outputs the alert based on an instruction from the image processing device 100. More specifically, when the speaker 501 receives an instruction to output an alert, it outputs a warning sound and / or a warning notification voice. The warning light 502 turns on when it receives an instruction to output an alert. The display device 300 may also function as the warning device 500. In that case, the display device 300 outputs the alert based on an instruction from the image processing device 100. For example, when the display device 300 receives an instruction to output an alert, it changes the color of the display frame of the flattened image N.

[0032] The warning device 500 may be installed in the room where the user Q views the flattened image N, or in the monitored area 120. The warning device 500 may also be installed in both the room where the user Q views the flattened image N and the monitored area 120. The warning device 500 may also be a light-emitting device installed in the monitored area 120. When receiving an alert output instruction, the light-emitting device emits light toward an object having a predetermined characteristic (person P1 in the example shown in FIG. 1 ). When the warning device 500 is installed in the room where the user Q views the flattened image N, the warning device issues a warning to the user Q. When the warning device 500 is installed in the monitored area 120, the warning device issues a warning to a person (e.g., person P1) present in the monitored area 120.

[0033] The input device 200 accepts an operation by the user Q and transmits an operation signal corresponding to the accepted operation to the image processing device 100. The input device 200 includes a mouse and may further include a keyboard. The input device 200 may include a touch panel instead of the mouse and keyboard. The input device 200 may also include a touch panel in addition to the mouse and keyboard.

[0034] 2 is a diagram showing a second example of an image processing system according to the present embodiment. System 12 is the second example of an image processing system according to the present embodiment. In system 11 (see FIG. 1), wide-angle image data representing a wide-angle image captured by camera 400 is stored in storage 103 of image processing device 100. In contrast, in system 12, wide-angle image data representing a wide-angle image captured by camera 400 is stored in VMS (Video Management System) 600.

[0035] More specifically, system 12 includes image processing device 100, input device 200, display device 300, camera 400, warning device 500, and VMS 600. Camera 400 transmits wide-angle image data representing a wide-angle image to VMS 600. Camera 400 is connected to VMS 600 via network 15. VMS 600 receives the wide-angle image data from camera 400. The wide-angle image data is stored in VMS 600. VMS 600 distributes the wide-angle image data to image processing device 100. VMS 600 is connected to image processing device 100 via network 15. Image processing device 100 receives the wide-angle image data from VMS 600 and generates flattened image N in the manner described with reference to FIG. 1 . In other respects, system 12 is similar to system 11.

[0036] In the following description, when there is no need to distinguish between the system 11 and the system 12, they will be referred to as "system 10."

[0037] <B. Hardware Configuration of Image Processing Device 100> The hardware configuration of the image processing device 100 will be described with reference to Fig. 3. Fig. 3 is a diagram showing an example of the hardware configuration of the image processing device 100. As shown in Fig. 3, the image processing device 100 includes, for example, a processor 101, a memory 102, a storage 103, an input interface 104, a display interface 105, and a communication interface 106. The processor 101, the memory 102, the storage 103, the input interface 104, the display interface 105, and the communication interface 106 are connected by a bus 199.

[0038] The processor 101 is, for example, a CPU (Central Processing Unit). The processor 101 reads a program 113 stored in the storage 103, loads it into the memory 102, and executes it.

[0039] The memory 102 is configured as a volatile storage device such as a dynamic random access memory (DRAM) or a static random access memory (SRAM).

[0040] The storage 103 is configured by a non-volatile storage device such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory.

[0041] The storage 103 stores a program 113. The program 113 includes a plurality of computer-readable instructions for controlling the image processing device 100. The processor 101 executes the program 113 to control each unit of the image processing device 100 and realize various processes according to the present embodiment.

[0042] The program 113 may be provided not as a standalone program but as part of an arbitrary program. In this case, the program 113 cooperates with the arbitrary program to realize the processing according to this embodiment. Even if the program does not include some of these modules, this does not deviate from the spirit of the image processing device 100 according to this embodiment. Furthermore, some or all of the functions provided by the program 113 may be realized by dedicated hardware.

[0043] The storage 103 further stores a management table 114 for managing detection targets and display specification information. The management table 114 will be described in detail later.

[0044] The storage 103 also stores planarization parameters 116 calculated by a process described below.

[0045] In the case of the image processing device 100 in the system 11 (see FIG. 1), the storage 103 further stores wide-angle image data 115 representing wide-angle images captured by the camera 400 .

[0046] The input interface 104 receives, from the input device 200, an operation signal corresponding to an operation accepted by the input device 200. The input interface 104 transmits the operation signal to the processor 101.

[0047] The display interface 105 transmits an instruction to display the image to the display device 300 based on an instruction from the processor 101 .

[0048] The communication interface 106 exchanges data with devices connected to the network 15 via the network 15. In the case of the image processing device 100 in the system 11, the device connected to the network 15 is, for example, the camera 400. In the case of the image processing device 100 in the system 12 (see FIG. 2), the device connected to the network 15 is, for example, the VMS 600.

[0049] <C. Example of Processing by System 10> (c1 Overview of Processing by System 10) An overview of processing by system 10 will be described with reference to Fig. 4 and Fig. 5. Fig. 4 is a diagram showing an overview of processing by system 11. System 11 includes an image processing device 100, an input device 200, a display device 300, a camera 400, and a warning device 500.

[0050] The camera 400 captures an image of the monitoring area 120 (see FIG. 1) and acquires a wide-angle image of the monitoring area 120. The camera 400 transmits wide-angle image data 115 representing the wide-angle image to the image processing device 100.

[0051] The image processing device 100 includes a reception unit 151, an acquisition unit 152, a detection unit 153, a generation unit 154, an output unit 155, a warning unit 156, and a storage 103. The reception unit 151, the acquisition unit 152, the detection unit 153, the generation unit 154, the output unit 155, and the warning unit 156 are realized by a processor 101 that executes a program 113 (see FIG. 3 ).

[0052] User Q (see FIG. 1 ) operates the input device 200 to specify a detection target and display specification information. The receiving unit 151 receives the detection target and display specification information specified by user Q based on an operation signal from the input device 200. User Q selects a desired set from among the sets of detection targets and display specification information included in the management table 114. Note that the management table 114 may be configured to be editable based on the operation of user Q. For example, if the desired set is not included in the management table 114, user Q can add the desired set to the management table 114. Here, the management table 114 will be described with reference to FIG. 5 .

[0053] FIG. 5 is a diagram illustrating an example of the management table 114. The management table 114 includes one or more sets of detection targets and display designation information corresponding to the detection targets. The detection targets are the predetermined characteristics described above. In the example illustrated in FIG. 5, the management table 114 includes display designation information d1 to d5, and d11 as display designation information when the detection target is a person's gesture of raising their hand. The management table 114 includes display designation information d6 and d7 as display designation information when the detection target is a person's gesture of running. The management table 114 includes display designation information d8 and d9 as display designation information when the detection target is luggage. The management table 114 includes display designation information d10 as display designation information when the detection target is a person's gesture of raising their hand and luggage. In the following description, when there is no need to distinguish between one or more pieces of display designation information, they will be referred to as "display designation information d."

[0054] Each of the display specification information d1 to d9 includes a display position of an object having a predetermined characteristic in the flattened image and an occupancy rate of the object having the predetermined characteristic in the flattened image. The display specification information d10 includes a display mode of a plurality of objects corresponding to a plurality of predetermined characteristics in the flattened image. The display specification information d11 includes a display mode of a plurality of objects having the predetermined characteristic in the flattened image.

[0055] User Q operates input device 200 (see FIG. 4) to specify a desired set of detection targets and display specification information from management table 114.

[0056] 4, acquisition unit 152 receives wide-angle image data 115 indicating a wide-angle image of monitoring area 120 from camera 400. Acquisition unit 152 stores wide-angle image data 115 in storage 103.

[0057] The detection unit 153 detects a characteristic portion in which a predetermined characteristic appears from the wide-angle image represented by the wide-angle image data 115. The predetermined characteristic is a detection target designated by the user Q.

[0058] The generation unit 154 generates a flattened image obtained by flattening a target region including a characteristic portion of the wide-angle image based on predetermined display specification information. The display specification information is specified in advance by the user Q. The generation unit 154 stores the flattening parameters 116 used to generate the current flattened image in the storage 103. Note that if the flattening parameters 116 have already been stored in the storage 103, the generation unit 154 updates the flattening parameters 116 stored in the storage 103 to the flattening parameters 116 used to generate the current flattened image.

[0059] The output unit 155 outputs flattened image data representing the flattened image to the display device 300. The display device 300 displays the flattened image based on the flattened image data.

[0060] When a characteristic portion showing a predetermined characteristic is detected from the wide-angle image, the warning unit 156 instructs the warning device 500 to output an alert. When the warning device 500 receives the instruction to output an alert, it outputs the alert.

[0061] In the case of system 12 (see FIG. 2), camera 400 transmits wide-angle image data 115 to VMS 600 (see FIG. 2), and wide-angle image data 115 is stored in VMS 600. Acquisition unit 152 receives wide-angle image data 115 from VMS 600. In other respects, the processing of system 12 is similar to the processing of system 11.

[0062] (c2 Details of Processing by the System 10) Details of processing by the system 10 will be described with reference to FIGS. 6 to 22.

[0063] 6 is a diagram for explaining the processing of the detection unit 153. The wide-angle image J1 is an example of a wide-angle image represented by the wide-angle image data 115 (see FIG. 4) acquired by the acquisition unit 152 (see FIG. 4). Note that the two-dot chain line in FIG. 6 is added for the purpose of explanation and is not included in the actual wide-angle image J1.

[0064] The detection unit 153 (see FIG. 4 ) detects a feature portion in the wide-angle image J1 that captures a predetermined feature. More specifically, the detection unit 153 first generates one or more feature detection images to be used to detect the feature portion from the wide-angle image J1. As an example, the detection unit 153 divides the wide-angle image J1 into partial images each having a predetermined horizontal rotation angle, and flattens each partial image to obtain one or more feature detection images. The predetermined horizontal rotation angle is an arbitrary angle. The predetermined horizontal rotation angle may be, for example, 90 degrees or 45 degrees. Each of the one or more feature detection images is an image obtained by flattening the wide-angle image J1.

[0065] In the example shown in FIG. 6, the detection unit 153 first obtains partial images J11 to J14 by dividing the wide-angle image J1 in four directions. That is, in the example shown in FIG. 6, the detection unit 153 divides the wide-angle image J1 into four partial images J11 to J14, each with a horizontal rotation angle of 90 degrees. Next, the detection unit 153 flattens each of the partial images J11 to J14 to generate feature detection images K11 to K14. The feature detection image K11 is an image obtained by flattening the partial image J11. The feature detection image K12 is an image obtained by flattening the partial image J12. The feature detection image K13 is an image obtained by flattening the partial image J13. The feature detection image K14 is an image obtained by flattening the partial image J14.

[0066] Next, the detection unit 153 performs a detection operation on each of the feature detection images K11 to K14. The detection operation is an operation of detecting a feature portion that shows a predetermined feature. If a gesture of a person raising their hand is specified as the predetermined feature, the detection unit 153 detects a portion of the feature detection image K11 that shows the person P1 raising their hand as the feature portion.

[0067] The detection unit 153 may generate an equirectangular image from the wide-angle image J1 and extract four areas specified at 90-degree intervals from the equirectangular image as partial images J11 to J14. The detection unit 153 may use the partial images J11 to J14 as feature detection images K11 to K14.

[0068] 7 is a diagram showing another example of a method for generating an image for feature detection. As another example of the method for generating an image for feature detection, first, the detection unit 153 (see FIG. 4) performs polar coordinate transformation on the wide-angle image J1 (see FIG. 6) to obtain an image K221. The image K221 is an example of an image obtained by performing polar coordinate transformation on the wide-angle image J1. The image K221 is an image obtained by flattening the wide-angle image J1.

[0069] Next, the detection unit 153 generates one or more feature detection images by trimming the image K221 at predetermined angles. The predetermined angle is any angle. For example, the predetermined angle may be 90 degrees or 45 degrees. The feature detection images K21 to K24 are examples of feature detection images obtained by trimming the image K221 at 90-degree intervals. Note that the detection unit 153 may use the image K221 as the feature detection image.

[0070] 8 is a diagram for explaining an example of the detection operation. The feature detection image K15 is an example of the feature detection image generated by the detection unit 153 (see FIG. 4) for detecting the feature portion.

[0071] When the predetermined feature is a specific identity of a person or a specific behavior of a person, the detection unit 153 first executes a skeleton detection (AI (Artificial Intelligence)) process on the feature detection image K15. When a skeleton is detected, the detection unit 153 executes a feature recognition process on the feature detection image K15 to detect a feature portion in which the predetermined feature appears from the feature detection image K15.

[0072] By performing skeletal detection and feature recognition processing, the detection unit 153 can detect feature portions that capture the following features, for example: Falling over, Crouching, Reaching for something from a shelf, Walking, Entering a predetermined area, Specific behavior of a person who has entered a predetermined area, Specific actions in sports (for example, a serving posture in table tennis or tennis, a pitching form, batting form, or catching form in baseball, etc.).

[0073] 8 shows an example of a detection operation when a gesture of a person raising a hand is specified as a predetermined feature. First, the detection unit 153 uses the skeleton of the largest bounding box in the feature detection video K15 as the target for hand-raising determination. In the example shown in FIG. 8, the feature detection video K15 includes only two bounding boxes b1 and b2. The bounding box b1 is larger than the bounding box b2. The detection unit 153 uses the skeleton of the bounding box b1 as the target for hand-raising determination.

[0074] Next, the detection unit 153 identifies the positional relationship between the right wrist and right elbow, and the positional relationship between the left wrist and left elbow. If the right wrist is higher than the right elbow, or if the left wrist is higher than the left elbow, the detection unit 153 determines that the hand is raised. In FIG. 8 , the left wrist ("l_wrist" in the figure) is higher than the left elbow ("l_elbow" in the figure), so the detection unit 153 determines that the hand is raised. Next, the detection unit 153 determines that the hand is raised if the hand is raised for a predetermined time (e.g., 1 second) or more. The detection unit 153 also determines that the hand is raised if the time during which the hand is raised out of the predetermined time (e.g., 1 second) is equal to or greater than a predetermined value (e.g., 80%). The detection unit 153 detects the area of ​​the bounding box b1 in which the hand is determined to be raised as a characteristic portion in which a predetermined characteristic is captured.

[0075] A skeleton with a small bounding box may not actually be a skeleton. Therefore, false detection can be prevented by using the skeleton with the largest bounding box in the feature detection video K15 as the target for hand-raising judgment. However, the detection unit 153 may also perform hand-raising judgment for all of the multiple bounding boxes in the feature detection video K15. In addition, user Q (see Figures 1 and 2) can arbitrarily change the predetermined time and predetermined value for the detection operation.

[0076] As shown in Fig. 8, the detector 153 performs detection operations on the feature detection image, which is a flattened version of the wide-angle image. The flattened image is more suitable for image recognition processing than the wide-angle image. This improves the accuracy of detecting feature portions that capture predetermined features.

[0077] In addition, if the predetermined feature is a specific identity of an object or a specific state of an object, the detection unit 153 detects a feature part in which the predetermined feature appears from the feature detection image by object detection (AI) and feature recognition processing.

[0078] FIG. 9 is a diagram illustrating the processing of the generation unit 154. Wide-angle image J1 is an example of a wide-angle image represented by wide-angle image data 115 (see FIG. 4) acquired by acquisition unit 152 (see FIG. 4). Feature detection image K11 is an example of a feature detection image generated by detection unit 153 (see FIG. 4) for detecting feature portions. In the example shown in FIG. 9, a gesture of a person raising their hand is designated as the predetermined feature. In the example shown in FIG. 9, display designation information d1 (see FIG. 5) is designated as the display designation information. Object R1 is an example of an object having a predetermined feature.

[0079] In the example shown in FIG. 9 , first, the generation unit 154 (see FIG. 4 ) converts coordinates indicating the center position F1 of the object R1 in the feature detection video K11 into coordinates indicating the center position F2 of the object R1 in the wide-angle video J1. The center position F1 is the center position of a bounding box b3 of the object R1 in the feature detection video K11. The center position F1 is an example of a "first display position" in the present disclosure. The coordinates indicating the center position F1 are coordinates indicating the display position of the object R1 in the feature detection video K11. In other words, the coordinates indicating the center position F1 are an example of a "first coordinate" in the present disclosure. The center position F2 is the center position of a bounding box b4 of the object R1 in the wide-angle video J1. The center position F2 is an example of a "second display position" in the present disclosure. The coordinates indicating the center position F2 are coordinates indicating the display position of the object R1 in the wide-angle video J1. That is, the coordinates indicating the center position F2 are an example of the "second coordinates" in this disclosure.

[0080] Next, the generation unit 154 determines the target region Z1 based on the coordinates indicating the center position F2 and the display specification information d1. Determining the target region Z1 involves calculating flattening parameters 116 (see FIG. 4 ). The flattening parameters 116 include a horizontal rotation angle, a vertical rotation angle, and a field of view (FOV). Next, the generation unit 154 generates a flattened image N1 obtained by flattening the target region Z1. As an example, the generation unit 154 converts a partial image of the wide-angle image J1 that indicates the target region Z1 into an equirectangular image and generates the flattened image N1 based on the equirectangular image. This results in a flattened image N1 in which the entire body of the object R1 is captured at 100% of the center of the flattened image N1.

[0081] Note that the generation unit 154 may generate the flattened image N1 without using an equirectangular image. Furthermore, the generation unit 154 may generate the flattened image N1 so that an object R1 having a predetermined characteristic is highlighted. As an example, as shown in FIG. 9 , the flattened image N1 may include a bounding box X1 that surrounds the object R1. The bounding box X1 may be displayed in a conspicuous color, such as red. The bounding box X1 may be displayed in a flashing manner.

[0082] Other examples of flattened images will be described with reference to FIGS. 10 to 12. FIG. 10 is a diagram illustrating an example of a flattened image when display designation information d2 is designated as the display designation information. In the example illustrated in FIG. 10, the wide-angle image represented by wide-angle image data 115 (see FIG. 4) acquired by acquisition unit 152 (see FIG. 4) is wide-angle image J1 (see FIG. 6). In the example illustrated in FIG. 10, a gesture of a person raising their hand is designated as the predetermined characteristic. In the example illustrated in FIG. 10, display designation information d2 (see FIG. 5) is designated as the display designation information. As illustrated in FIG. 10, when display designation information d2 is designated, generation unit 154 (see FIG. 4) generates flattened image N2 in which the entire body of object R1 having the predetermined characteristic is displayed in the center of flattened image N2 with an occupancy rate of 50%.

[0083] FIG. 11 is a diagram showing an example of a flattened image when display designation information d3 is designated as the display designation information. In the example shown in FIG. 11, the wide-angle image represented by wide-angle image data 115 (see FIG. 4) acquired by acquisition unit 152 (see FIG. 4) is wide-angle image J1 (see FIG. 6). In the example shown in FIG. 11, a gesture of a person raising their hand is designated as the predetermined characteristic. In the example shown in FIG. 11, display designation information d3 (see FIG. 5) is designated as the display designation information. As shown in FIG. 11, when display designation information d3 is designated, generation unit 154 (see FIG. 4) generates flattened image N3 in which the upper body of object R1 having the predetermined characteristic is displayed at the center of flattened image N3 with a 100% occupancy rate.

[0084] FIG. 12 is a diagram illustrating an example of a flattened image when display designation information d4 is designated as the display designation information. In the example illustrated in FIG. 12, the wide-angle image represented by wide-angle image data 115 (see FIG. 4) acquired by acquisition unit 152 (see FIG. 4) is wide-angle image J1 (see FIG. 6). In the example illustrated in FIG. 12, a gesture of a person raising their hand is designated as the predetermined characteristic. In the example illustrated in FIG. 12, display designation information d4 (see FIG. 5) is designated as the display designation information. As illustrated in FIG. 12, when display designation information d4 is designated, generation unit 154 (see FIG. 4) generates a flattened image N4 in which the entire body of object R1 having the predetermined characteristic is captured at position F3 in the flattened image N4 with a 50% occupancy rate. Position F3 is a position in flattened image N4 where the x coordinate is 25% and the y coordinate is 25%. If there is a high probability that the item to be confirmed (for example, luggage R11) will appear in the lower right area (area e1) of object R1, user Q (see Figures 1 and 2) can specify display specification information d4 as display specification information to confirm the item to be confirmed within flattened image N4 showing object R1.

[0085] Fig. 13 is a diagram illustrating another example of coordinate transformation by the generation unit 154. In the example shown in Fig. 9, coordinates indicating the center position of a bounding box of an object having a predetermined feature in the feature detection video are used as coordinates indicating the display position of the object having the predetermined feature in the feature detection video. In contrast, in the example shown in Fig. 13, coordinates indicating the positions of the four vertices of the bounding box of the object having the predetermined feature are used as coordinates indicating the display position of the object having the predetermined feature in the feature detection video.

[0086] More specifically, wide-angle image J3 is an example of a wide-angle image represented by wide-angle image data 115 (see FIG. 4) acquired by acquisition unit 152 (see FIG. 4). Feature detection image K31 is an example of a feature detection image generated by detection unit 153 (see FIG. 4) for detecting feature portions. In the example shown in FIG. 13, a gesture of a person raising their hand is designated as the predetermined feature. Object R2 is an example of an object having a predetermined feature.

[0087] The generation unit 154 uses coordinates indicating the positions F11 to F14 of the four vertices of the bounding box b5 of the object R2 as coordinates indicating the display position of the object R2 in the feature detection video K31. The generation unit 154 converts the coordinates indicating the positions F11 to F14 into coordinates indicating the positions F21 to F24 of the four vertices of the bounding box b6 of the object R2 in the wide-angle video J3. More specifically, the generation unit 154 converts the coordinates indicating the position F11 in the feature detection video K31 into coordinates indicating the position F21 in the wide-angle video J3. The generation unit 154 converts the coordinates indicating the position F12 in the feature detection video K31 into coordinates indicating the position F22 in the wide-angle video J3. The generation unit 154 converts the coordinates indicating the position F13 in the feature detection video K31 into coordinates indicating the position F23 in the wide-angle video J3. The generation unit 154 converts the coordinates indicating the position F14 in the feature detection image K31 into coordinates indicating the position F24 in the wide-angle image J3.

[0088] Positions F11 to F14 are an example of a "first display position" in the present disclosure. Coordinates indicating positions F11 to F14 are coordinates indicating the display position of object R2 in feature detection image K31. In other words, coordinates indicating positions F11 to F14 are an example of a "first coordinate" in the present disclosure. Positions F21 to F24 are an example of a "second display position" in the present disclosure. Coordinates indicating positions F21 to F24 are coordinates indicating the display position of object R2 in wide-angle image J3. In other words, coordinates indicating positions F21 to F24 are an example of a "second coordinate" in the present disclosure.

[0089] Once the coordinate transformation is complete, the generation unit 154 determines the target region Z2 based on the coordinates indicating the positions F21 to F24 and the display designation information previously designated by the user Q (see FIGS. 1 and 2). Next, the generation unit 154 generates a flattened image obtained by flattening the target region Z2. The coordinate transformation shown in FIG. 13 makes it possible to more accurately reproduce the position and shape of the object R2 in the wide-angle image J3 in the flattened image.

[0090] A method for calculating the flattening parameter 116 (see FIG. 4) by the generation unit 154 (see FIG. 4) will be described with reference to FIG. 14. FIG. 14 is a diagram showing an example of a method for calculating the horizontal rotation angle and the vertical rotation angle by the generation unit 154. FIG. 14 shows a coordinate system 1500 in which the fisheye video is viewed from above, and a coordinate system 1510 in which the fisheye video is viewed obliquely.

[0091] Using coordinate system 1500 as an example, the generation unit 154 calculates the horizontal rotation angle θ based on equations 1501, 1502, 1503, and 1504. px and py are coordinates indicating the display position of an object having predetermined characteristics in the fisheye image. For example, in the example shown in FIG. 9 , px and py are coordinates indicating center position F2. The generation unit 154 calculates the horizontal rotation angle θ using px and py and equations 1501 to 1504.

[0092] Taking coordinate system 1510 as an example, generation unit 154 calculates vertical rotation angle η based on equations 1511, 1512, 1513, and 1514. Generation unit 154 calculates vertical rotation angle η using px and py and equations 1511 to 1514.

[0093] FIG. 15 is a diagram illustrating an example of a method for calculating the FOV by the generation unit 154. FIG. 15 illustrates the relationship between a fisheye image 1600 and a conceptually illustrated side view 1610 of the camera 400. In the fisheye image 1600, objects 1601 and 1602 are captured. The object 1602 is an example of an object having predetermined characteristics in the fisheye image. In the side view 1610, for example, the object 1601 corresponds to the reflected light of an actual object 1603 passing through the fisheye lens 1612. Furthermore, the object 1602 corresponds to the reflected light of an actual object 1604 passing through the fisheye lens 1612.

[0094] The generation unit 154 calculates the FOV using the reference distance L to the object, the default viewing angle A, the flattened screen window size W, and equations 1621, 1622, 1623, and 1624. The flattened screen window size is the size of the window in which the flattened image is displayed. The reference distance L to the object, the default viewing angle A, and the flattened screen window size W may be input in advance to the image processing device 100 (see FIGS. 1 and 2). The image processing device 100 may also include a UI (User Interface) for user Q (see FIGS. 1 and 2) to input the reference distance L to the object, the default viewing angle A, and the flattened screen window size W.

[0095] FIG. 16 is a diagram showing a first example of an alert output by the warning device 500. In the example shown in FIG. 16, the warning device 500 includes a speaker 503 and a light 504. The speaker 503 and the light 504 are provided in the monitoring area 120 (see FIGS. 1 and 2). In the example shown in FIG. 16, a gesture of a person raising their hand is specified as the predetermined characteristic. When a person raising their hand is detected in the wide-angle image, the warning unit 156 (see FIG. 4) instructs the speaker 503 and the light 504 to output an alert. When the speaker 503 receives an instruction to output an alert, the speaker 503 outputs a warning sound and / or a voice warning notification. When the light 504 receives an instruction to output an alert, the light 504 shines light on an object R1 having the predetermined characteristic.

[0096] As a result, when an object R1 having predetermined characteristics is detected by the image processing device 100 (see FIGS. 1 and 2), a warning sound and / or a warning notification voice is output in the monitoring area 120. Furthermore, when an object R1 having predetermined characteristics is detected by the image processing device 100, a light is irradiated onto the object R1. Therefore, attention is drawn to the object R1 in the monitoring area 120.

[0097] The warning device 500 may include the speaker 503 but not the light 504. The warning device 500 may not include the speaker 503 but may include the light 504. The warning device 500 may include a laser pointer. When an instruction to output an alert is received, the laser pointer irradiates a laser beam onto an object R1 having predetermined characteristics. The warning device 500 may include a projector. When an instruction to output an alert is received, the projector projects a warning notification. The warning notification may be, for example, an image resembling a no-entry bar, or text information indicating the content of the warning.

[0098] Fig. 17 is a diagram showing a second example of an alert output by the warning device 500. In the example shown in Fig. 17, the warning device 500 includes a speaker 501 and a warning light 502. The speaker 501 and the warning light 502 are provided in a room where a user Q (see Figs. 1 and 2) checks the flattened image. In the example shown in Fig. 17, the display device 300 also functions as the warning device 500. In the example shown in Fig. 17, a person's approach to a restricted area e2 is specified as a predetermined characteristic.

[0099] When a person approaching the restricted area e2 is detected in the wide-angle image, the warning unit 156 (see FIG. 4 ) instructs the speaker 501, the warning light 502, and the display device 300 to output an alert. When the speaker 501 receives the instruction to output an alert, it outputs a warning sound and / or a warning notification voice. When the warning light 502 receives the instruction to output an alert, it turns on. When the display device 300 receives the instruction to output an alert, it changes the color of the display frame e3 of the flattened image N5. The flattened image N5 shows the restricted area e2 and a person approaching the restricted area e2 (i.e., an object R3 having predetermined characteristics).

[0100] As a result, when an object R3 having predetermined characteristics is detected by the image processing device 100 (see FIGS. 1 and 2), a warning sound and / or a warning notification voice is output in the room where the user Q (see FIGS. 1 and 2) is viewing the flattened image N5. Furthermore, when an object R3 having predetermined characteristics is detected by the image processing device 100, a warning light 502 is turned on in the room where the user Q is viewing the flattened image N5. Furthermore, when an object R3 having predetermined characteristics is detected by the image processing device 100, the color of the display frame e3 of the flattened image N5 is changed. Therefore, the user Q is alerted.

[0101] The warning device 500 may include the speaker 501 but not the warning light 502. The warning device 500 may not include the speaker 501 but include the warning light 502. The display device 300 may not have the function of the warning device 500.

[0102] Fig. 18 is a flowchart showing an example of the processing procedure of the image processing device 100. Fig. 18 shows the processing procedure of the image processing device 100 in the case where the predetermined feature is a specific identity of a person or a specific behavior of a person.

[0103] First, in step S1, the processor 101 (see FIG. 3) acquires a wide-angle image indicated by the wide-angle image data 115 (see FIG. 3) received from the camera 400 (see FIG. 1) or the VMS 600 (see FIG. 2).

[0104] Next, in step S2, the processor 101 generates one or more feature detection images (for example, feature detection images K11 to K14 shown in FIG. 6) to be used for detecting feature portions from the wide-angle image.

[0105] Next, in step S3, the processor 101 performs a skeleton detection process on each of the one or more feature detection images.

[0106] Next, in step S4, processor 101 determines whether a skeleton has been detected. If a skeleton has been detected (YES in step S4), processor 101 proceeds to step S5. On the other hand, if a skeleton has not been detected (NO in step S4), processor 101 proceeds to step S11.

[0107] In step S5, the processor 101 performs feature recognition processing on one or more feature detection images (for example, feature detection images K11 to K13 shown in FIG. 6) from which a skeleton has been detected.

[0108] Next, in step S6, processor 101 determines whether a characteristic portion showing a predetermined feature has been detected from one or more feature detection images from which a skeleton has been detected. If a characteristic portion showing a predetermined feature has been detected from one or more feature detection images from which a skeleton has been detected (YES in step S6), processor 101 proceeds to step S7. On the other hand, if a characteristic portion showing a predetermined feature has not been detected from one or more feature detection images from which a skeleton has been detected (NO in step S6), processor 101 proceeds to step S11.

[0109] In step S7, the processor 101 acquires coordinates indicating the display position of an object having a predetermined feature in the feature detection image in which the feature portion has been detected. For example, in the example shown in Fig. 9, the processor 101 acquires coordinates indicating the center position F1 of the object R1 in the feature detection image K11.

[0110] Next, in step S8, the processor 101 acquires the display specification information d (see FIG. 5) specified by the user Q (see FIGS. 1 and 2).

[0111] Next, in step S9, processor 101 converts the coordinates indicating the display position of the object having the predetermined feature in the feature detection image, in which the feature portion has been detected, into coordinates indicating the display position of the object having the predetermined feature in the wide-angle image. For example, in the example shown in Fig. 9, processor 101 converts the coordinates indicating the center position F1 of object R1 in feature detection image K11 into coordinates indicating the center position F2 of object R1 in wide-angle image J1.

[0112] Next, in step S10, the processor 101 calculates flattening parameters 116 (see FIG. 4 ) based on coordinates indicating the display position of an object having predetermined characteristics in the wide-angle image and the display specification information d acquired in step S8. For example, in the example shown in FIG. 9 , the processor 101 calculates the flattening parameters 116 based on coordinates indicating the center position F2 and the display specification information d1 (see FIG. 5 ). The processor 101 stores the calculated flattening parameters 116 in the storage 103 (see FIG. 4 ). Note that if the flattening parameters 116 have already been stored in the storage 103, the processor 101 updates the flattening parameters 116 stored in the storage 103 with the currently calculated flattening parameters 116.

[0113] In step S11, the processor 101 obtains the flattening parameters used to generate the previous flattened image from the storage 103. The storage 103 stores the flattening parameters used to generate the previous flattened image.

[0114] After step S10 or step S11, the processor 101 proceeds to step S12. In step S12, the processor 101 generates a flattened image using the flattening parameters 116 calculated in step S10 or the flattening parameters 116 acquired in step S11.

[0115] Next, in step S13, the processor 101 outputs flattened image data representing the flattened image to the display device 300.

[0116] Next, in step S14, processor 101 determines whether or not the conditions for issuing a warning notification are met. When a characteristic portion that captures a predetermined characteristic is detected in the wide-angle image, processor 101 determines that the conditions for issuing a warning notification are met. When the conditions for issuing a warning notification are met (YES in step S14), processor 101 proceeds to step S15. On the other hand, when the conditions for issuing a warning notification are not met (NO in step S14), processor 101 returns the process to step S1.

[0117] In step S15, processor 101 instructs warning device 500 to output an alert. After step S15, processor 101 returns the process to step S1.

[0118] Through the series of processes shown in FIG. 18, a flattened image is displayed on the display device 300 in real time.

[0119] Fig. 19 is a diagram illustrating a case where display designation information d6 or d7 is designated as the display designation information. Wide-angle image J4 is an example of a wide-angle image represented by wide-angle image data 115 (see Fig. 4) acquired by acquisition unit 152 (see Fig. 4). In the example shown in Fig. 19, a gesture of a person running is designated as the predetermined characteristic. Object R4 is an example of an object having a predetermined characteristic. Object R4 is moving.

[0120] The feature detection image K41 is an example of a feature detection image generated by the detection unit 153 (see FIG. 4) for detecting feature portions. The feature detection image K41 is an image obtained by flattening a partial image J41 of the wide-angle image J4. The person indicated by the dotted line f1 in the feature detection image K41 is added for explanatory purposes and is not actually captured in the feature detection image K41. The dotted line f1 indicates the position of the object R4 in the previous frame of the feature detection image K41.

[0121] Flattened image N6 is an example of a flattened image when display designation information d6 (see FIG. 5) is designated as the display designation information. As shown in FIG. 19, when display designation information d6 is designated, generation unit 154 (see FIG. 4) generates flattened image N6 in which the entire body of object R4 having predetermined characteristics is displayed in the center of flattened image N6 with an occupancy rate of 100%.

[0122] Flattened image N7 is an example of a flattened image when display designation information d7 (see FIG. 5) is designated as the display designation information. In flattened image N7, the person indicated by dotted line f3 and the obstacle indicated by dotted line f4 are added for explanatory purposes and are not actually captured in flattened image N7.

[0123] When display specification information d7 is specified, the generation unit 154 generates a flattened image N7 in which the entire body of an object R4 having predetermined characteristics is displayed at an automatically adjusted position with an occupancy rate of 50%.

[0124] More specifically, the display specification information d7 specifies that the display position of an object having a predetermined characteristic is automatically adjusted based on the moving speed of the object. When the display specification information d7 is specified, the generation unit 154 adjusts the display position of the object R4 in the flattened image N7 based on the moving direction and moving speed of the object R4.

[0125] The generation unit 154 determines the moving direction of the object R4 based on the coordinates indicating the position of the object R4 in the current frame of the feature detection video K41 and the coordinates indicating the position of the object R4 in the previous frame of the feature detection video K41. In the example shown in Fig. 19 , the object R4 is moving from left to right in the feature detection video K41.

[0126] The generation unit 154 calculates the moving distance f2 of the object R4 based on coordinates indicating the position of the object R4 in the current frame of the feature detection video K41 and coordinates indicating the position of the object R4 in the immediately previous frame of the feature detection video K41. The generation unit 154 calculates the time between a first timing when the current frame is acquired and a second timing when the immediately previous frame is acquired. The generation unit 154 calculates the moving speed of the object R4 based on the moving distance f2 and the time between the first timing and the second timing.

[0127] The generation unit 154 adjusts the display position of the object R4 in the flattened image N7 based on the moving direction and moving speed of the object R4. As an example, the generation unit 154 calculates the position of the object R4 in the next frame of the flattened image N7 based on the moving direction and moving speed of the object R4. The dotted line f3 indicates the position of the object R4 in the next frame of the flattened image N7. The generation unit 154 determines the display position of the object R4 in the flattened image N7 so that the object R4 appears in the flattened image N7 and the next frame of the flattened image N7. In the example shown in FIG. 19 , the display position of the object R4 is determined to be in an area behind the center line of the flattened image N7 in the traveling direction of the object R4.

[0128] Adjusting the display position of object R4 can prevent object R4 from going out of frame in the next frame of flattened image N7. Also, adjusting the display position of object R4 allows user Q (see FIGS. 1 and 2) to check flattened image N7 and confirm whether there is an obstacle (e.g., an obstacle indicated by dotted line f4) ahead in the traveling direction of object R4 that may obstruct the traveling of object R4.

[0129] 20 is a diagram illustrating a case where multiple people raise their hands at the same time. Wide-angle image J5 is an example of a wide-angle image represented by wide-angle image data 115 (see FIG. 4) acquired by acquisition unit 152 (see FIG. 4). Feature detection image K51 is an example of a feature detection image generated by detection unit 153 (see FIG. 4) for detecting feature portions. Feature detection image K51 is an image obtained by flattening partial image J51 of wide-angle image J5.

[0130] 20, a gesture of a person raising their hand is specified as the predetermined feature. Both the object R5 and the object R6 are examples of objects having the predetermined feature. The object R5 and the object R6 are raising their hands at the same time.

[0131] There are three patterns of flattened images that are generated and displayed when multiple people raise their hands at the same time. "Simultaneously" can mean either at the exact same time or within a certain period of time.

[0132] In the first pattern, the generation unit 154 (see FIG. 4) generates a flattened image for each of the multiple people who raise their hands, in accordance with the display specification information d (see FIG. 5) specified by the user Q (see FIGS. 1 and 2). That is, as many flattened images as there are people who raise their hands are generated. The generation unit 154 outputs a display instruction for the multiple flattened images to the display device 300 so that the multiple flattened images are displayed with a time difference.

[0133] For example, when display designation information d1 (see FIG. 5) is designated as the display designation information, the generation unit 154 generates a flattened image N8 in which the entire body of an object R5 having predetermined characteristics is displayed at the center of the flattened image N8 with a 100% occupancy rate. The generation unit 154 also generates a flattened image N9 in which the entire body of an object R6 having predetermined characteristics is displayed at the center of the flattened image N9 with a 100% occupancy rate. The generation unit 154 instructs the display device 300 to alternate between the flattened image N8 and the flattened image N9 every three seconds. As a result, the flattened image N8 and the flattened image N9 are displayed on the display device 300 while being alternated every three seconds. Note that three seconds is an example. The user Q can arbitrarily designate the display time of each flattened image.

[0134] In the second pattern, multiple flattened images are displayed simultaneously. More specifically, the generation unit 154 generates a flattened image for each of the multiple people who are raising their hands, in accordance with the display specification information d specified by the user Q. That is, the same number of flattened images as the number of people who are raising their hands are generated. The generation unit 154 outputs a display instruction for the multiple flattened images to the display device 300 so that the multiple flattened images are displayed simultaneously.

[0135] For example, when display designation information d1 is designated as the display designation information, the generation unit 154 generates a flattened image N8 in which the entire body of an object R5 having predetermined characteristics is displayed at the center of the flattened image N8 with a 100% occupancy rate. The generation unit 154 generates a flattened image N9 in which the entire body of an object R6 having predetermined characteristics is displayed at the center of the flattened image N9 with a 100% occupancy rate. The generation unit 154 instructs the display device 300 to simultaneously display the flattened image N8 and the flattened image N9. As a result, the display device 300 simultaneously displays a window w1 in which the flattened image N8 is displayed and a window w2 in which the flattened image N9 is displayed.

[0136] In the third pattern, the generation unit 154 generates a flattened image in which a plurality of people who raise their hands are surrounded by bounding boxes. The third pattern is implemented when the display designation information d11 (see FIG. 5) is designated as the display designation information.

[0137] For example, the generation unit 154 generates a flattened image N10 in which multiple people raising their hands (i.e., objects R5 and R6) are surrounded by a bounding box X2. The flattened image N10 may further include information X3 indicating the order in which the hands were raised. If someone lowers their hand during the flattened image N10, the flattened image N10 may further include information indicating the order in which the remaining people raised their hands. The generation unit 154 outputs flattened image data representing the flattened image N10 to the display device 300.

[0138] In addition, if multiple people raise their hands at the same time, the generation unit 154 may select, from among the multiple people, a person who was not featured in the most recently generated flattened image as a target for generating a flattened image.

[0139] FIG. 20 illustrates a case where two people raise their hands at the same time. However, the flattened image generation method illustrated in FIG. 20 can also be applied to a case where three or more people raise their hands at the same time. FIG. 20 also illustrates a case where a gesture of a person raising their hand is specified as the predetermined feature. However, the flattened image generation method illustrated in FIG. 20 can also be applied to a case where a specific human behavior other than a gesture of a person raising their hand is specified as the predetermined feature and multiple people are performing the specific behavior. The flattened image generation method illustrated in FIG. 20 can also be applied to a case where a specific state of an object is specified as the predetermined feature and multiple objects are in the specific state.

[0140] FIG. 21 is a diagram illustrating a case where multiple people raise their hands in turn. Wide-angle image J6 is an example of a wide-angle image represented by wide-angle image data 115 (see FIG. 4) acquired by acquisition unit 152 (see FIG. 4). Feature detection images K61 and K62 are both examples of feature detection images generated by detection unit 153 (see FIG. 4) for detecting feature portions. Feature detection image K61 is an image obtained by flattening partial image J61 of wide-angle image J6. Feature detection image K62 is an image obtained by flattening partial image J62 of wide-angle image J6.

[0141] In the example shown in Fig. 21 , a gesture of a person raising their hand is specified as the predetermined feature. Both the object R7 and the object R8 are examples of objects having the predetermined feature. In the example shown in Fig. 21 , after the object R7 raises its hand, the object R8 raises its hand.

[0142] When multiple people raise their hands in turn, the generation unit 154 (see FIG. 4) generates a flattened image for each of the multiple people who have raised their hands, in accordance with the display specification information d (see FIG. 5) specified by user Q (see FIGS. 1 and 2). That is, the generation unit 154 generates multiple flattened images. Furthermore, the generation unit 154 generates one or more intermediate images based on the horizontal rotation angles of the multiple flattened images. The generation unit 154 instructs the display device 300 to switch between the multiple flattened images and one or more intermediate images in order of decreasing horizontal rotation angle. Note that the generation unit 154 may also instruct the display device 300 to switch between the multiple flattened images and one or more intermediate images in order of increasing horizontal rotation angle.

[0143] For example, the generation unit 154 generates a flattened image N11 in which an object R7 appears, in accordance with the display specification information d (see FIG. 5 ) specified by the user Q. The generation unit 154 also generates a flattened image N13 in which an object R8 appears, in accordance with the display specification information d specified by the user Q.

[0144] Furthermore, the generation unit 154 further generates an intermediate image N12 whose horizontal rotation angle is an arbitrary angle between the horizontal rotation angle used to generate the flattened image N11 and the horizontal rotation angle used to generate the flattened image N13. The intermediate image N12 is an image in which a portion of the wide-angle image J6 has been flattened. The generation unit 154 instructs the display device 300 to switch between the flattened image N11, the intermediate image N12, and the flattened image N13 in this order.

[0145] When the image displayed on the display device 300 switches from the flattened image N11 to the flattened image N13, it is difficult for the user Q to grasp the positional relationship between the object R7 and the object R8. In contrast, when the image displayed on the display device 300 switches in the order of the flattened image N11, the intermediate image N12, and the flattened image N13, the user Q gets the impression that the camera 400 has panned along the arrow AR. Therefore, the user Q can easily grasp the positional relationship between the object R7 and the object R8. The flattened images N11 and N13 may further include information indicating the order in which the hands were raised.

[0146] FIG. 21 illustrates a case where two people raise their hands in turn. However, the flattened image generation method illustrated in FIG. 21 can also be applied to a case where three or more people raise their hands in turn. FIG. 21 also illustrates a case where a gesture of a person raising their hand is specified as the predetermined feature. However, the flattened image generation method illustrated in FIG. 21 can also be applied to a case where a specific human behavior other than a gesture of a person raising their hand is specified as the predetermined feature, and multiple people perform the specific behavior in turn. The flattened image generation method illustrated in FIG. 21 can also be applied to a case where a specific state of an object is specified as the predetermined feature, and multiple objects transition to the specific state in turn.

[0147] 22 is a diagram illustrating a case where multiple features are specified as detection targets and the multiple features are detected. Wide-angle image J7 is an example of a wide-angle image represented by wide-angle image data 115 (see FIG. 4) acquired by acquisition unit 152 (see FIG. 4). Feature detection image K71 is an example of a feature detection image generated by detection unit 153 (see FIG. 4) for detecting feature portions. Feature detection image K71 is an image obtained by flattening partial image J71 of wide-angle image J7.

[0148] In the example shown in Fig. 22, a gesture of a person raising their hand and luggage are specified as predetermined features. Objects R9 and R10 are both examples of objects having predetermined features. Object R9 is an example of an object having the feature of raising a hand. Object R10 is an example of an object having the feature of luggage.

[0149] When a plurality of features are specified as detection targets and the plurality of features are detected, the flattened images that are generated and displayed include the following two types.

[0150] In a first example, the generation unit 154 (see FIG. 4) generates a flattened image in which a plurality of objects corresponding to a plurality of detected features are enclosed by bounding boxes. For example, the generation unit 154 generates a flattened image N14 in which an object R9 and an object R10 are enclosed by a bounding box X4. The first example is implemented when display specification information d10 (see FIG. 5) is specified as the display specification information.

[0151] In the second example, for each of the detected features, the generation unit 154 generates a flattened image showing an object having the feature in accordance with the display specification information d (see FIG. 5) specified by the user Q (see FIGS. 1 and 2). That is, the generation unit 154 generates a plurality of flattened images. The generation unit 154 outputs a display instruction for the plurality of flattened images to the display device 300 so that the plurality of flattened images are displayed simultaneously.

[0152] For example, when display designation information d5 and d9 are designated as the display designation information, the generation unit 154 generates a flattened image N15 in which the entire body of an object R9 having predetermined characteristics is displayed at the center of the flattened image N15 with an occupancy rate of 90%. The generation unit 154 generates a flattened image N16 in which an object R10 having predetermined characteristics is displayed at the center of the flattened image N16 with an occupancy rate of 30%. The generation unit 154 instructs the display device 300 to simultaneously display the flattened image N15 and the flattened image N16. As a result, the display device 300 simultaneously displays a window w3 in which the flattened image N15 is displayed and a window w4 in which the flattened image N16 is displayed.

[0153] 22 has been described in connection with a case where two features are designated as detection targets and the two features are detected. However, the method for generating a flattened image described in connection with FIG. 22 can also be applied to a case where three or more features are designated as detection targets and the three or more features are detected.

[0154] In this manner, the image processing device 100 of the present embodiment detects a characteristic portion showing a predetermined characteristic from the wide-angle image represented by the wide-angle image data. The image processing device 100 generates a flattened image obtained by flattening a target area including the characteristic portion in the wide-angle image based on predetermined display specification information. The image processing device 100 outputs flattened image data representing the flattened image to the display device 300. People cannot intuitively understand wide-angle images. In contrast, flattened images are not curved and are therefore intuitively understandable to people. Therefore, the image processing device 100 of the present embodiment can generate an image from the wide-angle image that is easy for people to intuitively understand. In other words, the image processing device 100 of the present embodiment can generate an image suitable for visual confirmation by humans. Therefore, according to the image processing device 100 of the present embodiment, the user Q can easily grasp the current state of the monitoring area 120 by checking the flattened image displayed on the display device 300.

[0155] Furthermore, flattened images are more suitable for image recognition processing than wide-angle images, and therefore, image processing device 100 according to the present embodiment can improve the accuracy of detecting characteristic portions that show predetermined characteristics.

[0156] [Variation 1] When the predetermined characteristic is a specific behavior of a person and the person having the predetermined characteristic stops performing the specific behavior, the image processing device 100 may instruct the display device 300 to stop displaying the flattened image in which the person appears. When the predetermined characteristic is a specific behavior of a person and the person having the predetermined characteristic stops performing the specific behavior, the image processing device 100 may continue to generate flattened images in which the person appears until a predetermined event occurs, and instruct the display device 300 to display the flattened images. An example of the predetermined event is the person performing another specific behavior.

[0157] When a predetermined characteristic is a specific state of an object, and an object having the predetermined characteristic is no longer in the specific state, the image processing device 100 may instruct the display device 300 to stop displaying a flattened image in which the object appears. When a predetermined characteristic is a specific state of an object, and an object having the predetermined characteristic is no longer in the specific state, the image processing device 100 may continue to generate a flattened image in which the object appears until a predetermined event occurs, and instruct the display device 300 to display the flattened image. An example of a predetermined event is when the object enters another specific state.

[0158] [Variation 2] If a second object having predetermined characteristics is detected within a predetermined period of time after a first object having predetermined characteristics is detected, the generation unit 154 may generate a flattened image showing the first object, but may not generate a flattened image showing the second object. The predetermined period is, for example, a period of time required for the user Q to recognize the first object. This prevents multiple flattened images from being generated simultaneously, thereby suppressing chattering.

[0159] [Additional Notes] The above-described embodiment and modifications include the following technical ideas.

[0160] [Configuration 1] An image processing device comprising: a detection unit that detects a characteristic portion in a wide-angle image that shows a predetermined characteristic; a generation unit that generates a flattened image obtained by flattening a target area in the wide-angle image that includes the characteristic portion based on predetermined display specification information; and an output unit that outputs video data representing the flattened image.

[0161] [Configuration 2] The image processing device according to Configuration 1, wherein the feature is a specific identity of a person, a specific behavior of a person, a specific identity of an object, or a specific state of an object.

[0162] [Configuration 3] The image processing device according to configuration 1 or 2, wherein the display specification information includes a display position of the object having the characteristic in the flattened image and an occupancy rate of the object in the flattened image.

[0163] [Configuration 4] The image processing device according to Configuration 3, wherein the object is moving, and the generation unit adjusts the display position of the object in the flattened image based on a moving direction and a moving speed of the object.

[0164] [Configuration 5] The image processing device according to Configuration 1 or 2, wherein the generating unit generates the flattened image so that an object having the characteristic is highlighted.

[0165] [Configuration 6] The image processing device according to any one of configurations 1 to 5, further comprising a warning unit that instructs a warning device to output an alert when the characteristic portion is detected.

[0166] [Configuration 7] The image processing device according to Configuration 1 or 2, wherein the detection unit generates a feature detection image used to detect the feature portion from the wide-angle image, and the feature detection image is an image obtained by flattening the wide-angle image.

[0167] [Configuration 8] The image processing device described in Configuration 7, wherein the generation unit converts first coordinates indicating a first display position of the object having the feature in the feature detection image into second coordinates indicating a second display position of the object in the wide-angle image, and determines the object area based on the second coordinates and the display specification information.

[0168] [Configuration 9] An image processing method comprising: detecting a characteristic portion in a wide-angle image that shows a predetermined characteristic; generating a flattened image obtained by flattening a target area in the wide-angle image that includes the characteristic portion based on predetermined display specification information; and outputting video data representing the flattened image.

[0169] [Configuration 10] An image processing program that causes one or more computers to execute the image processing method according to Configuration 9.

[0170] [Configuration 11] An image processing system comprising the image processing device according to configuration 1 or 2, and a camera that captures the wide-angle image.

[0171] The embodiments disclosed herein should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined by the claims, not by the above description, and is intended to include all modifications within the meaning and scope of the claims.

[0172] 10, 11, 12 System, 15 Network, 100 Image processing device, 101 Processor, 102 Memory, 103 Storage, 104 Input interface, 105 Display interface, 106 Communication interface, 113 Program, 114 Management table, 115 Wide-angle image data, 116 Planarization parameters, 120 Monitoring area, 151 Reception unit, 152 Acquisition unit, 153 Detection unit, 154 Generation unit, 155 Output unit, 156 Warning unit, 199 Bus, 200 Input device, 300 Display device, 400 Camera, 500 Warning device, 501, 503 Speaker, 502 Warning light, 504 Light, 1500, 1510 Coordinate system, 1600 Fisheye image, 1601, 1602, R, R1, R2, R3, R4, R5, R6, R7, R8, R9, R10 Object, 1603, 1604 Object, 1610 Side, 1612 Fisheye lens, A Default field of view, AR Arrow, F1, F2 Center position, F3, F11, F12, F13, F14, F21, F22, F23, F24 Position, J1, J3, J4, J5, J6, J7 Wide-angle image, J11, J12, J13, J14, J41, J51, J61, J62, J71 Partial images, K11, K12, K13, K14, K15, K21, K24, K31, K41, K51, K61, K62, K71 Image for feature detection, K221 Image, L Reference distance, N, N1, N2, N3, N4, N5, N6, N7, N8, N9, N10, N11, N13, N14, N15, N16 Flattened image, N12 Intermediate image, P1 Person, Q User, W Flattened screen window size, X1, X2, X4, b1, b2, b3, b4, b5, b6 Bounding box, X3 Information, Z1, Z2 Target area, d, d1, d2, d3, d4, d5, d6, d7, d8, d9, d10, d11 Display specification information, e1 Area, e2 No entry area, e3 display frame, f1, f3, f4 dotted lines, f2 movement distance, w1, w2, w3, w4 window.

Claims

1. An image processing device comprising: a detection unit that detects a characteristic portion in a wide-angle image that shows a predetermined characteristic; a generation unit that generates a flattened image based on predetermined display specification information by flattening a target area in the wide-angle image that includes the characteristic portion; and an output unit that outputs image data representing the flattened image.

2. The image processing device of claim 1, wherein the feature is a particular identity of a person, a particular action of a person, a particular identity of an object, or a particular state of an object.

3. An image processing device according to claim 1 or 2, wherein the display specification information includes a display position of an object having the characteristic in the flattened image and an occupancy rate of the object in the flattened image.

4. The image processing device according to claim 3, wherein the object is moving, and the generation unit adjusts the display position of the object in the flattened image based on a moving direction and a moving speed of the object.

5. An image processing device according to claim 1 or 2, wherein the generating unit generates the flattened image so that an object having the characteristic is highlighted.

6. The image processing device according to claim 1 or 2, further comprising a warning unit that instructs a warning device to output an alert when the characteristic portion is detected.

7. An image processing device as described in claim 1 or 2, wherein the detection unit generates a feature detection image used to detect the characteristic portion from the wide-angle image, and the feature detection image is an image obtained by flattening the wide-angle image.

8. The image processing device described in claim 7, wherein the generation unit converts first coordinates indicating a first display position of the object having the feature in the feature detection image into second coordinates indicating a second display position of the object in the wide-angle image, and determines the object area based on the second coordinates and the display designation information.

9. An image processing method comprising: detecting a characteristic portion showing a predetermined feature from a wide-angle image; generating a flattened image by flattening a target area of ​​the wide-angle image including the characteristic portion based on predetermined display specification information; and outputting image data representing the flattened image.

10. An image processing program for causing one or more computers to execute the image processing method according to claim 9.

11. An image processing system comprising the image processing device according to claim 1 or 2 and a camera for acquiring the wide-angle image.

Citation Information

Patent Citations

  • Image monitoring system

    JP2005167604A

  • Fish-eye monitoring system

    JP2011061511A

  • Image processor and image processing method

    JP2015210702A

  • Moving image display device, moving image display method, and program

    JP2017034362A

  • Monitoring image processing device and monitoring image processing method

    JP2018042105A