Imaging processing device, imaging processing method, and program

The dual-frame imaging processing device enhances subject detection accuracy by optimizing frame rates and readout ranges for both display and detection frames, addressing limitations in existing imaging technologies.

WO2025197159A1PCT designated stage Publication Date: 2025-09-25FUJIFILM CORP
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/035861
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-18
Filing Date
2024-10-07
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing imaging technologies face challenges in accurately detecting subjects, particularly in dynamic environments, due to limitations in frame rates, readout ranges, and angles of view, which affect the precision of subject detection.

Method used

The proposed imaging processing device employs a dual-frame approach, where a first frame is optimized for display and a second frame is optimized for subject detection, with higher frame rates, narrower readout ranges, and tailored exposure methods to enhance subject detection accuracy.

Benefits of technology

This dual-frame strategy significantly improves subject detection accuracy by leveraging higher frame rates and focused readout ranges, enabling more precise subject identification and control in dynamic imaging scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024035861_25092025_PF_FP_ABST
    Figure JP2024035861_25092025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an imaging processing device, an imaging processing method, and a program capable of improving detection accuracy for a subject. This imaging processing device comprises a processor. The processor outputs a first frame for displaying at least a video on the basis of first imaging data read from an imaging element included in an imaging device, and outputs a second frame for detecting a subject on the basis of second imaging data read from the imaging element. A factor pertaining to the detection of the subject in the second frame is different from a factor in the first frame.
Need to check novelty before this filing date? Find Prior Art

Description

Image capture processing device, image capture processing method, and program

[0001] The technology of the present disclosure relates to an imaging processing device, an imaging processing method, and a program.

[0002] Japanese Patent Laid-Open Publication No. 2006-217355 discloses an imaging device for capturing an image of a moving subject. The disclosed imaging device includes an imaging unit that receives a subject image formed on the subject and photoelectrically converts the image to output an imaging signal, a movement detection unit that detects a feature amount indicating the movement state of the subject, and a unit that changes a readout mode of the imaging signal from the imaging unit in accordance with the movement speed or acceleration of the subject based on the feature amount of the subject image obtained from the movement detection unit.

[0003] Japanese Patent Application Laid-Open Publication No. 2021-192498 discloses an imaging device that includes an imaging means for imaging a subject, an estimation means for estimating the speed of the subject at the time of capturing subsequent images from the image of the subject captured by the imaging means using a trained model created by machine learning, and a control means for controlling the imaging operation of the imaging means for capturing subsequent images based on the estimated speed of the subject estimated by the estimation means.

[0004] Japanese Patent Application Laid-Open No. 2008-135862 discloses an image processing system that includes an imaging device with an electronic zoom function that captures images based on imaging conditions using parameters such as the magnification and center position of the electronic zoom and outputs image data of the captured image, a control means that sets a plurality of imaging conditions and controls the imaging device by switching between the set plurality of imaging conditions in a time-division manner, and an image synthesis means that uses a plurality of image data from the imaging device that correspond to each of the plurality of imaging conditions to synthesize a plurality of images into a single screen configuration.

[0005] International Publication No. WO 2016 / 125351 discloses an operating device that transmits control information for an automatic tracking function to a camera equipped with the automatic tracking function. The disclosed operating device includes a receiving unit that receives images continuously transmitted from the camera, a tracking target receiving unit that receives an identification of a tracking target to be tracked by the camera, a movement amount calculation unit that calculates an amount of movement of the tracking target between consecutive images based on the images continuously received by the receiving unit, a movement amount determination unit that determines whether the movement amount calculated by the movement amount calculation unit is equal to or greater than a first movement amount threshold, and a size instruction transmission unit that, when the movement amount determination unit determines that the movement amount is equal to or greater than the first movement amount threshold, transmits an instruction to the camera to change the size of images to be transmitted by the camera from a first size to a second size smaller than the first size.

[0006] Japanese Patent Laid-Open Publication No. 2017-108303 discloses an image processing device used in an imaging device equipped with an image sensor having multiple photoelectric conversion units. The disclosed image processing device includes an image generation unit that generates a captured image using outputs from the multiple photoelectric conversion units, a phase difference information acquisition unit that acquires phase difference information using outputs from multiple specific photoelectric conversion units that are part of the multiple photoelectric conversion units, a subject detection unit that detects a subject included in the captured image, a size detection unit that detects an image size of the detected subject, and a selection unit that selects from the multiple specific photoelectric conversion units. The subject detection unit is capable of detecting a first subject and a second subject that is part of the first subject as subjects, and the selection unit selects the specific photoelectric conversion unit according to an area in the captured image that includes the first subject when the image size of the first subject is smaller than a predetermined size, and selects the specific photoelectric conversion unit according to an area in the captured image that includes the second subject when the image size of the first subject is larger than the predetermined size.

[0007] One embodiment of the technique of the present disclosure provides an image capturing processing device, an image capturing processing method, and a program that can improve the detection accuracy of a subject.

[0008] A first aspect of the technology of the present disclosure is an imaging processing device that includes a processor, the processor outputs at least a first frame for displaying an image based on first imaging data read out from an imaging element included in the imaging device, and outputs a second frame for detecting a subject based on second imaging data read out from the imaging element, and factors related to the detection of the subject in the second frame are different from factors in the first frame.

[0009] A second aspect of the technique of the present disclosure is the imaging processing device according to the first aspect, in which the factors include a frame rate.

[0010] A third aspect of the technology of the present disclosure is an imaging processing device according to the second aspect, wherein the second frame rate, which is the frame rate of the second frame, is higher than the first frame rate, which is the frame rate of the first frame.

[0011] A fourth aspect of the technology of the present disclosure is an imaging processing device according to any one of the first to third aspects, wherein the factors include at least one of the readout range, angle of view, data volume, and bit depth of the imaging element.

[0012] A fifth aspect of the technology of the present disclosure is an imaging processing device according to the fourth aspect, wherein the factors include a readout range, and the second readout range, which is the readout range corresponding to the second frame, is narrower than the first readout range, which is the readout range corresponding to the first frame.

[0013] A sixth aspect of the technology of the present disclosure is an imaging processing device according to the fourth or fifth aspect, wherein the factor includes an angle of view, and the second angle of view, which is the angle of view corresponding to the second frame, is narrower than the first angle of view, which is the angle of view corresponding to the first frame.

[0014] A seventh aspect of the technology of the present disclosure is an imaging processing device according to any one of the fourth to sixth aspects, wherein the factors include a readout range, and the second readout range, which is the readout range corresponding to the second frame, is determined based on at least one of the subject detection result, the processing performance of the processor, and the communication performance of the imaging processing device.

[0015] An eighth aspect of the technology of the present disclosure is an imaging processing device according to any one of the fourth to seventh aspects, wherein the factors include an angle of view, and the second angle of view, which is the angle of view corresponding to the second frame, is determined based on at least one of the subject detection result, the processing performance of the processor, and the communication performance of the imaging processing device.

[0016] A ninth aspect of the technique of the present disclosure is the imaging processing device according to the seventh or eighth aspect, wherein the imaging processing device is communicably connected to a management device via a network device, and the communication performance of the imaging processing device is determined by the performance of at least one of the network device and the management device.

[0017] A tenth aspect of the technology of the present disclosure is an imaging processing device according to the seventh or eighth aspect, wherein the imaging processing device is communicatively connected to a learning model device having a learning model, and the communication performance of the imaging processing device is determined by the performance of the learning model device.

[0018] An eleventh aspect of the technology of the present disclosure is an imaging processing device according to any one of the first to tenth aspects, wherein the processor sets either a first output state in which the first frame is output or a second output state in which the first frame and the second frame are selectively output based on the imaging conditions.

[0019] A twelfth aspect of the technology of the present disclosure is an imaging processing device according to the eleventh aspect, wherein the imaging conditions include a first condition that the exposure time of the imaging element is longer than a time calculated as the reciprocal of a first frame rate, which is the frame rate of the first frame, and the processor sets a first output state when the first condition is met.

[0020] A thirteenth aspect of the technology of the present disclosure is an imaging processing device according to the eleventh or twelfth aspect, wherein the imaging conditions include conditions relating to at least one of the exposure time of the imaging element, the video format, the processing performance of the processor, and the communication performance of the imaging processing device.

[0021] A fourteenth aspect of the technology of the present disclosure is an imaging processing device according to the thirteenth aspect, wherein the imaging conditions include the communication performance of the imaging processing device, the imaging processing device is communicatively connected to a management device via a network device, the processor outputs a first frame to the management device via the network device, outputs a second frame to a learning model of the imaging processing device, and outputs a subject detection result of the learning model to the management device via the network device, and the communication performance of the imaging processing device is determined by the performance of at least one of the network device and the management device.

[0022] A fifteenth aspect of the technology of the present disclosure is an imaging processing device according to the thirteenth aspect, wherein the imaging conditions include the communication performance of the imaging processing device, the imaging processing device is communicatively connected to a learning model device having a learning model, the processor outputs a first frame to the learning model device and outputs a second frame to the learning model device, and the communication performance of the imaging processing device is determined by the performance of the learning model device.

[0023] A sixteenth aspect of the technology of the present disclosure is an imaging processing device according to any one of the first to fifteenth aspects, wherein the processor changes the exposure method of the imaging device between the first frame and the second frame.

[0024] A 17th aspect of the technology of the present disclosure is an imaging processing device according to the 16th aspect, wherein the exposure method of the imaging device corresponding to the first frame is a method of setting exposure based on an average exposure amount of a plurality of photosensitive pixels that constitute the light receiving surface of the imaging element, and the exposure method of the imaging device corresponding to the second frame is a method of setting exposure based on the exposure amount of a photosensitive pixel that corresponds to a subject among the plurality of photosensitive pixels.

[0025] An 18th aspect of the technology of the present disclosure is an imaging processing device relating to any one of the first to seventeenth aspects, wherein the second frame is a frame in which the readout range of the imaging element is narrowed to an area including the subject compared to the first frame.

[0026] A 19th aspect of the technology of the present disclosure is an imaging processing device according to any one of the first to eighteenth aspects, in which the pixel resolution of the imaging element when outputting the second frame corresponds to the pixel resolution of the imaging element when outputting the first frame.

[0027] A twentieth aspect of the technology of the present disclosure is an imaging processing device according to any one of the first to nineteenth aspects, in which the resolution of the second frame corresponds to the resolution of the first frame.

[0028] A twenty-first aspect of the technique of the present disclosure is an imaging processing device according to any one of the first to twentieth aspects, in which the resolution of the second frame is lower than the resolution of the first frame.

[0029] A 22nd aspect of the technology of the present disclosure is an imaging processing device relating to any one of the first to 21st aspects, wherein the second frame is a frame input to a second learning model.

[0030] A 23rd aspect of the technology of the present disclosure is an imaging processing device according to the 22nd aspect, in which the processor generates a second frame to be input to the second learning model by performing a trimming process on the read image indicated by the second imaging data to cut out an area including the subject.

[0031] A 24th aspect of the technique of the present disclosure is the imaging processing device according to the 22nd or 23rd aspect, wherein the first frame is a frame input to a first learning model.

[0032] A 25th aspect of the technique of the present disclosure is the imaging processing device according to the 24th aspect, wherein the first learning model and the second learning model are a common learning model.

[0033] A 26th aspect of the technique of the present disclosure is the imaging processing device according to the 24th aspect, wherein the first learning model is a learning model different from the second learning model.

[0034] A 27th aspect of the technology of the present disclosure is an imaging processing device according to the 24th aspect, in which the processor performs a trimming process to cut out an area including the subject from the first frame, which is a read image indicated by the first imaging data, thereby generating a third frame to be input to the third learning model.

[0035] A 28th aspect of the technique of the present disclosure is the imaging processing device according to the 27th aspect, wherein the first learning model and the third learning model are a common learning model.

[0036] A 29th aspect according to the technique of the present disclosure is the imaging processing device according to the 27th aspect, in which the first learning model is a learning model different from the third learning model.

[0037] A 30th aspect of the technology of the present disclosure is an imaging processing device according to any one of the first to twenty-ninth aspects, wherein the processor stops outputting the second frame and performs electronic shake correction processing on the first frame when the amplitude of vibration acting on the imaging device is greater than a threshold value.

[0038] A 31st aspect of the technology of the present disclosure is an imaging processing device according to any one of the first to 30th aspects, wherein the imaging element is an imaging element capable of imaging at a frame rate higher than a first frame rate, which is the frame rate of the first frame.

[0039] A 32nd aspect of the technology of the present disclosure is an imaging processing method that includes outputting at least a first frame for displaying an image based on first imaging data read out from an imaging element provided in an imaging device, and outputting a second frame for detecting a subject based on second imaging data read out from the imaging element, wherein factors related to the detection of the subject in the second frame are different from factors in the first frame.

[0040] A 33rd aspect of the technology of the present disclosure is a program for causing a computer to execute processing, the processing including outputting at least a first frame for displaying an image based on first imaging data read out from an imaging element provided in an imaging device, and outputting a second frame for detecting a subject based on second imaging data read out from the imaging element, wherein factors related to the detection of the subject in the second frame are different from factors in the first frame.

[0041] The terms "first," "second," and "third" are used to identify components and do not specify the number of components.

[0042] According to the technology of the present disclosure, an imaging processing device, an imaging processing method, and a program are provided that can improve the detection accuracy of a subject.

[0043] 1 is a perspective view showing an example of a monitoring system according to a first embodiment. FIG. 2 is a block diagram showing an example of the hardware configuration of a monitoring camera according to the first embodiment. FIG. 3 is a block diagram showing an example of the functional configuration of a computer according to the first embodiment. FIG. 4 is an explanatory diagram showing an example of the operation of an imaging processing unit according to the first embodiment. FIG. 5 is an explanatory diagram showing an example of a mode in which video display frames and subject detection frames according to the first embodiment are selectively output as time passes. FIG. 6 is an explanatory diagram showing an example of the operation of a subject detection processing unit according to the first embodiment. FIG. 7 is a flowchart showing an example of the flow of imaging processing according to the first embodiment. FIG. 8 is a flowchart showing an example of the flow of subject detection processing according to the first embodiment. FIG. 9 is an explanatory diagram showing an example of a mode in which video display frames and subject detection frames according to the second embodiment are selectively output as time passes. FIG. 10 is an explanatory diagram showing an example of a mode in which video display frames and subject detection frames according to the third embodiment are selectively output as time passes. FIG. 11 is an explanatory diagram showing an example of a mode in which video display frames and subject detection frames according to a fourth embodiment are selectively output as time passes. FIG. 12 is an explanatory diagram showing an example of a mode in which video display frames and subject detection frames according to a modification of the fourth embodiment are selectively output as time passes. FIG. 13 is a block diagram showing an example of the hardware configuration of a monitoring camera according to a fifth embodiment. FIG. 14 is an explanatory diagram showing an example of a mode in which video display frames and subject detection frames according to the fifth embodiment are selectively output. FIG. 10 is an explanatory diagram showing an example of a manner in which a video display frame according to a fifth embodiment is output. FIG. 11 is an explanatory diagram showing an example of an operation of an exposure setting unit according to a sixth embodiment. FIG. 12 is an explanatory diagram showing an example of an operation of an imaging control unit according to a seventh embodiment. FIG. 13 is an explanatory diagram showing an example of an operation of an imaging processing unit according to an eighth embodiment. FIG. 14 is a block diagram showing an example of a monitoring system according to a ninth embodiment. FIG. 15 is an explanatory diagram showing an example of a manner in which a video display frame and a subject detection frame according to a comparative example are selectively output as time passes.

[0044] Hereinafter, exemplary embodiments of an imaging processing device, an imaging processing method, and a program according to the techniques of the present disclosure will be described with reference to the accompanying drawings.

[0045] First, the terms used in the following description will be explained.

[0046] UI is an abbreviation for "User Interface." I / F is an abbreviation for "Interface." CPU is an abbreviation for "Central Processing Unit." RAM is an abbreviation for "Random Access Memory." EEPROM is an abbreviation for "Electrically Erasable and Programmable Read Only Memory." HDD is an abbreviation for "Hard Disk Drive." SSD is an abbreviation for "Solid State Drive." CMOS is an abbreviation for "Complementary Metal Oxide Semiconductor." A / D is an abbreviation for "Analog / Digital." WAN is an abbreviation for "Wide Area Network." EL is an abbreviation for "Electro Luminescence." ASIC is an abbreviation for "Application Specific Integrated Circuit." FPGA is an abbreviation for "Field-Programmable Gate Array." PLD is an abbreviation for "Programmable Logic Device." GPU is an abbreviation for "Graphics Processing Unit." TPU is an abbreviation for "Tensor Processing Unit." USB is an abbreviation for "Universal Serial Bus." SoC is an abbreviation for "System-on-a-Chip." IC is an abbreviation for "Integrated Circuit."

[0047] In the description of this specification, "same" refers not only to the exact sameness but also to the sameness including a tolerance generally accepted in the technical field to which the technology of the present disclosure belongs.

[0048] First Embodiment First, the first embodiment will be described.

[0049] 1 shows an example of a monitoring system S according to a first embodiment. The monitoring system S includes a monitoring camera 10 and a management device 12. The monitoring camera 10 is an example of an "imaging device" according to the technology of the present disclosure.

[0050] The surveillance camera 10 is installed on a building, a wall, a pillar, or the like, and captures images of a monitored area 1 (see FIG. 2 ) to generate video (i.e., moving images). The surveillance camera 10 transmits the generated video to the management device 12 via a communication line 14 or the like.

[0051] The management device 12 includes a display 16 and a secondary storage device 18. The display 16 may be, for example, a liquid crystal display or an organic EL display. The secondary storage device 18 may be, for example, a hard disk drive (HDD). The secondary storage device 18 may not be an HDD, but may be a non-volatile memory such as a flash memory, an SSD, or an EEPROM.

[0052] The management device 12 receives the video images transmitted by the surveillance camera 10, and displays the received video images on the display 16 and stores them in the secondary storage device 18. The surveillance camera 10 also detects a subject 2 (see FIG. 2 ) included in the monitored area 1, and transmits the results of detecting the subject 2 to the management device 12. The management device 12 receives the results transmitted from the surveillance camera 10, and controls the panning, tilting, zooming, tracking imaging, and other aspects of the surveillance camera 10 based on the received results.

[0053] The surveillance camera 10 is attached to a rotation mechanism 20. The rotation mechanism 20 supports the surveillance camera 10 so that it can rotate. A pitch axis P, a yaw axis Y, and a roll axis R are set for the surveillance camera 10. The rotation mechanism 20 is a two-axis rotation mechanism that supports the surveillance camera 10 so that it can rotate around the pitch axis P and the yaw axis Y. The pan function of the surveillance camera 10 is achieved by rotating the surveillance camera 10 around the yaw axis Y. The tilt function of the surveillance camera 10 is achieved by rotating the surveillance camera 10 around the pitch axis P. Note that the rotation mechanism 20 may be a three-axis rotation mechanism in addition to a two-axis rotation mechanism.

[0054] 2 shows an example of the hardware configuration of the surveillance camera 10 according to the first embodiment. The surveillance camera 10 includes an optical system 30 and an image sensor 32. The image sensor 32 is located after the optical system 30. The optical system 30 includes an objective lens 36 and a lens group 38. The objective lens 36 and the lens group 38 are arranged in this order along the optical axis OA of the optical system 30, from the subject 2 side (object side) to the light receiving surface 32A side (image side) of the image sensor 32. The lens group 38 includes a zoom lens, a focus lens, an aperture, an imaging lens, etc.

[0055] The surveillance camera 10 includes a computer 40, a UI device 42, and a communication I / F 44. The computer 40 is an example of an "imaging processing device" according to the technology of the present disclosure. The computer 40 includes a CPU 50, a memory 52, and a storage 54. The image sensor 32, the UI device 42, the communication I / F 44, the CPU 50, the memory 52, and the storage 54 are connected to each other via a bus 56 so as to be able to communicate with each other.

[0056] The CPU 50 reads various programs from the storage 54 and executes the read programs on the memory 52, thereby controlling the entire surveillance camera 10. The CPU 50 is an example of a "processor" according to the technology of the present disclosure.

[0057] The memory 52 temporarily stores various types of information and is used as a work memory. An example of the memory 52 is a RAM, but the memory 52 is not limited to this and may be other types of storage devices.

[0058] The storage 54 is a non-volatile storage device. Here, a flash memory is used as an example of the storage 54. The flash memory is merely an example, and the storage 54 may be, for example, a variety of non-volatile memories such as a magnetoresistive memory and / or a ferroelectric memory instead of or in addition to the flash memory. The non-volatile storage device may also be an EEPROM, a HDD, and / or an SSD. Various programs for the surveillance camera 10 are stored in the storage 54.

[0059] The imaging element 32 is, for example, a CMOS image sensor. The imaging element 32 captures an image of the monitoring target area 1 at a predetermined frame rate under the instruction of the CPU 50. The imaging element 32 has a light receiving surface 32A.

[0060] The light receiving surface 32A is formed by a plurality of photosensitive pixels 34 arranged in a matrix. In the imaging element 32, each photosensitive pixel 34 is exposed to light, and photoelectric conversion is performed for each photosensitive pixel 34. The charge obtained by photoelectric conversion for each photosensitive pixel 34 is an analog imaging signal that indicates the monitored area 1. The imaging element 32 performs signal processing such as A / D conversion on the analog imaging signal to generate a digital image, which is a digital imaging signal. Hereinafter, the digital image data will be referred to as "imaging data." The plurality of photosensitive pixels 34 may be a plurality of photoelectric conversion elements that are sensitive to visible light, or a plurality of photoelectric conversion elements that are sensitive to infrared light.

[0061] The communication I / F 44 is, for example, a network interface, and controls the transmission of various information between the management device 12 and the management device 12 via a network including the network device 22. An example of the network is a WAN such as the Internet or a public communication network.

[0062] The UI device 42 includes a reception device 46 and a display 48. The reception device 46 is, for example, a hard key or a touch panel, and receives various instructions from a user of the monitoring system S. The display 48 displays various information under the control of the CPU 50. Examples of the information displayed on the display 48 include the contents of the various instructions received by the reception device 46 and captured images. Examples of the display 48 include a liquid crystal display or an organic EL display.

[0063] The monitored area 1 is an area imaged by the surveillance camera 10. The monitored area 1 includes a subject 2. The subject 2 is, for example, a moving object such as a person. An example of imaging the monitored area 1 including the subject 2 by the surveillance camera 10 will be described below.

[0064] 3 shows an example of the functional configuration of the computer 40 according to the first embodiment. A program 58 is stored in the storage 54. The CPU 50 reads the program 58 from the storage 54 and executes the read program 58 on the memory 52. ​​The program 58 is an example of a "program" according to the technology of the present disclosure. The CPU 50 executes an image capture process and a subject detection process in accordance with the program 58 executed on the memory 52.

[0065] The imaging process is a process of imaging the monitored area 1 (see Figure 2), generating frames for displaying an image including the monitored area 1 and frames for detecting a subject 2 (see Figure 2) included in the monitored area 1, and outputting the generated frames.

[0066] Hereinafter, a frame for displaying an image including the area to be monitored 1 will be referred to as an image display frame 120 (see FIGS. 4 and 5). Also, a frame for detecting a subject 2 included in the area to be monitored 1 will be referred to as a subject detection frame 122 (see FIGS. 4 and 5). The image display frame 120 is an example of a "first frame" according to the technology of the present disclosure. The subject detection frame 122 is an example of a "second frame" according to the technology of the present disclosure.

[0067] The imaging process is realized by the CPU 50 operating as the imaging processing section 60. The imaging processing section 60 has an imaging determination section 62, a first imaging processing section 64, and a second imaging processing section 66.

[0068] The subject detection process is a process for detecting a subject 2 based on the subject detection frame 122 generated in the imaging process. The subject detection process is realized by the CPU 50 operating as a subject detection processing unit 70. The subject detection processing unit 70 has a frame input unit 72 and a subject detection unit 74.

[0069] A learning model 82 is stored in the storage 54. The learning model 82 is an example of a "second learning model" according to the technology of the present disclosure. The learning model 82 is constructed by a learning method such as deep learning using training data in which the subject detection frame 122 is input data and the result of detecting the subject 2 (hereinafter also referred to as the "subject detection result") is output data. For example, a neural network or the like is used for the learning model 82. The subject detection result includes, for example, a result of determining the type, features, size, or properties of the subject 2.

[0070] 4 shows an example of the operation of the imaging processing unit 60 according to the first embodiment. The imaging determination unit 62 determines, based on imaging conditions, whether to set a first output state in which the image display frame 120 (for example, only the image display frame 120) is output, or a second output state in which the image display frame 120 and the subject detection frame 122 are selectively output. The imaging conditions include, for example, a condition related to the exposure time of the image sensor 32. Specifically, the imaging conditions include a condition that the exposure time of the image sensor 32 is longer than a threshold value. The threshold value is set to, for example, a time calculated as the reciprocal of the frame rate of the image display frame 120.

[0071] The frame rate of the subject detection frames 122 is set to a frame rate higher than the frame rate of the video display frames 120. For example, the frame rate of the video display frames 120 is set to 30 fps (frames per second), and the frame rate of the subject detection frames 122 is set to 60 fps. Alternatively, for example, the frame rate of the video display frames 120 is set to 60 fps, and the frame rate of the subject detection frames 122 is set to 120 fps. The frame rates of the video display frames 120 and the subject detection frames 122 may be set to frame rates other than those described above. The frame rate of the video display frames 120 is an example of a "first frame rate" according to the technology of the present disclosure. The frame rate of the subject detection frames 122 is an example of a "second frame rate" according to the technology of the present disclosure.

[0072] The image pickup element 32 is an image pickup element 32 that can capture images at a frame rate higher than the frame rate of the video display frame 120 .

[0073] The exposure time of the image sensor 32 is determined by, for example, the shutter speed. The shutter speed may be that of a mechanical shutter or that of an electronic shutter. The surveillance camera 10 is configured to be able to change the shutter speed.

[0074] As an example, when a first condition is met that the exposure time of the image sensor 32 is longer than a threshold, the imaging determination unit 62 determines to set a first output state in which the image display frame 120 is output. On the other hand, when a second condition is met that the exposure time of the image sensor 32 is equal to or shorter than a threshold, the imaging determination unit 62 determines to set a second output state in which the subject detection frame 122 is output. When the imaging determination unit 62 determines to set the first output state, the first imaging processing unit 64 executes a first imaging process, which is a process for outputting the image display frame 120. On the other hand, when the imaging determination unit 62 determines to set the second output state, the second imaging processing unit 66 executes a second imaging process, which is a process for selectively outputting the image display frame 120 and the subject detection frame 122.

[0075] The first imaging processing unit 64 has a first imaging control unit 64A and a first frame output unit 64B. The first imaging control unit 64A controls the imaging element 32 to read imaging data for video display. Specifically, the first imaging control unit 64A controls the imaging element 32 to read imaging data for video display from photosensitive pixels 34 (hereinafter referred to as "effective photosensitive pixels") that correspond to the resolution of the video display frame 120, among the multiple photosensitive pixels 34 that the imaging element 32 has. The effective photosensitive pixels may be all or some of the multiple photosensitive pixels 34 that the imaging element 32 has. The imaging data for video display is an example of "first imaging data" according to the technology of the present disclosure.

[0076] The first frame output unit 64B generates a video display frame 120 for displaying a video based on the image data for video display read out from the valid photosensitive pixels. The first frame output unit 64B then outputs the generated video display frame 120 to the management device 12. As a result, the video display frame 120 (i.e., a video) including the monitored area 1 (see FIG. 2 ) is displayed on the display 16 of the management device 12.

[0077] The second imaging processing unit 66 has a second imaging control unit 66A, a frame determination unit 66B, a trimming unit 66C, and a second frame output unit 66D. The second imaging control unit 66A executes control over the imaging element 32 to selectively read out imaging data for video display and imaging data for subject detection. The control to selectively read out imaging data for video display and imaging data for subject detection may be control to alternately read out imaging data for video display and imaging data for subject detection, control to alternately read out a plurality of imaging data for video display and imaging data for subject detection, or control to alternately read out imaging data for video display and a plurality of imaging data for subject detection.

[0078] Under the control of the second imaging control unit 66A, the angle of view of the surveillance camera 10 is kept the same when reading out imaging data for video display and when reading out imaging data for subject detection. The control by the second imaging control unit 66A to read out imaging data for video display on the imaging element 32 is similar to the control by the first imaging control unit 64A to read out imaging data for video display on the imaging element 32.

[0079] The second imaging control unit 66A controls the image sensor 32 to read imaging data for subject detection as follows. That is, the second imaging control unit 66A identifies the subject 2 appearing in the image display frame 120 by performing image processing for identifying the subject 2 on the video display frame 120 generated by the first imaging processing unit 64 or the second imaging processing unit 66 in the previous routine (i.e., the routine from step ST10 to step ST26 described below). The second imaging control unit 66A also specifies a row of valid photosensitive pixels including an area where light from the subject 2 is focused, and controls the image sensor 32 to read imaging data for subject detection from the photosensitive pixels 34 in the specified row. Examples of control for reading imaging data for subject detection from the photosensitive pixels 34 in the specified row include a crop function. The imaging data for subject detection is an example of "second imaging data" according to the technology of the present disclosure.

[0080] The frame determination unit 66B determines whether the second imaging control unit 66A has executed control for the imaging element 32 to read imaging data for image display, or whether it has executed control for the imaging element 32 to read imaging data for subject detection. Here, if the frame determination unit 66B determines that the second imaging control unit 66A has executed control for the imaging element 32 to read imaging data for image display, the second frame output unit 66D generates an image display frame 120 for displaying an image based on the imaging data for image display read from the valid photosensitive pixels. The second frame output unit 66D then outputs the generated image display frame 120 to the management device 12. As a result, the image display frame 120 (i.e., an image) including the monitored area 1 (see FIG. 2 ) is displayed on the display 16 of the management device 12.

[0081] On the other hand, when the frame determination unit 66B determines that the second imaging control unit 66A has executed control of the image sensor 32 to read out imaging data for subject detection, the trimming unit 66C executes a trimming process to cut out an image 112A including the subject 2 from the readout image 112 (see FIG. 5 ) indicated by the imaging data for subject detection, thereby generating a subject detection frame 122 indicating an area including the subject 2 from the readout image 112. Then, the second frame output unit 66D outputs the subject detection frame 122 to the subject detection processing unit 70. For example, the second frame output unit 66D stores the subject detection frame 122 in the memory 52 or the storage 54 and notifies the subject detection processing unit 70 by, for example, setting a flag in the memory 52 that the subject detection frame 122 has been stored in the memory 52 or the storage 54.

[0082] 5 shows an example of how the video display frame 120 and the subject detection frame 122 according to the first embodiment are selectively output over time. The subject detection frame 122 is a frame that is input to the learning model 82 for detecting the subject 2. Therefore, in order to improve the detection accuracy when the learning model 82 is used to detect the subject 2, it is preferable that the factors related to the detection of the subject 2 in the subject detection frame 122 are set to values ​​that can improve the detection accuracy when the learning model 82 is used to detect the subject 2.

[0083] The factors related to the detection of the subject 2 are, for example, factors related to the detection accuracy of the subject 2. The factors related to the detection accuracy of the subject 2 in the subject detection frames 122 are different from the factors in the video display frames 120. The factors may include, for example, the frame rate. As described above, the frame rate of the subject detection frames 122 is higher than the frame rate of the video display frames 120.

[0084] The factors may also include, for example, the readout range of the image sensor 32. The readout range is defined, for example, by the number of rows of the plurality of photosensitive pixels 34. The readout range 102 corresponding to the subject detection frame 122 is narrower than the readout range 100 corresponding to the video display frame 120. In other words, the subject detection frame 122 is a frame in which the readout range of the image sensor 32 is narrowed to an area including the subject 2 compared to the video display frame 120. The readout range 100 corresponding to the video display frame 120 is an example of a "first readout range" according to the technology of the present disclosure. The readout range 102 corresponding to the subject detection frame 122 is an example of a "second readout range" according to the technology of the present disclosure.

[0085] The pixel resolution of the image sensor 32 when outputting the subject detection frame 122 corresponds to the pixel resolution of the image sensor 32 when outputting the video display frame 120. As an example, the pixel resolution of the image sensor 32 when outputting the subject detection frame 122 is the same as the pixel resolution of the image sensor 32 when outputting the video display frame 120. The resolution of the subject detection frame 122 corresponds to the resolution of the video display frame 120. As an example, the resolution of the subject detection frame 122 is the same as the resolution of the video display frame 120.

[0086] Note that the pixel resolution of the image sensor 32 when outputting the subject detection frame 122 may be lower than the pixel resolution of the image sensor 32 when outputting the video display frame 120, as long as the accuracy of the subject detection result is ensured. Also, the resolution of the subject detection frame 122 may be lower than the resolution of the video display frame 120, as long as the accuracy of the subject detection result is ensured.

[0087] 6 shows an example of the operation of the subject detection processing unit 70 according to the first embodiment. The frame input unit 72 acquires the subject detection frame 122 output from the second frame output unit 66D, and inputs the acquired subject detection frame 122 to the learning model 82.

[0088] The subject detection unit 74 outputs to the learning model 82 the subject detection result corresponding to the subject detection frame 122 input to the learning model 82. Then, the subject detection unit 74 acquires the subject detection result output from the learning model 82 and outputs the subject detection result to the management device 12. As a result, in the management device 12, for example, based on the subject detection result, control of the surveillance camera 10, such as panning, tilting, zooming, and tracking imaging, is executed on the surveillance camera 10 in order to fit the subject 2 within the imaging range of the surveillance camera 10.

[0089] 7 shows an example of the flow of the imaging process according to the first embodiment. In the imaging process, first, in step ST10, the imaging determination unit 62 determines whether the exposure time of the imaging element 32 is longer than a threshold value. If the exposure time of the imaging element 32 is longer than the threshold value in step ST10, the determination is affirmative, and the imaging process proceeds to step ST12. On the other hand, if the exposure time of the imaging element 32 is equal to or shorter than the threshold value in step ST10, the determination is negative, and the imaging process proceeds to step ST16.

[0090] In step ST12, the first imaging control unit 64A executes control to read out imaging data for video display on the imaging element 32. After the process of step ST12 is executed, the imaging process proceeds to step ST14.

[0091] In step ST14, the first frame output unit 64B generates a video display frame 120 for displaying a video based on the imaging data for video display read out in step ST12, and outputs the generated video display frame 120 to the management device 12. As a result, the video display frame 120 including the monitored area 1 is displayed on the display 16 of the management device 12. After the processing of step ST14 is executed, the imaging processing proceeds to step ST26.

[0092] In step ST16, the second imaging control unit 66A executes control over the imaging element 32 to selectively read out imaging data for image display and imaging data for subject detection. As a result, either imaging data for image display or imaging data for subject detection is read out each time step ST16 is processed. After step ST16 is executed, the imaging process proceeds to step ST18.

[0093] In step ST18, the frame determination unit 66B determines whether or not control to read out imaging data for video display was executed in step ST16. If control to read out imaging data for video display was executed in step ST16, the determination in step ST18 is affirmative, and the imaging process proceeds to step ST20. On the other hand, if control to read out imaging data for subject detection was executed in step ST16, the determination in step ST18 is negative, and the imaging process proceeds to step ST22.

[0094] In step ST20, the second frame output unit 66D generates a video display frame 120 for displaying a video based on the imaging data for video display read out in step ST16, and outputs the generated video display frame 120 to the management device 12. As a result, the video display frame 120 including the monitored area 1 is displayed on the display 16 of the management device 12. After the processing of step ST20 is executed, the imaging processing proceeds to step ST26.

[0095] In step ST22, the trimming unit 66C performs a trimming process on the read image 112 indicated by the imaging data for subject detection read out in step ST16, to cut out an image 112A including the subject 2. As a result, a subject detection frame 122 indicating an area including the subject 2 is generated based on the imaging data for subject detection. After the process of step ST22 is performed, the imaging process proceeds to step ST24.

[0096] In step ST24, the second frame output section 66D outputs the subject detection frame 122 to the subject detection processing section 70. After the processing of step ST24 is executed, the imaging processing proceeds to step ST26.

[0097] In step ST26, the CPU 50 determines whether a termination condition, which is a condition for terminating the image capture process, is met. An example of the termination condition is that an end instruction is given to the surveillance camera 10 by the management device 12 or a user. If the termination condition is not met in step ST26, the determination is negative, and the image capture process proceeds to step ST10. On the other hand, if the termination condition is met in step ST26, the determination is positive, and the image capture process ends.

[0098] 8 shows an example of the flow of the subject detection process according to the first embodiment. In the subject detection process, first, in step ST30, the frame input unit 72 acquires the subject detection frame 122 generated in the imaging process, and inputs the acquired subject detection frame 122 into the learning model 82. After the process of step ST30 is executed, the subject detection process proceeds to step ST32.

[0099] In step ST32, the subject detection unit 74 causes the learning model 82 to output the subject detection result corresponding to the subject detection frame 122 input to the learning model 82. Then, the subject detection unit 74 acquires the subject detection result output from the learning model 82 and outputs the subject detection result to the management device 12. As a result, in the management device 12, for example, based on the subject detection result, control of the surveillance camera 10, such as panning, tilting, zooming, and tracking imaging, is executed on the surveillance camera 10 in order to fit the subject 2 within the imaging range of the surveillance camera 10. After the processing of step ST32 is executed, the subject detection processing ends.

[0100] The imaging processing method executed by the imaging process and the subject detection process is an example of the "imaging processing method" according to the technology of the present disclosure.

[0101] Next, the effects of the first embodiment will be described. First, a comparative example will be described to clarify the effects of the first embodiment.

[0102] 20 shows an example of a manner in which a video display frame 120 and a subject detection frame 122 according to a comparative example are selectively output over time. In the comparative example, the video display frame 120 for displaying a video is generated based on image data for video display read from the image sensor 32. The video display frame 120 is then resized to a resolution that can be input to the learning model 82, thereby generating the subject detection frame 122. In other words, the resolution of the subject detection frame 122 is lower than the resolution of the video display frame 120.

[0103] However, if the resolution of the subject detection frame 122 is lower than the resolution of the video display frame 120, the subject 2 captured in the subject detection frame 122 will appear smaller than in the video display frame 120, resulting in a decrease in the accuracy of the subject detection results for detecting the subject 2. On the other hand, if the video display frame 120 is generated by limiting the readout range of the image sensor 32 to the area surrounding the subject 2 in order to maintain the resolution of the subject detection frame 122, areas other than the area surrounding the subject 2 will not be captured in the video display frame 120, making it impossible to monitor the area that is actually desired to be monitored. Therefore, it is desirable to ensure the accuracy of the subject detection results while maintaining the area captured in the video display frame 120.

[0104] In this regard, in the first embodiment shown in FIG. 5 , the image display frame 120 for displaying an image is generated based on image display imaging data read from the image sensor 32. Therefore, compared to, for example, when the readout range of the image sensor 32 is limited to the range surrounding the subject 2, the range captured in the image display frame 120 can be maintained. Furthermore, in the first embodiment, image data for subject detection is read from a row of photosensitive pixels 34 that includes an area where light from the subject 2 is focused, among the multiple photosensitive pixels 34, and the subject detection frame 122 is generated based on the image detection imaging data. Therefore, compared to when the subject detection frame 122 is generated by resizing the image display frame 120 to a resolution that can be input to the learning model 82, the resolution of the subject detection frame 122 can be ensured, thereby ensuring the accuracy of the subject detection result. In this way, in the first embodiment, the range captured in the image display frame 120 can be maintained while ensuring the accuracy of the subject detection result.

[0105] Moreover, the frame rate of the subject detection frames 122 is higher than the frame rate of the video display frames 120. This allows the subject 2 to be detected more frequently than when the frame rate of the subject detection frames 122 is the same as the frame rate of the video display frames 120, for example, and therefore the accuracy of the subject detection results can be improved.

[0106] Furthermore, the readout range 102 corresponding to the subject detection frame 122 is narrower than the readout range 100 corresponding to the video display frame 120. This allows the amount of data in the subject detection frame 122 generated based on the imaging data for subject detection to be smaller than the amount of data in the video display frame 120 generated based on the imaging data for video display, thereby reducing the burden on the CPU 50 when detecting the subject 2 using the learning model 82. This ensures the accuracy of the subject detection results.

[0107] In the first embodiment, factors related to the detection accuracy of the subject 2 in the subject detection frame 122 include, for example, the frame rate and the readout range, but may also include the angle of view of the surveillance camera 10. The angle of view is represented, for example, by an angle corresponding to the imaging range. For example, the angle of view corresponding to the subject detection frame 122 may be narrower than the angle of view corresponding to the video display frame 120. If the angle of view corresponding to the subject detection frame 122 is narrower than the angle of view corresponding to the video display frame 120, the load on the CPU 50 when detecting the subject 2 using the learning model 82 can be reduced, thereby ensuring the accuracy of the subject detection result.

[0108] Furthermore, factors related to the detection accuracy of the subject 2 in the subject detection frame 122 may include, for example, the amount of data. The amount of data is expressed, for example, in bytes. For example, the amount of data in the subject detection frame 122 may be smaller than the amount of data in the video display frame 120. If the amount of data in the subject detection frame 122 is smaller than the amount of data in the video display frame 120, the burden on the CPU 50 can be reduced when detecting the subject 2 using the learning model 82. This ensures the accuracy of the subject detection results.

[0109] Furthermore, factors related to the detection accuracy of the subject 2 in the subject detection frame 122 may include, for example, bit depth. Bit depth is expressed, for example, by bpp (bits per pixel). For example, the bit depth corresponding to the subject detection frame 122 may be lower than the bit depth corresponding to the video display frame 120. If the bit depth corresponding to the subject detection frame 122 is lower than the bit depth corresponding to the video display frame 120, the amount of data in the subject detection frame 122 can be smaller than the amount of data in the video display frame 120, thereby reducing the burden on the CPU 50 when detecting the subject 2 using the learning model 82. This ensures the accuracy of the subject detection result.

[0110] Second Embodiment Next, a second embodiment will be described.

[0111] 9 shows an example of how the image display frame 120 and the object detection frame 122 according to the second embodiment are selectively output over time. The second embodiment is modified from the first embodiment as follows: In the second embodiment, in addition to the object detection frame 122, the image display frame 120 is also input to the learning model 82.

[0112] The resolution of the subject detection frame 122 is set to be the same as the resolution of the video display frame 120. The learning model 82 is constructed by a learning method such as deep learning using training data in which the video display frame 120 and the subject detection frame 122 are input data and the subject detection result is output data. The learning model 82 is an example of a "first learning model" and a "second learning model" according to the technology of the present disclosure. The learning model 82 then outputs a subject detection result corresponding to the input video display frame 120 and subject detection frame 122.

[0113] In the second embodiment, in addition to the subject detection frame 122, the video display frame 120 is also input to the learning model 82. Therefore, for example, the accuracy of the subject detection result can be improved compared to when only the subject detection frame 122 is input to the learning model 82.

[0114] Furthermore, the video display frame 120 and the subject detection frame 122 are input to a common learning model 82. Therefore, the number of learning models can be reduced compared to when the video display frame 120 and the subject detection frame 122 are input to different learning models.

[0115] Third Embodiment Next, a third embodiment will be described.

[0116] FIG. 10 shows an example of how a video display frame 120 and a subject detection frame 122 according to the third embodiment are selectively output over time. The third embodiment is modified from the second embodiment as follows. That is, in the third embodiment, a trimming process is performed to cut out an image 120A including a subject 2 from a video display frame 120, which is a read image indicated by imaging data for video display, to generate a learning model input frame 124. Then, in addition to the subject detection frame 122, the learning model input frame 124 is also input to the learning model 82. The learning model 82 is an example of a "first learning model" and a "third learning model" according to the technology of the present disclosure. The learning model input frame 124 is an example of a "third frame" according to the technology of the present disclosure.

[0117] The learning model input frame 124 is resized and digitally zoomed relative to the video display frame 120 so that the size of the subject 2 shown in the learning model input frame 124 is the same as the size of the subject 2 shown in the subject detection frame 122. For this reason, the subject 2 shown in the learning model input frame 124 is blurred compared to the subject 2 shown in the subject detection frame 122.

[0118] The learning model 82 is constructed by a learning method such as deep learning using training data in which the learning model input frame 124 and the subject detection frame 122 are input data and the subject detection result is output data. The learning model 82 then outputs the subject detection result corresponding to the input learning model input frame 124 and subject detection frame 122.

[0119] In the third embodiment, in addition to the subject detection frame 122, the learning model input frame 124 is also input to the learning model 82, so that the accuracy of the subject detection results can be improved compared to, for example, when only the subject detection frame 122 is input to the learning model 82.

[0120] Furthermore, a trimming process is performed on the video display frame 120 to cut out the image 120A including the subject 2, thereby generating a learning model input frame 124 to be input to the learning model 82. Therefore, the amount of data can be reduced compared to when the video display frame 120 is input to the learning model 82.

[0121] Furthermore, the video display frame 120 and the subject detection frame 122 are input to a common learning model 82. Therefore, the number of learning models can be reduced compared to when the video display frame 120 and the subject detection frame 122 are input to different learning models 82.

[0122] Fourth Embodiment Next, a fourth embodiment will be described.

[0123] 11 shows an example of how a video display frame 120 and a subject detection frame 122 according to the fourth embodiment are selectively output over time. The fourth embodiment is modified from the second embodiment as follows. That is, in the fourth embodiment, a learning model 80 is used in addition to a learning model 82. The video display frame 120 is input to the learning model 80, and the subject detection frame 122 is input to the learning model 82.

[0124] The learning model 80 is constructed by a learning method such as deep learning using training data in which the video display frame 120 is input data and the subject detection result is output data. The learning model 80 is an example of a "first learning model" according to the technology of the present disclosure. The learning model 80 outputs a subject detection result corresponding to the input video display frame 120, and the learning model 82 outputs a subject detection result corresponding to the input subject detection frame 122. Then, a subject detection result that combines the subject detection result output from the learning model 80 and the subject detection result output from the learning model 82 is output to the management device 12.

[0125] 11, the image display frame 120 output from the second image capturing processing unit 66 (see FIG. 4) is input to the learning model 80, but both the image display frame 120 output from the second image capturing processing unit 66 (see FIG. 4) and the image display frame 120 output from the first image capturing processing unit 64 (see FIG. 4) may be input to the learning model 80. Also, instead of the image display frame 120 output from the second image capturing processing unit 66 (see FIG. 4), the image display frame 120 output from the first image capturing processing unit 64 may be input to the learning model 80.

[0126] In the fourth embodiment, the learning model 80 to which the video display frame 120 is input is different from the learning model 82 to which the subject detection frame 122 is input. Therefore, the learning model 80 can output a subject detection result corresponding to the video display frame 120, and the learning model 82 can output a subject detection result corresponding to the subject detection frame 122. Furthermore, since it is possible to obtain a subject detection result that combines the subject detection result output from the learning model 80 and the subject detection result output from the learning model 82, the accuracy of the subject detection result can be improved compared to when a common learning model is used.

[0127] 12 shows an example of a mode in which a video display frame 120 and a subject detection frame 122 according to a modified example of the fourth embodiment are selectively output over time. The modified example of the fourth embodiment is modified as follows compared to the third embodiment. That is, in the modified example of the fourth embodiment, a learning model 80 is used in addition to a learning model 82. A learning model input frame 124 generated by performing a trimming process on the video display frame 120 is input to the learning model 80, and the subject detection frame 122 is input to the learning model 82.

[0128] The learning model 80 is constructed by a learning method such as deep learning using training data in which the learning model input frame 124 is input data and the subject detection result is output data. The learning model 80 is an example of a "third learning model" according to the technology of the present disclosure. The learning model 80 outputs a subject detection result corresponding to the input learning model input frame 124, and the learning model 82 outputs a subject detection result corresponding to the input subject detection frame 122. Then, a subject detection result that combines the subject detection result output from the learning model 80 and the subject detection result output from the learning model 82 is output to the management device 12.

[0129] 12, the video display frame 120 output from the second imaging processing unit 66 (see FIG. 4) is input to the learning model 80, but both the video display frame 120 output from the second imaging processing unit 66 (see FIG. 4) and the video display frame 120 output from the first imaging processing unit 64 (see FIG. 4) may be input to the learning model 80. Also, instead of the learning model input frame 124 generated from the video display frame 120 output from the second imaging processing unit 66, the learning model input frame 124 generated from the video display frame 120 output from the first imaging processing unit 64 may be input to the learning model 82.

[0130] In a modified example of the fourth embodiment, the learning model 80 to which the learning model input frame 124 is input is different from the learning model 82 to which the subject detection frame 122 is input. Therefore, the learning model 80 can output a subject detection result corresponding to the learning model input frame 124, and the learning model 82 can output a subject detection result corresponding to the subject detection frame 122. Furthermore, since it is possible to obtain a subject detection result that combines the subject detection result output from the learning model 80 and the subject detection result output from the learning model 82, the accuracy of the subject detection result can be improved compared to when a common learning model is used.

[0131] Fifth Embodiment Next, a fifth embodiment will be described.

[0132] 13 shows an example of the hardware configuration of a surveillance camera 10 according to the fifth embodiment. The fifth embodiment is modified from the first embodiment as follows. Specifically, the surveillance camera 10 includes a vibration sensor 130 and an electronic shake correction unit 132. The vibration sensor 130 and the electronic shake correction unit 132 are connected to the bus 56.

[0133] The vibration sensor 130 is, for example, a device including a gyro sensor, and detects the direction and amplitude of vibration acting on the surveillance camera 10. Note that instead of the vibration sensor 130, the direction and amplitude of vibration may be detected based on a plurality of video display frames 120 obtained successively in time series from the imaging element 32. For example, a motion vector may be detected by comparing a plurality of video display frames 120 obtained successively in time series from the imaging element 32, and the direction and amplitude of vibration may be detected based on the detected motion vector.

[0134] The electronic shake correction unit 132 is a device for executing electronic shake correction processing, which is image processing for correcting shake on the video display frame 120. The electronic shake correction unit 132 is, for example, a device including an ASIC. Note that the electronic shake correction unit 132 may also be, for example, a device including an FPGA or a PLD. The electronic shake correction unit 132 may also be, for example, a device including a combination of an ASIC, an FPGA, and a PLD. The electronic shake correction unit 132 may also be realized by a computer 40 including a CPU 50, a storage 54, and a memory 52.

[0135] 14 shows an example of a mode in which the image display frame 120 and the object detection frame 122 according to the fifth embodiment are selectively output. The vibration sensor 130 outputs a vibration detection result that is the result of detecting the direction and amplitude of vibration acting on the surveillance camera 10. The imaging determination unit 62 determines, based on the vibration detection result, whether the amplitude of the vibration detected by the vibration sensor 130 is greater than a threshold value. The threshold value is set to a value corresponding to the upper limit of the blur allowed for the image display frame 120.

[0136] When the imaging determination unit 62 determines that the amplitude of the vibration detected by the vibration sensor 130 is equal to or less than the threshold, the second imaging control unit 66A specifies the center row of the light receiving surface 32A and executes control over the imaging element 32 to read out imaging data for video display from the photosensitive pixels 34 of the specified row. A trimming process is executed to cut out the center image 110A from the read-out image 110 indicated by the imaging data for video display, and a video display frame 120 is generated from the read-out image 110.

[0137] Furthermore, when the amplitude of vibration detected by the vibration sensor 130 is equal to or less than a threshold value, the second imaging control unit 66A performs image processing for identifying the subject 2 on the video display frame 120 generated in the previous routine (i.e., the routine from step ST10 to step ST26), thereby identifying the subject 2 appearing in the video display frame 120. The second imaging control unit 66A also specifies a row of the light receiving surface 32A including an area where light from the subject 2 is imaged, and controls the image sensor 32 to read out imaging data for subject detection from the photosensitive pixels 34 of the specified row. A trimming process is performed to cut out an image 112A including the subject 2 from the readout image 112 indicated by the imaging data for subject detection, thereby generating a subject detection frame 122 from the readout image 112.

[0138] 15 shows an example of a manner in which the video display frame 120 according to the fifth embodiment is output. When the amplitude of the vibration detected by the vibration sensor 130 is greater than a threshold, the output of the subject detection frame 122 is stopped, and electronic image stabilization is performed on the video display frame 120.

[0139] Specifically, when the imaging determination unit 62 determines that the amplitude of vibration detected by the vibration sensor 130 is greater than the threshold, the second imaging control unit 66A sets the entire light receiving surface 32A as the readout range and controls the imaging element 32 to read out imaging data for video display from the photosensitive pixels 34 in the specified readout range. Then, the electronic shake correction unit 132 performs electronic shake correction processing, which is image processing for correcting shake on the imaging data for video display, to generate the video display frame 120. The electronic shake correction processing compares captured images represented by multiple pieces of imaging data to detect shake and generates the video display frame 120 with the shake corrected.

[0140] In the fifth embodiment, when the amplitude of vibration detected by the vibration sensor 130 is greater than a threshold, output of the subject detection frame 122 is stopped. This makes it possible to avoid outputting a subject detection result from the learning model 82 based on the subject detection frame 122 obtained when the amplitude of vibration detected by the vibration sensor 130 is greater than the threshold. This makes it possible to avoid a decrease in detection accuracy compared to a subject detection result obtained when the amplitude of vibration detected by the vibration sensor 130 is equal to or less than the threshold.

[0141] Furthermore, if the amplitude of the vibration detected by the vibration sensor 130 is greater than the threshold, electronic image stabilization is performed on the image display frame 120. Therefore, it is possible to obtain an image display frame 120 with corrected shake. This allows the surveillance camera 10 to continue monitoring the target even if the amplitude of the vibration detected by the vibration sensor 130 is greater than the threshold.

[0142] The fifth embodiment may be combined with any one of the second to fourth embodiments.

[0143] Sixth Embodiment Next, a sixth embodiment will be described.

[0144] 16 shows an example of the operation of the exposure setting unit 140 according to the sixth embodiment. The sixth embodiment is modified from the first embodiment as follows: The CPU 50 operates as the exposure setting unit 140.

[0145] The exposure setting unit 140 changes the exposure method of the surveillance camera 10 between the video display frame 120 and the subject detection frame 122. For example, as an exposure method corresponding to the video display frame 120, the exposure setting unit 140 derives an average exposure amount of the plurality of photosensitive pixels 34 based on the imaging data for video display, and sets the exposure based on the derived average exposure amount. On the other hand, as an exposure method corresponding to the subject detection frame 122, the exposure setting unit 140 derives an exposure amount of the photosensitive pixel 34 corresponding to the subject 2 out of the plurality of photosensitive pixels 34 based on the imaging data for subject detection, and sets the exposure based on the derived exposure amount. The exposure of the surveillance camera 10 is determined by an aperture value and a shutter speed.

[0146] In the sixth embodiment, for the video display frame 120, the average exposure of the plurality of photosensitive pixels 34 is derived based on the imaging data for video display, and the exposure is set based on the derived average exposure. Therefore, compared to the case where the exposure is set based on the exposure of the photosensitive pixels 34 corresponding to the subject 2, for example, it is possible to set the exposure appropriate for the video display frame 120.

[0147] On the other hand, for the subject detection frame 122, the exposure amount of the photosensitive pixel 34 that corresponds to the subject 2 among the multiple photosensitive pixels 34 is derived based on the imaging data for subject detection, and the exposure is set based on the derived exposure amount. Therefore, compared to when the exposure is set based on, for example, the average exposure amount of the multiple photosensitive pixels 34, it is possible to set an exposure that is more suitable for the subject detection frame 122. This makes it possible to improve the detection accuracy when the subject 2 is detected using the learning model 82 based on the subject detection frame 122.

[0148] Note that the sixth embodiment may be combined with any one of the second to fifth embodiments.

[0149] Seventh Embodiment Next, a seventh embodiment will be described.

[0150] 17 shows an example of the operation of the second imaging control unit 66A according to the seventh embodiment. The seventh embodiment is modified from the first embodiment as follows: The second imaging control unit 66A sets a factor related to the detection accuracy of the subject 2 in the subject detection frame 122 based on at least one of the subject detection result, the processing performance of the CPU 50, and the communication performance of the surveillance camera 10.

[0151] The processing performance of the CPU 50 is determined based on, for example, the specifications of the CPU 50 and / or the capacity of the memory 52. ​​The communication performance of the surveillance camera 10 is determined based on, for example, the communication performance of the bus 56 and / or the communication performance of the communication I / F 44. The communication performance of the surveillance camera 10 may be determined based on the performance of the network device 22 and / or the management device 12. The performance of the network device 22 and / or the management device 12 may be determined based on, for example, the processing performance of the CPU installed in the network device 22 and / or the CPU installed in the management device 12, or may be determined based on the communication performance of the network device 22 and / or the management device 12. The communication performance of the surveillance camera 10 is an example of the "communication performance of an imaging processing device" according to the technology of the present disclosure.

[0152] For example, the second imaging control unit 66A sets a read range 102 for reading the subject detection frame 122 from the imaging element 32 as an example of a factor based on at least one of the subject detection result, the processing performance of the CPU 50, and the communication performance of the surveillance camera 10.

[0153] Specifically, when the accuracy of the subject detection result is lower than the default accuracy, the second imaging control unit 66A widens the readout range 102. This widens the area in which the subject 2 is detected compared to when the readout range 102 is not widened, thereby improving the accuracy of the subject detection result. The default accuracy can be determined arbitrarily, for example, depending on the type of subject 2, etc.

[0154] Furthermore, when the processing load of the CPU 50 is higher than the default processing load, the second imaging control unit 66A narrows the readout range 102. This reduces the amount of data in the subject detection frame 122 compared to when the readout range 102 is not narrowed, thereby reducing the burden on the CPU 50. This ensures the accuracy of the subject detection results. The default processing load is set, for example, to the upper limit of the processing load when the accuracy of the subject detection results can be ensured.

[0155] Furthermore, when the communication load of the surveillance camera 10 is higher than the default communication load, the second imaging control unit 66A narrows the readout range 102. This reduces the amount of data in the subject detection frame 122 compared to when the readout range 102 is not narrowed, thereby reducing the processing load on the CPU 50 and ultimately the communication load on the surveillance camera 10. The default communication load is set, for example, to the upper limit of the communication load when the surveillance camera 10 can communicate normally.

[0156] In addition, the second imaging control unit 66A may set the angle of view of the surveillance camera 10 as an example of a factor based on at least one of the subject detection result, the processing performance of the CPU 50, and the communication performance of the surveillance camera 10.

[0157] Specifically, when the accuracy of the subject detection result is lower than the default accuracy, the second imaging control unit 66A widens the angle of view, which increases the area in which the subject 2 is detected compared to when the angle of view is not widened, thereby improving the accuracy of the subject detection result.

[0158] Furthermore, the second imaging control unit 66A narrows the angle of view when the processing load of the CPU 50 is higher than the default processing load. This reduces the load on the CPU 50 when detecting the subject 2 using the learning model 82 compared to when the angle of view is not narrowed, thereby ensuring the accuracy of the subject detection result.

[0159] Furthermore, the second imaging control unit 66A narrows the angle of view when the communication load of the surveillance camera 10 is higher than the default communication load. This reduces the burden on the CPU 50 when detecting the subject 2 using the learning model 82 compared to when the angle of view is not narrowed, and ultimately reduces the communication load on the surveillance camera 10.

[0160] Furthermore, the second imaging control unit 66A may set the data amount of the subject detection frame 122, as an example of a factor, based on at least one of the subject detection result, the processing performance of the CPU 50, and the communication performance of the surveillance camera 10. The data amount of the subject detection frame 122 may be adjusted, for example, by the resolution of the subject detection frame 122.

[0161] Specifically, when the accuracy of the subject detection result is lower than the default accuracy, the second imaging control unit 66A increases the resolution of the subject detection frame 122. This makes it possible to improve the accuracy of the subject detection result compared to when the resolution is not increased.

[0162] Furthermore, the second imaging control unit 66A lowers the resolution when the processing load of the CPU 50 is higher than the default processing load. This reduces the amount of data in the subject detection frame 122 compared to when the resolution is not lowered, thereby reducing the burden on the CPU 50.

[0163] Furthermore, the second imaging control unit 66A lowers the resolution when the communication load of the surveillance camera 10 is higher than the default communication load. This reduces the amount of data in the subject detection frame 122 compared to when the resolution is not lowered, thereby reducing the burden on the CPU 50 and, ultimately, the communication load on the surveillance camera 10.

[0164] In addition, the second imaging control unit 66A may set the frame rate of the subject detection frame 122 as an example of a factor based on at least one of the subject detection result, the processing performance of the CPU 50, and the communication performance of the surveillance camera 10.

[0165] Specifically, when the accuracy of the subject detection result is lower than the default accuracy, the second imaging control unit 66A increases the frame rate of the subject detection frames 122. This increases the frequency of subject 2 detection compared to when the frame rate is not increased, thereby improving the accuracy of subject 2 detection.

[0166] Furthermore, the second imaging control unit 66A lowers the frame rate when the processing load of the CPU 50 is higher than the default processing load. This reduces the frequency of detection of the subject 2 compared to when the frame rate is not lowered, thereby reducing the burden on the CPU 50.

[0167] Furthermore, the second imaging control unit 66A lowers the frame rate when the communication load of the surveillance camera 10 is higher than the default communication load. This reduces the frequency of detection of the subject 2 compared to when the frame rate is not lowered, thereby reducing the burden on the CPU 50 and, ultimately, the communication load of the surveillance camera 10.

[0168] The seventh embodiment may be combined with any one of the second to sixth embodiments.

[0169] Eighth Embodiment Next, an eighth embodiment will be described.

[0170] FIG. 18 shows an example of the operation of the imaging processing unit 60 according to the eighth embodiment. The eighth embodiment is modified from the first embodiment as follows. That is, in the first embodiment, either a first output state in which the image display frame 120 is output or a second output state in which the image display frame 120 and the subject detection frame 122 are selectively output is set based on the exposure time of the image sensor 32. In contrast, in the eighth embodiment, either the first output state or the second output state is set based on at least one of the video format, the processing performance of the CPU 50, and the communication performance of the surveillance camera 10. The video format may be, for example, the frame rate. The processing performance of the CPU 50 and the communication performance of the surveillance camera 10 are as described in the seventh embodiment.

[0171] For example, when the frame rate of the subject detection frames 122 is set to be less than a default frame rate, a first output state may be set in which the video display frames 120 are output, and when the frame rate of the subject detection frames 122 is set to be greater than the default frame rate, a second output state may be set in which the video display frames 120 and the subject detection frames 122 are selectively output. The default frame rate is set to, for example, a frame rate that corresponds to the upper limit of the data amount of the subject detection frames 122 when the accuracy of the subject detection results can be ensured.

[0172] In this way, for example, when the frame rate of the subject detection frames 122 is less than the default frame rate, the first output state is set to output the video display frames 120, so that it is possible to avoid detecting the subject 2 using the learning model 82 even if the data amount of the subject detection frames 122 exceeds the upper limit. On the other hand, when the frame rate of the subject detection frames 122 is equal to or greater than the default frame rate, it is possible to keep the data amount of the subject detection frames 122 below the upper limit. This reduces the burden on the CPU 50 compared to when the data amount of the subject detection frames 122 exceeds the upper limit, for example, and ensures the accuracy of the subject detection results.

[0173] Furthermore, when the processing load of the CPU 50 is higher than a default processing load, a first output state may be set in which the video display frame 120 is output, and when the processing load of the CPU 50 is equal to or lower than the default processing load, a second output state may be set in which the video display frame 120 and the subject detection frame 122 are selectively output. The default processing load is set, for example, to an upper limit value of the processing load when the accuracy of the subject detection result can be ensured.

[0174] In this way, for example, when the processing load of the CPU 50 is higher than the default processing load, the first output state for outputting the video display frame 120 is set, so that it is possible to avoid detecting the subject 2 using the learning model 82 even when the processing load of the CPU 50 exceeds the upper limit. On the other hand, when the processing load of the CPU 50 is equal to or lower than the default processing load, it is possible to reduce the burden on the CPU 50 compared to when the processing load of the CPU 50 is higher than the default processing load, so that it is possible to ensure the accuracy of the subject detection result.

[0175] Furthermore, when the communication load of the surveillance camera 10 is higher than a default communication load, a first output state in which the video display frame 120 is output may be set, and when the communication load of the surveillance camera 10 is equal to or lower than the default communication load, a second output state in which the video display frame 120 and the subject detection frame 122 are selectively output may be set. The default communication load is set to, for example, an upper limit value of the communication load in a state in which communication is established between the surveillance camera 10 and the management device 12 when the video display frame 120 and the subject detection result are selectively output from the surveillance camera 10 to the management device 12.

[0176] In this way, for example, when the communication load of the surveillance camera 10 is higher than a default communication load, a first output state in which the video display frame 120 is output is set, so that it is possible to avoid the surveillance camera 10 selectively outputting the video display frame 120 and the subject detection results to the management device 12 even when the communication load of the surveillance camera 10 exceeds an upper limit. On the other hand, when the communication load of the surveillance camera 10 is equal to or lower than the default communication load, a second output state in which the video display frame 120 and the subject detection frame 122 are selectively output is set, so that the video display frame 120 and the subject detection results generated in the surveillance camera 10 can be selectively output from the surveillance camera 10 to the management device 12.

[0177] The eighth embodiment may be combined with any one of the second to seventh embodiments.

[0178] Ninth Embodiment Next, a ninth embodiment will be described.

[0179] FIG. 19 shows an example of a monitoring system S according to the ninth embodiment. The ninth embodiment is modified from the first embodiment as follows. That is, in the first embodiment, the monitoring camera 10 is communicatively connected to the management device 12 via the network device 22, and a learning model 82 is installed in the monitoring camera 10. In contrast, in the ninth embodiment, the monitoring camera 10 is connected to the learning model device 200, and the learning model 82 is installed in the learning model device 200. Furthermore, the learning model device 200 is communicatively connected to the management device 12.

[0180] The learning model device 200 includes a computer having a CPU, memory, and storage, similar to the computer 40 (see FIG. 2) installed in the surveillance camera 10 according to the first embodiment. The processing function using the learning model 82 in the learning model device 200 is similar to the processing function using the learning model 82 in the surveillance camera 10 according to the first embodiment (for example, the subject detection processing unit 70 shown in FIG. 6).

[0181] The subject detection frame 122 output from the surveillance camera 10 is input to the learning model 82 of the learning model device 200, and the learning model 82 of the learning model device 200 generates a subject detection result corresponding to the input subject detection frame 122 and outputs the generated subject detection result to the management device 12. Even with this configuration, it is possible to obtain the same effects as in the first embodiment.

[0182] The ninth embodiment may be combined with any one of the second to eighth embodiments.

[0183] Furthermore, for example, when the seventh or eighth embodiment is combined with the ninth embodiment, the communication performance of the surveillance camera 10 may be determined based on the performance of the learning model device 200. The performance of the learning model device 200 may be determined based on the processing performance of the CPU mounted on the learning model device 200, or may be determined based on the communication performance of the learning model device 200, for example.

[0184] In the seventh embodiment described above, a factor (e.g., the readout range 102, etc.) relating to the detection accuracy of the subject 2 in the subject detection frame 122 may be set in the surveillance camera 10 based on the performance of the learning model device 200. Also, in the eighth embodiment described above, either a first output state in which the video display frame 120 is output or a second output state in which the video display frame 120 and the subject detection frame 122 are selectively output may be set in the surveillance camera 10 based on the performance of the learning model device 200.

[0185] Furthermore, in each of the above embodiments, a CPU 50 is exemplified for the surveillance camera 10, but instead of or together with the CPU 50, at least one other CPU, at least one GPU, and / or at least one TPU may be used.

[0186] In addition, in the above-described embodiments, the surveillance camera 10 has been described with an example in which the program 58 is stored in the storage 54, but the technology of the present disclosure is not limited to this. For example, the program 58 may be stored in a portable, non-transitory, computer-readable storage medium (hereinafter simply referred to as a "non-transitory storage medium") such as an SSD or USB memory. The program 58 stored in the non-transitory storage medium may be installed in the computer 40 of the surveillance camera 10.

[0187] In addition, the program 58 may be stored in a storage device such as another computer or server device connected to the surveillance camera 10 via a network, and the program 58 may be downloaded in response to a request from the surveillance camera 10 and installed on the computer 40 of the surveillance camera 10.

[0188] Furthermore, it is not necessary to store the entire program 58 in a storage device such as another computer or server device connected to the surveillance camera 10, or in the storage 54; only a part of the program 58 may be stored therein.

[0189] Furthermore, although the surveillance camera 10 has a built-in computer 40 , the technology of the present disclosure is not limited to this, and for example, the computer 40 may be provided outside the surveillance camera 10 .

[0190] In addition, in each of the above embodiments, the computer 40 including the CPU 50, memory 52, and storage 54 is exemplified for the surveillance camera 10, but the technology of the present disclosure is not limited to this, and a device including an ASIC, FPGA, and / or PLD may be applied instead of the computer 40. Furthermore, a combination of a hardware configuration and a software configuration may be used instead of the computer 40.

[0191] Furthermore, the hardware resources that execute the various processes described in the above embodiments can be various processors, as listed below. Examples of processors include a CPU, which is a general-purpose processor that functions as a hardware resource that executes various processes by executing software, i.e., a program. Examples of processors include dedicated electronic circuits, such as FPGAs, PLDs, and ASICs, which are processors with a circuit configuration specifically designed to execute specific processes. Each processor has built-in or connected memory, and each processor executes various processes by using the memory.

[0192] The hardware resources that execute various processes may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resources that execute various processes may be a single processor.

[0193] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes various processes. Second, there is a system that uses a processor that realizes the functions of the entire system, including multiple hardware resources that execute various processes, on a single IC chip, as typified by SoC. In this way, various processes are realized using one or more of the above-mentioned various processors as hardware resources.

[0194] Furthermore, the hardware structure of these various processors can be, more specifically, electronic circuits that combine circuit elements such as semiconductor devices. The various processes described above are merely examples. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the order of processes may be rearranged, without departing from the spirit of the invention.

[0195] The technology of the present disclosure extends to all program products. Program products include all manner of products for providing programs. For example, program products include programs provided over a network such as the Internet, and non-transitory computer-readable recording media such as CD-ROMs, DVDs, and USB memory sticks that store programs.

[0196] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[0197] In this specification, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."

[0198] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

Claims

1. An imaging processing device comprising a processor, the processor outputs at least a first frame for displaying an image based on first imaging data read out from an imaging element included in the imaging device, and outputs a second frame for detecting an object based on second imaging data read out from the imaging element, wherein factors related to the detection of the object in the second frame are different from the factors in the first frame.

2. The imaging processing device according to claim 1, wherein the factors include a frame rate.

3. The imaging processing device according to claim 2, wherein the second frame rate, which is the frame rate of the second frames, is higher than the first frame rate, which is the frame rate of the first frames.

4. The imaging processing device according to any one of claims 1 to 3, wherein the factors include at least one of the readout range, angle of view, data amount, and bit depth of the imaging element.

5. The imaging processing device according to claim 4, wherein the factors include the readout range, and the second readout range corresponding to the second frame is narrower than the first readout range corresponding to the first frame.

6. An imaging processing device as described in claim 4 or claim 5, wherein the factors include the angle of view, and the second angle of view corresponding to the second frame is narrower than the first angle of view corresponding to the first frame.

7. An imaging processing device as described in any one of claims 4 to 6, wherein the factors include the readout range, and the second readout range, which is the readout range corresponding to the second frame, is determined based on at least one of the subject detection result, the processing performance of the processor, and the communication performance of the imaging processing device.

8. An imaging processing device as described in any one of claims 4 to 7, wherein the factors include the angle of view, and the second angle of view, which is the angle of view corresponding to the second frame, is determined based on at least one of the subject detection result, the processing performance of the processor, and the communication performance of the imaging processing device.

9. The imaging processing device according to claim 7 or claim 8, wherein the imaging processing device is communicably connected to a management device via a network device, and the communication performance of the imaging processing device is determined by the performance of at least one of the network device and the management device.

10. An imaging processing device as described in claim 7 or claim 8, wherein the imaging processing device is communicatively connected to a learning model device having a learning model, and the communication performance of the imaging processing device is determined by the performance of the learning model device.

11. An imaging processing device according to any one of claims 1 to 10, wherein the processor sets either a first output state in which the first frame is output or a second output state in which the first frame and the second frame are selectively output based on imaging conditions.

12. The imaging processing device described in claim 11, wherein the imaging conditions include a first condition that the exposure time of the imaging element is longer than a time calculated as the reciprocal of a first frame rate, which is the frame rate of the first frame, and the processor sets the first output state when the first condition is met.

13. An imaging processing device according to claim 11 or 12, wherein the imaging conditions include conditions relating to at least one of the exposure time of the imaging element, the format of the image, the processing performance of the processor, and the communication performance of the imaging processing device.

14. The imaging processing device described in claim 13, wherein the imaging conditions include communication performance of the imaging processing device, the imaging processing device is communicatively connected to a management device via a network device, the processor outputs the first frame to the management device via the network device, outputs the second frame to a learning model of the imaging processing device, and outputs the subject detection result of the learning model to the management device via the network device, and the communication performance of the imaging processing device is determined by the performance of at least one of the network device and the management device.

15. The imaging processing device described in claim 13, wherein the imaging conditions include communication performance of the imaging processing device, the imaging processing device is communicatively connected to a learning model device having a learning model, the processor outputs the first frame to the learning model device, and outputs the second frame to the learning model device, and the communication performance of the imaging processing device is determined by the performance of the learning model device.

16. The imaging processing device according to any one of claims 1 to 15, wherein the processor changes an exposure method of the imaging device between the first frame and the second frame.

17. An imaging processing device as described in claim 16, wherein the exposure method of the imaging device corresponding to the first frame is a method of setting exposure based on the average exposure of a plurality of photosensitive pixels that constitute the light receiving surface of the imaging element, and the exposure method of the imaging device corresponding to the second frame is a method of setting exposure based on the exposure of a photosensitive pixel among the plurality of photosensitive pixels that corresponds to the subject.

18. An imaging processing device according to any one of claims 1 to 17, wherein the second frame is a frame in which the readout range of the imaging element is narrowed to an area including the subject, compared to the first frame.

19. An imaging processing device according to any one of claims 1 to 18, wherein the pixel resolution of the imaging element when outputting the second frame corresponds to the pixel resolution of the imaging element when outputting the first frame.

20. An imaging processing device according to any one of claims 1 to 19, wherein the resolution of the second frame corresponds to the resolution of the first frame.

21. An imaging processing device according to any one of claims 1 to 20, wherein the resolution of the second frame is lower than the resolution of the first frame.

22. The imaging processing device according to any one of claims 1 to 21, wherein the second frame is a frame input to a second learning model.

23. The imaging processing device described in claim 22, wherein the processor generates the second frame to be input to the second learning model by performing a cropping process on the read image indicated by the second imaging data to cut out an area including the subject.

24. The imaging processing device according to claim 22 or 23, wherein the first frame is a frame input to a first learning model.

25. The imaging processing device according to claim 24, wherein the first learning model and the second learning model are a common learning model.

26. The imaging processing device according to claim 24, wherein the first learning model is a learning model different from the second learning model.

27. The imaging processing device described in claim 24, wherein the processor generates a third frame to be input to a third learning model by performing a cropping process on the first frame, which is a read image indicated by the first imaging data, to cut out an area including the subject.

28. The imaging processing device according to claim 27, wherein the first learning model and the third learning model are a common learning model.

29. The imaging processing device according to claim 27, wherein the first learning model is a learning model different from the third learning model.

30. An imaging processing device as described in any one of claims 1 to 29, wherein the processor stops outputting the second frame and performs electronic shake correction processing on the first frame when the amplitude of vibration acting on the imaging device is greater than a threshold value.

31. An imaging processing device according to any one of claims 1 to 30, wherein the imaging element is an imaging element capable of imaging at a frame rate higher than a first frame rate, which is the frame rate of the first frame.

32. An imaging processing method comprising: outputting at least a first frame for displaying an image based on first imaging data read out from an imaging element included in an imaging device; and outputting a second frame for detecting an object based on second imaging data read out from the imaging element, wherein factors related to the detection of the object in the second frame are different from the factors in the first frame.

33. A program for causing a computer to execute processing, the processing including: outputting at least a first frame for displaying an image based on first imaging data read from an imaging element included in an imaging device; and outputting a second frame for detecting an object based on second imaging data read from the imaging element, wherein factors related to the detection of the object in the second frame are different from the factors in the first frame.

Citation Information

Patent Citations

  • Photographing device and method

    JP2007235640A

  • Image sensing apparatus, image sensing system, and image sensing method

    JP2007295525A

  • Imaging apparatus, image playback device, and program for them

    JP2009141538A

  • Imaging device and control method thereof

    JP2010068439A

  • Control device, imaging device, control method, and program

    JP2020191606A