Information processing system, information processing method, and information processing device.
The information processing system effectively utilizes recognition processing results outside the CCU by generating and outputting recognition metadata, enhancing focus adjustment and imaging accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2026-03-25
AI Technical Summary
Existing systems fail to effectively utilize recognition processing results outside the Camera Control Unit (CCU), limiting their application and efficiency.
An information processing system and method that includes an imaging device and a control device, where recognition processing is performed, recognition metadata is generated, and output to the imaging device or subsequent devices, enabling effective utilization of recognition processing results.
Enhances the utilization of recognition processing results, improving focus adjustment, visibility, and accuracy in imaging systems by integrating recognition metadata into the imaging process.
Smart Images

Figure 0007835217000001 
Figure 0007835217000002 
Figure 0007835217000003
Abstract
Description
Technical Field
[0001] The present technology relates to an information processing system, an information processing method, and an information processing apparatus, and particularly relates to an information processing system, an information processing method, and an information processing apparatus suitable for use when an information processing apparatus that controls an imaging apparatus performs recognition processing on an imaging image.
Background Art
[0002] Conventionally, a system including a CCU (Camera Control Unit) that performs recognition processing on an image captured by a camera has been proposed (see, for example, Patent Documents 1 and 2).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the inventions described in Patent Documents 1 and 2, the results of recognition processing are used within the CCU, but using the results of recognition processing outside the CCU has not been considered.
[0005] The present technology has been made in view of such a situation, and enables effective use of the results of recognition processing on an imaging image by an information processing apparatus that controls an imaging apparatus.
Means for Solving the Problems
[0006] The first aspect of this technology is an information processing system comprising an imaging device for capturing images and an information processing device for controlling the imaging device, wherein the information processing device comprises a recognition unit for performing recognition processing on the captured images, a recognition metadata generation unit for generating recognition metadata including data based on the results of the recognition processing, and an output unit for outputting the recognition metadata to the imaging device.
[0007] In the first aspect of this technology, recognition processing is performed on the captured image, recognition metadata including data based on the results of the recognition processing is generated, and the recognition metadata is output to the imaging device.
[0008] The second aspect of this technology is an information processing method in which an information processing device that controls an imaging device that captures an image performs recognition processing on the captured image, generates recognition metadata including data based on the results of the recognition processing, and outputs the recognition metadata to the imaging device.
[0009] In the second aspect of this technology, recognition processing is performed on the captured image, recognition metadata including data based on the results of the recognition processing is generated, and the recognition metadata is output to the imaging device.
[0010] The third aspect of this technology is an information processing system comprising an imaging device for capturing images and an information processing device for controlling the imaging device, wherein the information processing device comprises a recognition unit for performing recognition processing on the captured images, a recognition metadata generation unit for generating recognition metadata including data based on the results of the recognition processing, and an output unit for outputting the recognition metadata to a subsequent device.
[0011] In the third aspect of this technology, recognition processing is performed on the captured image, recognition metadata including data based on the results of the recognition processing is generated, and the recognition metadata is output to a subsequent device.
[0012] The fourth aspect of this technology is the information processing method, in which an information processing device that controls an imaging device that captures an image performs recognition processing on the captured image, generates recognition metadata including data based on the results of the recognition processing, and outputs the recognition metadata to a subsequent device.
[0013] In the fourth aspect of this technology, recognition processing is performed on the captured image, recognition metadata including data based on the results of the recognition processing is generated, and the recognition metadata is output to a subsequent device.
[0014] The fifth aspect of this technology is an information processing device comprising: a recognition unit that performs recognition processing on an image captured by an imaging device; a recognition metadata generation unit that generates recognition metadata including data based on the results of the recognition processing; and an output unit that outputs the recognition metadata.
[0015] In the fifth aspect of this technology, recognition processing is performed on the captured image captured by the imaging device, recognition metadata including data based on the results of the recognition processing is generated, and the recognition metadata is output. [Brief explanation of the drawing]
[0016] [Figure 1] This is a block diagram showing one embodiment of an information processing system to which this technology is applied. [Figure 2] This is a block diagram showing an example of the functional configuration of a camera's CPU. [Figure 3] This block diagram shows an example of the CPU functional configuration of a CCU. [Figure 4] This is a block diagram showing an example of the functional configuration of the information processing unit of a CCU. [Figure 5] This is a flowchart explaining the process of displaying the focus indicator. [Figure 6] This figure shows an example of a focus indicator display. [Figure 7] This is a flowchart to explain the peaking highlighting process. [Figure 8] This figure shows an example of peaking highlighting. [Figure 9] It is a flowchart for explaining the image mask process. [Figure 10] It is a diagram showing an example of an image frame. [Figure 11] It is a diagram showing an example of region recognition. [Figure 12] It is a diagram for explaining the mask process. [Figure 13] It is a diagram showing an example of the luminance waveform and vector scope display of the image frame before mask processing. [Figure 14] It is a diagram showing an example of the luminance waveform and vector scope display of the image frame after mask processing by the first method. [Figure 15] It is a diagram showing an example of the luminance waveform and vector scope display of the image frame after mask processing by the second method. [Figure 16] It is a diagram showing an example of the luminance waveform and vector scope display of the image frame after mask processing by the third method. [Figure 17] It is a flowchart for explaining the reference direction correction process. [Figure 18] It is a diagram showing an example of a feature point map. [Figure 19] It is a diagram for explaining the method of detecting the imaging direction based on feature points. [Figure 20] It is a diagram for explaining the method of detecting the imaging direction based on feature points. [Figure 21] It is a flowchart for explaining the subject recognition and embedding process. [Figure 22] It is a diagram showing an example of an image with the information indicating the result of subject recognition superimposed. [Figure 23] It is a diagram showing an example of the configuration of a computer.
Modes for Carrying Out the Invention
[0017] Hereinafter, the modes for carrying out this technology will be described. The description will be made in the following order. 1. Embodiment 2. Variation 3. Others
[0018] <<1. Embodiment>> Embodiments of this technology will be described with reference to Figures 1 to 22.
[0019] <Example Configuration of Information Processing System 1> Figure 1 is a block diagram showing one embodiment of an information processing system 1 to which this technology is applied.
[0020] The information processing system 1 comprises a camera 11, a tripod 12, a pan / tilt head 13, a camera cable 14, a CCU (Camera Control Unit) 15 for controlling the camera 11, an operation panel 16, and a monitor 17. The camera 11 is mounted on a pan / tilt head 13 attached to the tripod 12, allowing it to rotate in the pan, tilt, and roll directions. The camera 11 and the CCU 15 are connected by a camera cable 14.
[0021] The camera 11 comprises a main unit 21, a lens 22, and a viewfinder 23. The lens 22 and the viewfinder 23 are mounted on the main unit 21. The main unit 21 comprises a signal processing unit 31, a motion sensor 32, and a CPU 33.
[0022] Lens 22 supplies lens information to the CPU 33. This lens information includes, for example, the focal length, focusing distance, and control values and specifications of the lens, such as the iris value.
[0023] The signal processing unit 31 shares video signal processing duties with the signal processing unit 51 of the CCU 15. For example, the signal processing unit 31 performs predetermined signal processing on the video signal obtained when an image sensor (not shown) captures a subject through the lens 22, and generates a video frame consisting of the captured image taken by the image sensor. The signal processing unit 31 supplies the video frame to the viewfinder 23 and also outputs it to the signal processing unit 51 of the CCU 15 via the camera cable 14.
[0024] The motion sensor 32 includes, for example, an angular velocity sensor and an acceleration sensor, and detects the angular velocity and acceleration of the camera 11. The motion sensor 32 supplies data indicating the detection results of the angular velocity and acceleration of the camera 11 to the CPU 33.
[0025] The CPU 33 controls the processing of each part of the camera 11. For example, the CPU 33 changes the control values of the camera 11 or displays information about the control values on the viewfinder 23 based on the control signals input from the CCU 15.
[0026] The CPU 33 detects the orientation (pan angle, tilt angle, roll angle), i.e., the imaging direction of the camera 11, based on the detection result of the camera 11's angular velocity. For example, the CPU 33 sets a reference direction in advance and detects the imaging direction (orientation) of the camera 11 by accumulating (integrating) the amount of change in the orientation of the camera 11 based on the reference direction. In addition, the CPU 33 may also use the detection result of the camera 11's acceleration to detect the imaging direction of the camera 11.
[0027] Here, the reference direction of camera 11 is the direction in which the pan angle, tilt angle, and roll angle of camera 11 are set to 0 degrees. The CPU 33 corrects the reference direction it holds internally based on the correction data included in the recognition metadata input from the CCU 15.
[0028] The CPU 33 acquires control information for the main unit 21, such as shutter speed and color balance. The CPU 33 generates camera metadata, including imaging direction information, control information, and lens information for the camera 11. The CPU 33 outputs the camera metadata to the CPU 52 of the CCU 15 via the camera cable 14.
[0029] The CPU 33 controls the display of the live view shown in the viewfinder 23. The CPU 33 also controls the display of information superimposed on the live view based on recognition metadata and control signals input from the CCU 15.
[0030] The viewfinder 23, under the control of the CPU 33, displays a through-image or various information superimposed on the through-image based on the video frame supplied from the signal processing unit 31.
[0031] The CCU15 comprises a signal processing unit 51, a CPU 52, an information processing unit 53, an output unit 54, and a mask processing unit 55.
[0032] The signal processing unit 51 performs predetermined video signal processing on the video frames generated by the signal processing unit 31 of the camera 11. The signal processing unit 51 then supplies the video frames after video signal processing to the information processing unit 53, the output unit 54, and the mask processing unit 55.
[0033] The CPU 52 controls the processing of each part of the CCU 15. The CPU 52 also communicates with the operation panel 16 and acquires control signals input from the operation panel 16. The CPU 52 outputs the acquired control signals to the camera 11 via the camera cable 14 or supplies them to the mask processing unit 55 as needed.
[0034] The CPU 52 supplies camera metadata input from the camera 11 to the information processing unit 53 and the mask processing unit 55. The CPU 52 outputs the recognition metadata supplied from the information processing unit 53 to the camera 11 via the camera cable 14, to the operation panel 16, or to the mask processing unit 55. The CPU 52 generates supplementary metadata based on the camera metadata and recognition metadata and supplies it to the output unit 54.
[0035] The information processing unit 53 performs various recognition processes on video frames using computer vision, AI (Artificial Intelligence), machine learning, etc. For example, the information processing unit 53 performs subject recognition and region recognition within video frames. More specifically, for example, the information processing unit 53 performs feature point extraction, matching, detection of the imaging direction of the camera 11 based on tracking (pose detection), skeletal detection, face detection, face recognition, pupil detection, object detection, action recognition, semantic segmentation, etc. using machine learning. The information processing unit 53 also detects deviations in the imaging direction detected by the camera 11 based on the video frames. The information processing unit 53 generates recognition metadata that includes data based on the results of the recognition processing. The information processing unit 53 supplies the recognition metadata to the CPU 52.
[0036] The output unit 54 arranges (adds) video frames and associated metadata to an output signal in a predetermined format (for example, an SDI (Serial Digital Interface) signal) and outputs it to the monitor 17 in the subsequent stage.
[0037] The mask processing unit 55 performs mask processing on the video frame based on the control signals and recognition metadata supplied from the CPU 52. Mask processing is the process of masking areas of the video frame other than the area of a predetermined type of subject (hereinafter referred to as the mask area), as will be described later. The output unit 54 places (adds) the masked video frame to an output signal of a predetermined format (for example, an SDI signal) and outputs it to the monitor 17.
[0038] The control panel 16 is composed of, for example, an MSU (Master Setup Unit) and an RCP (Remote Control Panel). The control panel 16 is used by a user such as a VE (Video Engineer), and generates control signals based on user operations, which are then output to the CPU 52.
[0039] The monitor 17 is used, for example, by a user such as in a VE to view the video captured by the camera 11. For example, the monitor 17 displays the video based on the output signal from the output unit 54. The monitor 17 displays the masked video based on the output signal from the mask processing unit 55. The monitor 17 displays the luminance waveform and vector scope of the masked video frame and the like.
[0040] In the following, in the signal and data transmission processing between the camera 11 and the CCU 15, the description of the camera cable 14 will be appropriately omitted. For example, when the camera 11 outputs a video frame to the CCU 15 via the camera cable 14, the description of the camera cable 14 may be omitted and it may simply be described that the camera 11 outputs the video frame to the CCU 15.
[0041] <Functional configuration example of the CPU 33> FIG. 2 shows a configuration example of the functions realized by the CPU 33 of the camera 11. For example, by executing a predetermined control program, functions including the control unit 71, the imaging direction detection unit 72, the camera metadata generation unit 73, and the display control unit 74 are realized.
[0042] The control unit 71 controls the processing of each part of the camera 11.
[0043] The imaging direction detection unit 72 detects the imaging direction of the camera 11 based on the detection result of the angular velocity of the camera 11. Note that the imaging direction detection unit 72 may use the detection result of the acceleration of the camera 11 to detect the imaging direction of the camera 11. Also, the imaging direction detection unit 72 corrects the reference direction of the camera 11 based on the recognition metadata input from the CCU 15.
[0044] The camera metadata generation unit 73 generates camera metadata including the imaging direction information, control information, and lens information of the camera 11. The camera metadata generation unit 73 outputs the camera metadata to the CPU 52 of the CCU 15.
[0045] The display control unit 74 controls the display of the through-image by the viewfinder 23. Further, the display control unit 74 controls the display of information superimposed on the through-image by the viewfinder 23 based on the recognition metadata input from the CCU 15.
[0046] <Functional configuration example of the CPU 52> FIG. 3 shows a configuration example of the functions realized by the CPU 52 of the CCU 15. For example, when the CPU 52 executes a predetermined control program, functions including the control unit 101 and the metadata output unit 102 are realized.
[0047] The control unit 101 controls the processing of each part of the CCU 15.
[0048] The metadata output unit 102 supplies the camera metadata input from the camera 11 to the information processing unit 53 and the mask processing unit 55. The metadata output unit 102 outputs the recognition metadata supplied from the information processing unit 53 to the camera 11, outputs it to the operation panel 16, or supplies it to the mask processing unit 55. The metadata output unit 102 generates additional metadata based on the camera metadata and the recognition metadata supplied from the information processing unit 53, and supplies it to the output unit 54.
[0049] <Configuration example of the information processing unit 53> FIG. 4 shows a configuration example of the information processing unit 53 of the CCU 15. The information processing unit 53 includes a recognition unit 131 and a recognition metadata generation unit 132.
[0050] The recognition unit 131 performs various recognition processes on the video frame.
[0051] The recognition metadata generation unit 132 generates recognition metadata including data based on the recognition process by the recognition unit 131. The recognition metadata generation unit 132 supplies the recognition metadata to the CPU 52.
[0052] <Processing of the information processing system 1> Next, the processing of the information processing system 1 will be described.
[0053] <Focus indicator display processing> First, we will explain the focus indicator display process performed by the information processing system 1, referring to the flowchart in Figure 5.
[0054] This process starts, for example, when the user inputs an instruction to start displaying the focus index value using the control panel 16, and ends when the user inputs an instruction to stop displaying the focus index value.
[0055] In step S1, the information processing system 1 performs imaging processing.
[0056] Specifically, the image sensor (not shown) captures an image of the subject and supplies the obtained video signal to the signal processing unit 31. The signal processing unit 31 performs predetermined video signal processing on the video signal supplied from the image sensor and generates video frames. The signal processing unit 31 supplies the video frames to the viewfinder 23 and also outputs them to the signal processing unit 51 of the CCU 15. Under the control of the display control unit 74, the viewfinder 23 displays a through-image based on the video frames.
[0057] Lens 22 supplies lens information related to lens 22 to CPU 33. Motion sensor 32 detects the angular velocity and acceleration of camera 11 and supplies data indicating the detection results to CPU 33.
[0058] The imaging direction detection unit 72 detects the imaging direction of the camera 11 based on the detection results of the angular velocity and acceleration of the camera 11. For example, the imaging direction detection unit 72 detects the imaging direction (attitude) of the camera 11 by accumulating (integrating) the amount of change in the orientation (angle) of the camera 11 based on the angular velocity detected by the motion sensor 32, using a pre-set reference direction as a reference.
[0059] The camera metadata generation unit 73 generates camera metadata including imaging direction information, lens information, and control information for the camera 11. The camera metadata generation unit 73 outputs camera metadata corresponding to the video frame to the CPU 52 of the CCU 15 in synchronization with the output of the video frame by the signal processing unit 31. This associates the video frame with the camera metadata including imaging direction information, control information, and lens information of the camera 11 around the time the video frame was captured.
[0060] The signal processing unit 51 of the CCU15 performs predetermined video signal processing on the video frames acquired from the camera 11, and supplies the processed video frames to the information processing unit 53, the output unit 54, and the mask processing unit 55.
[0061] The metadata output unit 102 of the CCU15 supplies camera metadata acquired from the camera 11 to the information processing unit 53 and the mask processing unit 55.
[0062] In step S2, the recognition unit 131 of the CCU 15 performs subject recognition. For example, the recognition unit 131 uses methods such as skeletal detection, face detection, pupil detection, and object detection to recognize subjects of the type that are eligible for display of the focus index value within the video frame. If there are multiple subjects of the type eligible for display of the focus index value within the video frame, the recognition unit 131 recognizes each subject individually.
[0063] In step S3, the recognition unit 131 of the CCU 15 calculates the focus index value. Specifically, the recognition unit 131 calculates the focus index value for the region containing each recognized subject.
[0064] The method for calculating the focus index is not particularly limited. For example, frequency analysis using Fourier transform, cepstrum analysis, and DfD (Depth from Defocus) techniques can be used to calculate the focus index.
[0065] In step S4, the CCU 15 generates recognition metadata. Specifically, the recognition metadata generation unit 132 generates recognition metadata including the position and focus index value of each subject recognized by the recognition unit 131, and supplies it to the CPU 52. The metadata output unit 102 outputs the recognition metadata to the CPU 33 of the camera 11.
[0066] In step S5, the viewfinder 23 of the camera 11 displays a focus indicator under the control of the display control unit 74.
[0067] Figure 6 schematically shows an example of the focus indicator display. Figure 6A shows an example of the through-image displayed in the viewfinder 23 before the focus indicator is displayed. Figure 6B shows an example of the through-image displayed in the viewfinder 23 after the focus indicator is displayed.
[0068] In this example, the through-image shows people 201a through 201c. Person 201a is closest to camera 11, and person 201c is furthest from camera 11. Camera 11 is in focus on person 201a.
[0069] In this example, the right eye of person 201a through person 201c is set as the target for displaying the focus index value. As shown in Figure 6B, indicator 202a, a circular image indicating the position of person 201a's right eye, is displayed around person 201a's right eye. Indicator 202b, a circular image indicating the position of person 201b's right eye, is displayed around person 201b's right eye. Indicator 202c, a circular image indicating the position of person 201c's right eye, is displayed around person 201c's right eye.
[0070] Additionally, bars 203a to 203c, which indicate the focus index value for the right eye of person 201a to person 201c, are displayed below the through-image. Bar 203a indicates the focus index value for the right eye of person 201a. Bar 203b indicates the focus index value for the right eye of person 201b. Bar 203c indicates the focus index value for the right eye of person 201c. The length of bars 203a to 203c indicates the value of the focus index.
[0071] Bars 203a through 203c are each set to a different display mode (e.g., different color). On the other hand, indicator 202a and bar 203a are set to the same display mode (e.g., the same color). Indicator 202b and bar 203b are set to the same display mode (e.g., the same color). Indicator 202c and bar 203c are set to the same display mode (e.g., the same color). This makes it easy for the user (e.g., a photographer) to grasp the correspondence between each subject and the focus index value.
[0072] For example, if the area where the focus indicator value is displayed is fixed to the center of the viewfinder 23, the focus indicator value becomes unusable if the subject you want to focus on moves outside that area.
[0073] In contrast, this technology automatically tracks a subject of a desired type and displays the focus index value for that subject. Furthermore, if there are multiple subjects for which the focus index value is to be displayed, the focus index values are displayed individually for each subject. Additionally, the subject and the focus index value are associated in a different display manner for each subject.
[0074] This allows users (e.g., photographers) to easily adjust the focus on their desired subject.
[0075] After that, the process returns to step S1, and the processes from step S1 onward are executed.
[0076] <Peaking highlighting process> Next, the peaking highlighting process performed by the information processing system 1 will be explained with reference to the flowchart in Figure 7.
[0077] This process starts, for example, when the user inputs an instruction to start peaking highlighting using the control panel 16, and ends when the user inputs an instruction to stop peaking highlighting.
[0078] Here, peaking highlighting is a function that highlights high-frequency components within an image frame, and is also called detail highlighting. Peaking highlighting is used, for example, to assist with manual focus operation.
[0079] In step S21, the imaging process is performed in the same manner as in step S1 in Figure 5.
[0080] In step S22, the recognition unit 131 of the CCU 15 performs subject recognition. For example, the recognition unit 131 recognizes the region and type of each subject within the video frame using object detection, semantic segmentation, etc.
[0081] In step S23, the CCU 15 generates recognition metadata. Specifically, the recognition metadata generation unit 132 generates recognition metadata including the position and type of each subject recognized by the recognition unit 131 and supplies it to the CPU 52. The metadata output unit 102 outputs the recognition metadata to the CPU 33 of the camera 11.
[0082] In step S24, the viewfinder 23 of the camera 11 performs peaking highlighting based on recognition metadata, with the control of the display control unit 74.
[0083] Figure 8 schematically illustrates an example of peaking highlighting for a golf tee shot scene. Figure 8A shows an example of the through-shot image displayed in the viewfinder 23 before peaking highlighting. Figure 8B shows an example of the through-shot image displayed in the viewfinder 23 after peaking highlighting, with the highlighted area indicated by diagonal lines.
[0084] For example, if peaking highlighting is applied to the entire through-image, high-frequency components in the background will also be highlighted, which may reduce visibility.
[0085] On the other hand, this technology allows for limiting the subjects to which peaking highlighting is applied. For example, as shown in Figure 8B, the subjects to which peaking highlighting is applied can be limited to the area where a person is visible, indicated by the diagonal lines. In this case, in the actual through-image, high-frequency components such as edges in the area indicated by the diagonal lines are highlighted using auxiliary lines or the like.
[0086] This improves the visibility of peaking highlighting, making it easier for users (e.g., photographers) to manually focus on a desired subject.
[0087] After that, the process returns to step S21, and the processes from step S21 onward are executed.
[0088] <Image masking> Next, referring to the flowchart in Figure 9, we will explain the video masking process performed by the information processing system 1.
[0089] This process starts, for example, when the user inputs an instruction to start the video masking process using the control panel 16, and ends when the user inputs an instruction to stop the video masking process.
[0090] In step S41, the imaging process is performed in the same manner as in step S1 in Figure 5.
[0091] In step S42, the recognition unit 131 of the CCU 15 performs region recognition. For example, the recognition unit 131 divides the video frame into multiple regions according to the type of subject by performing semantic segmentation on the video frame.
[0092] In step S43, the CCU 15 generates recognition metadata. Specifically, the recognition metadata generation unit 132 generates recognition metadata including the region within the video frame recognized by the recognition unit 131 and its type, and supplies it to the CPU 52. The metadata output unit 102 supplies the recognition metadata to the mask processing unit 55.
[0093] In step S44, the mask processing unit 55 performs mask processing.
[0094] For example, the user uses the control panel 16 to select the type of subject they want to keep without masking. The control unit 101 supplies data indicating the type of subject selected by the user to the masking processing unit 55.
[0095] The mask processing unit 55 performs masking on areas of the video frame other than the type of subject selected by the user (mask area).
[0096] In the following, the area of the subject selected by the user will be referred to as the recognition target area.
[0097] Here, we will explain specific examples of mask processing with reference to Figures 10 to 12.
[0098] Figure 10 schematically shows an example of a video frame captured during a golf tee shot.
[0099] Figure 11 shows an example of the results of performing region recognition on the video frame in Figure 10. In this example, the video frame is divided into regions 251 to 255, and each region is shown in a different pattern. Region 251 is the region where a person is visible (hereinafter referred to as the person region). Region 252 is the region where the ground is visible. Region 253 is the region where the forest is visible. Region 254 is the region where the sky is visible. Region 255 is the region where the tee marker is visible.
[0100] Figure 12 schematically shows an example of setting a recognition target area and a mask area for the video frame in Figure 10. In this example, the area indicated by the diagonal lines (corresponding to areas 252 to 255 in Figure 11) is set as the mask area. The area not marked with diagonal lines (corresponding to area 251 in Figure 11) is set as the recognition target area.
[0101] Furthermore, it is possible to set the recognition target area to include regions of multiple types of subjects.
[0102] Here, we will explain three types of masking methods.
[0103] In the first masking method, the pixel signals in the masked area are replaced with black signals; that is, the masked area is filled with black. On the other hand, the pixel signals in the recognition target area are not modified.
[0104] In the second masking method, the chroma component of the pixel signal in the masked area is reduced. For example, the U and V components of the chroma component of the pixel signal in the masked area are set to 0. On the other hand, the luminance component of the pixel signal in the masked area is not changed. Also, the pixel signal of the recognition target area is not changed.
[0105] In the third masking method, similar to the second masking method, the chroma component of the pixel signal in the masked area is reduced. For example, the U and V components of the chroma component of the pixel signal in the masked area are set to 0. In addition, the luminance component of the masked area is reduced. For example, the luminance component of the masked area is transformed by equation (1) below, and the contrast of the luminance component of the masked area is compressed. On the other hand, the pixel signal of the recognition target area is not particularly changed.
[0106] Yout=Yin×gain+offset ···(1)
[0107] Yin represents the luminance component before masking, while Yout represents the luminance component after masking. Gain indicates a predetermined gain, set to a value less than 1.0. Offset indicates the offset value.
[0108] The mask processing unit 55 places (adds) the masked video frame to the output signal in a predetermined format and outputs the output signal to the monitor 17.
[0109] In step S45, the monitor 17 displays the masked video and waveform. Specifically, the monitor 17 displays the video based on the masked video frame, based on the output signal obtained from the mask processing unit 55. The monitor 17 also displays the luminance waveform of the masked video frame for brightness adjustment. Furthermore, the monitor 17 displays the vectorscope of the masked video frame for color adjustment.
[0110] Now, with reference to Figures 13 to 16, we compare the masking processes of the first to third methods described above.
[0111] Figures 13 to 16 show examples of the brightness waveform and vectorscope display for the video frame in Figure 10.
[0112] Figure 13A shows an example of the brightness waveform of a video frame before masking, and Figure 13B shows an example of the vectorscope of a video frame before masking.
[0113] In the luminance waveform, the horizontal axis indicates the horizontal position of the video frame, and the vertical axis indicates the amplitude of the luminance. In the vectorscope, the circumferential direction indicates hue, and the radial direction indicates saturation. This is also true for Figures 14 to 16.
[0114] The luminance waveform before masking displays the luminance waveform for the entire video frame. Similarly, the vectorscope before masking displays the hue and saturation waveforms for the entire video frame.
[0115] In luminance waveforms and vectorscopes before masking, luminance and chroma components in areas outside the recognition target region become noise. Furthermore, for example, when matching color balance between multiple cameras, the luminance waveform and vectorscope waveform for the same subject will differ significantly depending on whether it is lit from the front or back. Therefore, adjusting the brightness and hue of the recognition target region while viewing the luminance waveform and vectorscope before masking is difficult, especially for inexperienced users.
[0116] Figure 14A shows an example of displaying the luminance waveform of a video frame after masking by the first method, and Figure 14B shows an example of displaying the vectorscope of a video frame after masking by the first method.
[0117] In the luminance waveform after masking using the first method, the luminance waveform of only the human region, which is the recognition target area, is displayed. Therefore, for example, it becomes easy to adjust the brightness targeting only the human.
[0118] In the vectorscope after masking using the first method, the hue and saturation waveforms are displayed only for the human area, which is the recognition target region. Therefore, for example, it becomes easy to adjust the color tone targeting only the human.
[0119] However, in the video frame after masking using the first method, the masked area is blacked out, reducing the visibility of the video frame. In other words, the user will not be able to see the video outside the recognition target area.
[0120] Figure 15A shows an example of the display of the luminance waveform of the video frame after masking by the second method, and Figure 15B shows an example of the display of the vectorscope of the video frame after masking by the second method.
[0121] The luminance waveform after masking using the second method is similar to the luminance waveform before masking in Figure 13A. Therefore, it becomes difficult to adjust the brightness of, for example, only a person.
[0122] The waveform of the vectorscope after masking using the second method is similar to the waveform of the vectorscope after masking using the first method, shown in Figure 14B. Therefore, for example, it becomes easier to adjust the color tone of only the person.
[0123] Furthermore, since the luminance components of the masked area remain intact in the video frame processed using the second method, visibility is improved compared to the video frame processed using the first method.
[0124] Figure 16A shows an example of the display of the luminance waveform of the video frame after masking by the third method, and Figure 16B shows an example of the display of the vectorscope of the video frame after masking by the third method.
[0125] In the luminance waveform after masking using the third method, the contrast of the masked area is compressed, making the waveform of the human area, which is the recognition target area, stand out. Therefore, for example, it becomes easier to adjust the brightness targeting only the human.
[0126] The waveform of the vectorscope after masking using the third method is similar to the waveform of the vectorscope after masking using the first method, shown in Figure 14B. Therefore, it becomes easier to adjust the color tone, for example, only for people.
[0127] Furthermore, the video frame after masking using the third method retains the luminance components of the masked area, albeit with compressed contrast, resulting in improved visibility compared to the video frame after masking using the first method.
[0128] Thus, the masking process of the third method makes it possible to easily adjust the brightness and hue of the recognition target area while ensuring the visibility of the masked area of the video frame.
[0129] Furthermore, the brightness of the video frame may be displayed using other methods, such as a parade display or a histogram. In this case as well, the brightness of the recognition target area can be easily adjusted by using the masking process of the first or third method.
[0130] Subsequently, the process returns to step S41, and the processes from step S41 onward are executed.
[0131] In this way, it is possible to easily adjust the brightness and color of the desired subject while maintaining the visibility of the video frame. Furthermore, since the monitor 17 does not require any special processing, it is possible to use an existing monitor as the monitor 17.
[0132] For example, in step S43, the metadata output unit 102 may also output the recognition metadata to the camera 11. The results of the area recognition in the camera 11 may then be used for selecting the detection area for functions such as auto iris and white balance adjustment.
[0133] <Reference direction correction process> Next, the reference direction correction process performed by the information processing system 1 will be explained with reference to the flowchart in Figure 17.
[0134] This process starts, for example, when camera 11 begins capturing images, and ends when camera 11 finishes capturing images.
[0135] In step S61, the information processing system 1 starts the imaging process. That is, the same imaging process as in step S1 of Figure 5 described above is started.
[0136] In step S62, the CCU 15 starts the process of embedding video frames and metadata into the output signal and outputting it. Specifically, the metadata output unit 102 starts the process of organizing the camera metadata acquired from the camera 11 to generate supplemental metadata and supplying it to the output unit 54. The output unit 54 starts the process of arranging (adding) the video frames and supplemental metadata to the output signal in a predetermined format and outputting it to the monitor 17.
[0137] In step S63, the recognition unit 131 of the CCU 15 starts updating the feature point map. Specifically, the recognition unit 131 detects feature points in the video frame and, based on the detection results, starts updating the feature point map that shows the distribution of feature points around the camera 11.
[0138] Figure 18 shows an example of a feature point map. The "X" marks in the figure indicate the location of the feature points.
[0139] For example, the recognition unit 131 generates and updates a feature point map that shows the positions and feature vectors of feature points in the scene around the camera 11 by stitching together the detection results of feature points in video frames captured around the camera 11. In this feature point map, the position of a feature point is represented, for example, by the distance in the direction relative to the reference direction of the camera 11 and in the depth direction.
[0140] In step S64, the recognition unit 131 of the CCU 15 detects a deviation in the imaging direction. Specifically, the recognition unit 131 detects the imaging direction of the camera 11 by matching feature points detected from the video frame with a feature point map.
[0141] For example, Figure 19 shows an example of a video frame when camera 11 is facing the reference direction. Figure 20 shows an example of a video frame when camera 11 is facing -7 degrees (7 degrees counterclockwise) from the reference direction in the panning direction.
[0142] For example, the recognition unit 131 detects the imaging direction of the camera 11 by matching the feature points of the feature point map in Figure 18 with the feature points of the video frame in Figure 19 or Figure 20.
[0143] The recognition unit 131 then detects the difference between the imaging direction detected based on the video frame and the imaging direction detected by the camera 11 using the motion sensor 32 as a deviation in the imaging direction. In other words, the detected deviation corresponds to the cumulative error that arises when the imaging direction detection unit 72 of the camera 11 calculates the cumulative angular velocity detected by the motion sensor 32.
[0144] In step S65, the CCU 15 generates recognition metadata. Specifically, the recognition metadata generation unit 132 generates recognition metadata that includes data based on the detected shift in the imaging direction. For example, the recognition metadata generation unit 132 calculates a correction value for the reference direction based on the detected shift in the imaging direction and generates recognition metadata that includes the correction value for the reference direction. The recognition metadata generation unit 132 supplies the generated recognition metadata to the CPU 52.
[0145] The metadata output unit 102 outputs the recognition metadata to the camera 11.
[0146] In step S66, the imaging direction detection unit 72 of the camera 11 corrects the reference direction based on the correction value of the reference direction included in the recognition metadata. At this time, the imaging direction detection unit 72 corrects the reference direction continuously in multiple steps, for example, using α blending (IIR (Infinite Impulse Response) processing). As a result, the reference direction changes gradually and smoothly.
[0147] After that, the process returns to step S64, and the processes from step S64 onward are executed.
[0148] In this way, the reference direction of the camera 11 is appropriately corrected, thereby improving the accuracy of the camera 11's detection of the imaging direction.
[0149] Furthermore, the camera 11 corrects the reference direction based on the results of the video frame recognition processing by the CCU 15. This reduces the delay in correcting the camera 11's imaging direction deviation compared to when the CCU 15 directly corrects the imaging direction using recognition processing that requires processing time.
[0150] <Subject Recognition and Metadata Embedding Processing> Next, referring to the flowchart in Figure 21, we will explain the subject recognition and metadata embedding process performed by the information processing system 1.
[0151] This process starts, for example, when the user inputs an instruction to start the subject recognition and embedding process using the control panel 16, and ends when the user inputs an instruction to stop the subject recognition and embedding process.
[0152] In step S81, imaging is performed in the same manner as in step S1 in Figure 5.
[0153] In step S82, the recognition unit 131 of the CCU 15 performs subject recognition. For example, the recognition unit 131 recognizes the position, type, and action of each object within the video frame by performing object recognition and action recognition on the video frame.
[0154] In step S83, the CCU 15 generates recognition metadata. Specifically, the recognition metadata generation unit 132 generates recognition metadata including the position, type, and action of each object recognized by the recognition unit 131, and supplies it to the CPU 52.
[0155] The metadata output unit 102 generates supplementary metadata based on camera metadata acquired from camera 11 and recognition metadata acquired from recognition metadata generation unit 132. The supplementary metadata includes, for example, imaging direction information, lens information, and control information of camera 11, as well as recognition results of the position, type, and action of each object in the video frame. The metadata output unit 102 supplies the supplementary metadata to the output unit 54.
[0156] In step S84, the output unit 54 embeds video frames and metadata into the output signal and outputs it. Specifically, the output unit 54 arranges (adds) video frames and associated metadata to an output signal in a predetermined format and outputs it to the monitor 17.
[0157] Monitor 17 displays the image shown in Figure 22, for example, based on the output signal. The image in Figure 22 is the image in Figure 10 with information indicating the location, type, and action recognition results of objects included in the accompanying metadata superimposed.
[0158] In this example, the positions of the person, golf club, ball, and mountain are displayed in the video. It also indicates that the person is performing a tee shot.
[0159] After that, the process returns to step S81, and the processes from step S81 onward are executed.
[0160] In this way, metadata including the results of subject recognition for video frames can be embedded into the output signal in real time without human intervention. This makes it possible to quickly present the subject recognition results, for example, as shown in Figure 22.
[0161] Furthermore, it becomes possible to omit the process of recognizing and analyzing video frames and adding metadata in subsequent devices.
[0162] <Summary of the effects of this technology> As described above, the CCU 15 performs recognition processing on video frames while the camera 11 is capturing images, and the camera 11 and monitor 17 outside the CCU 15 can use the results of the recognition processing in real time.
[0163] For example, the viewfinder 23 of camera 11 can display information based on the recognition processing results superimposed on the through-image in real time. The monitor 17 can display information based on the recognition processing results superimposed on the video based on the video frame in real time, or display the masked video in real time. This improves the operability for users such as cameramen and video engineers.
[0164] Furthermore, the camera 11 can correct the detection result of the imaging direction in real time based on the correction value of the reference direction obtained through recognition processing. This improves the accuracy of imaging direction detection.
[0165] <<2. Variant>> The following describes some modifications of the embodiments of the present technology described above.
[0166] <Variations regarding the division of labor> For example, it is possible to change the division of processing between camera 11 and CCU 15. For instance, camera 11 may be configured to perform some or all of the processing of the information processing unit 53 of CCU 15.
[0167] However, if, for example, the camera 11 were to perform all of the processing of the information processing unit 53, the processing load on the camera 11 would increase, leading to a larger camera housing and increased power consumption and heat generation. The increased size of the camera housing and increased heat generation are undesirable because they would hinder cable management and other aspects of the camera 11. Furthermore, if, for example, the information processing system 1 performs signal processing using the Baseband Processing Unit for 4K / 8K shooting or high frame rate shooting, it would be difficult for the camera 11 to develop the entire video frame and perform recognition processing, as the information processing unit 53 does.
[0168] Furthermore, it is also possible to configure a device downstream of the CCU15, such as a PC (Personal Computer) or a server, to execute the processing of the information processing unit 53. In this case, the CCU15 outputs video frames and camera metadata to the downstream device, which then performs the recognition processing described above, generates recognition metadata, and outputs it to the CCU15. Therefore, processing delays and ensuring sufficient transmission bandwidth between the CCU15 and the downstream device become challenges. In particular, delays in processing related to the operation of the camera 11, such as focus operations, become a problem.
[0169] Therefore, considering the addition of metadata to the output signal, the output of recognition metadata to the camera 11, and the display of the recognition processing results in the viewfinder 23 and monitor 17, it is optimal to place the information processing unit 53 in the CCU 15 as described above.
[0170] <Other variations> For example, the output unit 54 may output the output signal in association with the output signal, rather than embedding the associated metadata in the output signal.
[0171] For example, the recognition metadata generation unit 132 of the CCU 15 may generate recognition metadata that includes a detected value of the deviation in the imaging direction instead of a correction value in the reference direction, as data used for correcting the reference direction. The imaging direction detection unit 72 of the camera 11 may then correct the reference direction based on the detected value of the deviation in the imaging direction.
[0172] <<3.B>> <Example of computer configuration> The series of processes described above can be executed by hardware or by software. When the series of processes are executed by software, the programs that make up that software are installed on a computer. Here, a computer includes computers built into dedicated hardware, as well as general-purpose personal computers that can perform various functions by installing various programs.
[0173] Figure 23 is a block diagram showing an example of the hardware configuration of a computer that executes the series of processes described above by a program.
[0174] In computer 1000, the CPU (Central Processing Unit) 1001, ROM (Read Only Memory) 1002, and RAM (Random Access Memory) 1003 are interconnected by a bus 1004.
[0175] An input / output interface 1005 is further connected to the bus 1004. An input / output interface 1005 is connected to an input unit 1006, an output unit 1007, a recording unit 1008, a communication unit 1009, and a drive 1010.
[0176] The input section 1006 consists of input switches, buttons, a microphone, an image sensor, etc. The output section 1007 consists of a display, a speaker, etc. The recording section 1008 consists of a hard disk or non-volatile memory, etc. The communication section 1009 consists of a network interface, etc. The drive 1010 drives removable media 1011 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.
[0177] In the computer 1000 configured as described above, the CPU 1001 loads, for example, a program stored in the recording unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004, and executes it, thereby performing the series of processes described above.
[0178] The program executed by computer 1000 (CPU 1001) can be provided by recording it on removable media 1011, such as a packaged media. The program can also be provided via wired or wireless transmission media, such as a local area network, the internet, or digital satellite broadcasting.
[0179] In computer 1000, programs can be installed in the recording unit 1008 via the input / output interface 1005 by inserting the removable media 1011 into the drive 1010. Alternatively, programs can be received by the communication unit 1009 via a wired or wireless transmission medium and installed in the recording unit 1008. Furthermore, programs can be pre-installed in the ROM 1002 or the recording unit 1008.
[0180] The programs executed by the computer may be programs that are processed chronologically in the order described herein, or they may be programs that are processed in parallel or at necessary times, such as when a call is made.
[0181] Furthermore, in this specification, a system means a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure or not. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device in which multiple modules are housed in one enclosure, are both considered systems.
[0182] Furthermore, the embodiments of this technology are not limited to those described above, and various modifications are possible without departing from the spirit of this technology.
[0183] For example, this technology can be configured as cloud computing, where a single function is shared and processed collaboratively by multiple devices via a network.
[0184] Furthermore, each step described in the flowchart above can be performed by a single device, or it can be divided and performed by multiple devices.
[0185] Furthermore, if a single step includes multiple processes, those processes can be executed by a single device or shared among multiple devices.
[0186] <Examples of configuration combinations> This technology can also be configured as follows:
[0187] (1) An imaging device that captures images, An information processing device that controls the imaging device and Equipped with, The aforementioned information processing device is A recognition unit that performs recognition processing on the captured image, A recognition metadata generation unit that generates recognition metadata including data based on the results of the recognition process, An output unit that outputs the aforementioned recognition metadata to the imaging device. An information processing system equipped with the following features. (2) The recognition unit performs at least one of the following: subject recognition and region recognition within the captured image. The recognition metadata includes at least one of the subject recognition results and the region recognition results. The information processing system described in (1) above. (3) The imaging device is A display unit that shows a through-image, Based on the recognition metadata, a display control unit controls the display of the through image. The information processing system described in (2) above, comprising: (4) The recognition unit calculates a focus index value for a predetermined type of subject recognized by the subject recognition, The recognition metadata further includes the focus index value. The display control unit superimposes an image indicating the position of the subject and the focus index value for the subject onto the through-image. The information processing system described in (3) above. (5) The display control unit superimposes an image indicating the position of the subject and the focus index value onto the through-image in a different display manner for each subject. The information processing system described in (4) above. (6) The display control unit performs peaking highlighting of the through-image, limited to the area of a predetermined type of subject, based on the recognition metadata. An information processing system as described in any of (3) to (5) above. (7) The imaging device is An imaging direction detection unit that detects the imaging direction of the imaging device with respect to a predetermined reference direction, A camera metadata generation unit generates camera metadata including the detected imaging direction and outputs it to the information processing device. Equipped with, The recognition unit detects the shift in the imaging direction included in the camera metadata based on the captured image, The recognition metadata includes data based on the detected deviation in the imaging direction. An information processing system as described in any of (1) to (6) above. (8) The recognition metadata generation unit generates the recognition metadata, which includes data used for correcting the reference direction, based on the detected deviation in the imaging direction. The imaging direction detection unit corrects the reference direction based on the recognition metadata. The information processing system described in (7) above. (9) An information processing device that controls the imaging device that captures images, Recognition processing is performed on the captured image, Recognition metadata is generated, which includes data based on the results of the recognition process. The recognition metadata is output to the imaging device. Information processing methods. (10) An imaging device that captures images, An information processing device that controls the imaging device and Equipped with, The aforementioned information processing device is A recognition unit that performs recognition processing on the captured image, A recognition metadata generation unit that generates recognition metadata including data based on the results of the recognition process, An output unit that outputs the aforementioned recognition metadata to a subsequent device. An information processing system equipped with the following features. (11) The recognition unit performs at least one of the following: subject recognition and region recognition within the captured image. The recognition metadata includes at least one of the subject recognition results and the region recognition results. The information processing system described in (10) above. (12) A mask processing unit performs masking on a mask region, which is the region of the captured image other than the region of a predetermined type of subject, and outputs the captured image after the masking process to the subsequent device. The information processing system described in (11) further comprises the above. (13) The mask processing unit reduces the chroma component of the mask region and compresses the contrast of the luminance component of the mask region. The information processing system described in (12) above. (14) The output unit adds at least a portion of the recognition metadata to the output signal, which includes the captured image, and outputs the output signal to the subsequent device. The information processing system described in any of (10) to (13) above. (15) The imaging device is A camera metadata generation unit generates camera metadata including the detection result of the imaging direction of the imaging device and outputs it to the information processing device. Prepare, The output unit further adds at least a portion of the camera metadata to the output signal. The information processing system described in (14) above. (16) The camera metadata further includes at least one of the control information of the imaging device and lens information relating to the lens of the imaging device. The information processing system described in (15) above. (17) An information processing device that controls the imaging device that captures images, Recognition processing is performed on the captured image, Recognition metadata is generated, which includes data based on the results of the recognition process. The aforementioned recognition metadata is output to a subsequent device. Information processing methods. (18) A recognition unit that performs recognition processing on an image captured by an imaging device, A recognition metadata generation unit that generates recognition metadata including data based on the results of the recognition process, The output unit that outputs the aforementioned recognition metadata An information processing device equipped with the following features. (19) The output unit outputs the recognition metadata to the imaging device. The information processing device described in (18) above. (20) The output unit outputs the recognition metadata to a subsequent device. The information processing apparatus described in (18) or (19) above.
[0188] Furthermore, the effects described herein are merely illustrative and not limiting; other effects may also occur. [Explanation of Symbols]
[0189] 1 Information Processing System, 11 Camera, 15 CCU, 16 Operation Panel, 17 Monitor, 21 Main Unit, 22 Lens, 23 Viewfinder, 31 Signal Processing Unit, 32 Motion Sensor, 33 CPU, 51 Signal Processing Unit, 52 CPU, 53 Information Processing Unit, 54 Output Unit, 55 Mask Processing Unit, 71 Control Unit, 72 Imaging Direction Detection Unit, 73 Camera Metadata Generation Unit, 74 Display Control Unit, 101 Control Unit, 102 Metadata Output Unit, 131 Recognition Unit, 132 Recognition Metadata Generation Unit
Claims
1. An imaging device that captures images, An information processing device that controls the imaging device and Equipped with, The imaging device is An imaging direction detection unit that detects the imaging direction of the imaging device with respect to a predetermined reference direction, A camera metadata generation unit generates camera metadata including the detected imaging direction and outputs it to the information processing device. Equipped with, The aforementioned information processing device is A recognition unit that performs recognition processing on the captured image and detects the discrepancy between the imaging direction of the imaging device based on the recognition processing and the imaging direction included in the camera metadata, A recognition metadata generation unit that generates recognition metadata including data based on the detected shift in imaging direction, An output unit that outputs the aforementioned recognition metadata to the imaging device. Equipped with Information processing system.
2. The recognition metadata generation unit generates the recognition metadata, which includes data used for correcting the reference direction, based on the detected deviation in the imaging direction. The imaging direction detection unit corrects the reference direction based on the recognition metadata. The information processing system according to claim 1.
3. The imaging device that captures the image is To detect the imaging direction of the imaging device with respect to a predetermined reference direction, To generate camera metadata including the detected imaging direction and output it to the information processing device. Includes, The aforementioned information processing device The process involves performing recognition processing on the captured image, detecting the discrepancy between the imaging direction of the imaging device based on the recognition processing and the imaging direction included in the camera metadata, To generate recognition metadata that includes data based on the detected shift in imaging direction, Outputting the aforementioned recognition metadata to the imaging device Information processing methods including
4. A recognition unit performs recognition processing on an image captured by an imaging device, and detects the difference between the imaging direction of the imaging device based on the recognition processing and the imaging direction of the imaging device, which is based on a predetermined reference direction and is included in the camera metadata output from the imaging device. A recognition metadata generation unit that generates recognition metadata including data based on the detected shift in imaging direction, An output unit that outputs the aforementioned recognition metadata to the imaging device. An information processing device equipped with the following features.
Citation Information
Patent Citations
Device and method for image recognition, and storage medium
JP2000113097A
Imaging apparatus
JP2015049294A
Image processor and control method thereof
JP2015156054A
Imaging apparatus and verification system
JP2015233261A
Display controller, control method and program of display controller and storage medium
JP2018067787A