Processing apparatus, photographing system, processing method, and program
The processing device detects and corrects moire and other abnormalities in virtual production by comparing display and captured videos, enhancing efficiency and accuracy without requiring a separate camera.
Patent Information
- Application Number
- JP2023222807
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-10
AI Technical Summary
Existing methods for detecting moire during virtual production require a separate camera for moire detection, which is inconvenient and inefficient.
A processing device that uses a display device to display a first video and captures a second video with an imaging device, analyzing the images to detect abnormalities such as moire by comparing the display area and the captured video, and performing control to notify the user of any detected issues.
Enables real-time detection and correction of moire and other abnormalities during shooting, improving the efficiency and accuracy of virtual production without the need for additional equipment.
Smart Images

Figure 2025104764000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a processing device, a photographing system, a processing method, and a program.
Background Art
[0002] In recent years, virtual production in which a background image is displayed on an LED display (LED wall) and a real subject is photographed at the same time has been rapidly spreading. If minute moire (interference fringes) occur on the LED display, it may not be noticed during shooting, and it may be necessary to re-shoot after noticing it in post-production. Therefore, it is desirable to be able to detect moire during shooting. Patent Document 1 discloses a method in which a signal on the spatial axis representing an image in which moire has occurred is converted into a signal on the frequency axis, and then the signal on the frequency axis from which the frequency component corresponding to moire has been removed is converted into a signal on the spatial axis.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the method of Patent Document 1, since moire is specified using a plurality of images taken at different shooting angles, it is necessary to prepare a camera different from the shooting camera for moire detection.
[0005] An object of the present invention is to provide a processing device capable of detecting an abnormality such as moire that occurs when shooting with a background image displayed on a display.
Means for Solving the Problems
[0006] A processing device according to one aspect of the present invention is a processing device that causes a display device to display a first video and acquires a second video captured by an imaging device that images an angle of view including the display area of the display device, the processing device comprising: acquisition means for acquiring a discrimination result obtained by discriminating an area in which a subject existing between the display area and the imaging device in the second video is imaged and the display area; detection means for detecting a change in the second video by comparing the first video and the video of the display area in the second video; and control means for performing first control for notifying information regarding the detection result by the detection means.
Advantages of the Invention
[0007] According to the present invention, it is possible to provide a processing device capable of detecting an abnormality such as moire that occurs when shooting with a background video displayed on a display.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Best Mode for Carrying Out the Invention
[0009] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In each figure, the same members are denoted by the same reference numerals, and duplicate descriptions are omitted. <First Embodiment> FIG. 1 is a block diagram of the imaging device 100. The imaging device 100 can input, output, and further record an image. Each unit connected to the internal bus 101 can exchange data with each other via the internal bus 101.
[0010] The lens unit 106 includes a lens group including a zoom lens and a focus lens, a diaphragm mechanism, and a drive motor. The optical image that has passed through the lens unit 106 is received by the imaging unit 107. The imaging unit 107 uses a CCD, a CMOS sensor, or the like, and replaces an optical signal with an electrical signal.
[0011] The CPU 102 controls each unit of the imaging device 100 according to a program stored in the ROM 103, using the RAM 104 as a work memory. The ROM 103 is a non-volatile recording element and stores programs for operating the CPU 102 and various adjustment parameters. The RAM 104 is a volatile memory using semiconductor elements, and generally, a low-speed and small-capacity one is used compared to the frame memory 111. The frame memory 111 is an element that can temporarily store an image signal and read it out when necessary. Since the image signal has a huge data volume, the frame memory 111 is required to be high-bandwidth and large-capacity. In recent years, DDR4-SDRAM (Dual Data Rate 4-Synchronous Dynamic RAM) or the like has been used as the frame memory 111. By using the frame memory 111, for example, it becomes possible to perform processes such as synthesizing images that are different in time or cutting out only a necessary area.
[0012] Based on the control of the CPU 102, the image processing unit 105 performs various image processes on the data from the imaging unit 107 or the image data stored in the frame memory 111 and the recording medium 112. The image processes performed by the image processing unit 105 include pixel interpolation of image data, encoding process, compression process, decoding process, enlargement / reduction process (resizing), noise reduction process, color conversion process, and the like. Further, the image processing unit 105 performs processes such as correction of performance variations of pixels in the imaging unit 107, correction of defective pixels, correction of white balance, correction of luminance, and correction of distortion and peripheral light quantity drop caused by lens characteristics. Furthermore, the image processing unit 105 performs a process of generating distance information regarding the distance from the imaging device 100 to the subject. Note that the image processing unit 105 may be composed of dedicated circuit blocks for performing specific image processes. Depending on the type of image process, the CPU 102 may perform image processes according to a program without using the image processing unit 105.
[0013] Based on the calculation result obtained by the image processing unit 105, the CPU 102 controls the lens unit 106, and can perform optical image enlargement, focal length adjustment, and adjustment of the aperture for adjusting the light quantity. Further, it may be configured to perform shake correction by moving a part of the lens group on a plane orthogonal to the optical axis.
[0014] The operation unit 113 is an interface with the external device that receives the user's operation. The operation unit 113 uses elements such as mechanical buttons and switches, and is composed of a power switch, a mode switch, and the like.
[0015] The display unit 114 is a display device that can be visually recognized by the user. By the display unit 114 displaying, for example, an image processed by the image processing unit 105, a setting menu, etc., it is possible for the user to check the operation status of the imaging device 100. In recent years, as the display unit 114, small-sized and low-power consumption devices such as LCD (Liquid Crystal Display) and organic EL (Electroluminescence) have been used. Also, the display unit 114 may be equipped with thin-film elements such as a resistive film type or a capacitive type called a touch panel, and may be used as an alternative to the operation unit 113. The CPU 102 generates a character string for notifying the user of the setting state, etc. of the imaging device 100, and a menu for setting the imaging device 100, superimposes them on the image processed by the image processing unit 105, and displays them on the display unit 114. In addition to character information, it is also possible to superimpose shooting assist displays such as a histogram, a vector scope, a waveform monitor, zebra, peaking, and false color.
[0016] The video terminal unit 109 is composed of a plurality of video terminals. The video terminals are, for example, SDI (Serial Digital Interface), HDMI (registered trademark) (High Definition Multimedia Interface), and DisplayPort (registered trademark), etc. By outputting a video signal via the video terminal unit 109, it becomes possible to display an image in real time on an external monitor (not shown). Also, the CPU 102 can process the image input via the video terminal unit 109 by the image processing unit 105 and display it on the display unit 114.
[0017] The network module 108 is an interface for inputting and outputting an image signal and an audio signal. The network module 108 communicates with external devices via the Internet or the like, and can also transmit and receive various data such as files, commands, video signals, and metadata. The transmission method by the network module 108 may be wireless or wired.
[0018] The recording medium 112 can record image data and various setting data, and a large-capacity storage element is used. For example, an HDD (Hard Disc Drive), an SSD (Solid State Drive), etc. are used as the recording medium 112 and are mounted on the recording medium I / F 110. When the user presses a predetermined button (hereinafter referred to as the recording button) of the operation unit 113, the CPU 102 starts recording the image data processed by the image processing unit 105 on the recording medium 112 via the recording medium I / F 110. When the user presses the recording button again, the recording stops. During recording, a signal indicating that the imaging device 100 is recording is output via the network module 108 and the video terminal unit 109. Thereby, it is also possible to notify an external device that recording is in progress or to record the image data output from the imaging device 100 in the external device in conjunction with the recording operation in the imaging device 100.
[0019] The object detection unit 115 detects an object using artificial intelligence typified by deep learning using, for example, a neural network. Taking object detection by deep learning as an example, the CPU 102 transmits a program for processing, a network structure, weight parameters, etc. stored in the ROM 103 to the object detection unit 115. The object detection unit 115 performs processing for detecting an object from the image signal based on various parameters obtained from the CPU 102 and develops the processing result in the RAM 104.
[0020] The attitude detection unit 116 detects the attitude state of the imaging device 100 using, for example, a gyro sensor or an acceleration sensor. Thereby, it becomes possible to detect that the camera is tilted or shaking.
[0021] FIG. 2 is a diagram showing a part of the light-receiving surface of the image sensor of the imaging unit 107. In the imaging unit 107, in order to enable imaging plane phase difference autofocus, a pixel unit that holds two photodiodes as light-receiving units as photoelectric conversion means for one microlens is arranged in an array. Thereby, each pixel unit can receive a light beam obtained by splitting the exit pupil of the lens unit 106.
[0022] Figure 2(A) is a schematic diagram of a part of the surface of an image sensor with a Bayer array example of red (R), blue (B), and green (Gb, Gr). Figure 2(B) shows an example of a pixel portion that holds two photodiodes for one microlens corresponding to the array of color filters in Figure 2(A).
[0023] The image sensor of this embodiment can output two signals for phase difference detection (hereinafter also referred to as A image signal and B image signal) from each pixel portion. Also, an imaging signal (A image signal + B image signal) obtained by adding the signals of the two photodiodes can be output. When outputting the added signal, a signal equivalent to the output of the image sensor with the Bayer array example schematically described in Figure 2(A) is output.
[0024] The imaging unit 107 can output a phase difference detection signal for each pixel portion, but can also output a value obtained by adding and averaging the phase difference detection signals of a plurality of adjacent pixel portions. By outputting the added and averaged value, the time for reading the signal from the imaging unit 107 can be shortened, and the bandwidth used in the internal bus 101 can be reduced. Using the output signal from such an imaging unit 107, the CPU 102 performs a correlation operation on the two image signals and calculates information such as the defocus amount, parallax information, and various reliabilities.
[0025] The defocus amount on the image plane is calculated based on the shift between the A image signal and the B image signal. The defocus amount has positive and negative values, and whether it is a front focus or a rear focus can be determined by whether the defocus amount is a positive value or a negative value. Also, the degree of defocus (the degree of out-of-focus shift) until focus can be known from the absolute value of the defocus amount, and if the defocus amount is 0, it is in focus. That is, the CPU 102 calculates information on whether it is a front focus or a rear focus based on the positive or negative of the defocus amount. Also, based on the absolute value of the defocus amount, defocus degree information, which is the degree of focus, is calculated. Information on whether it is a front focus or a rear focus is output when the defocus amount exceeds a predetermined value, and information indicating that it is in focus is output when the absolute value of the defocus amount is within a predetermined value.
[0026] The CPU 102 controls the lens unit 106 according to the defocus amount to perform focus adjustment. Further, the CPU 102 calculates the distance to the subject using the principle of triangulation from the parallax information and the lens information of the lens unit 106.
[0027] In FIG. 2, an example is shown in which pixel portions each holding two photodiodes for one microlens are arranged in an array, but the present invention is not limited to this. Pixel portions each holding three or more photodiodes for one microlens may be arranged in an array. Also, there may be a plurality of pixel portions having different light-receiving part opening positions with respect to the microlens. That is, any configuration other than the configuration shown in FIG. 2 may be used as long as it is a configuration capable of detecting a phase difference between two signals such as an A image signal and a B image signal.
[0028] Hereinafter, the distance information generation process performed by the image processing unit 105 will be described. FIG. 3 is a flowchart showing the distance information generation process.
[0029] In step S301, the image processing unit 105 calculates the B image signal for phase difference detection by obtaining the difference between the two signals of the imaging (A image signal + B image signal) output from the imaging unit 107 and the A image signal for phase difference detection. In the present embodiment, the description is made using a method in which the imaging (A image signal + B image signal) and the A image signal for phase difference detection are output, but the present invention is not limited to this. The imaging unit 107 may output the A image signal and the B image signal. In this case, the imaging (A image signal + B image signal) can be calculated by adding the A image signal and the B image signal. Also, in the case of having two sensors like a stereo camera, the image signals output from the respective sensor outputs may be used as the A image signal and the B image signal.
[0030] In step S302, the image processing unit 105 corrects the shading due to optical factors for each of the A image signal for phase difference detection and the B image signal for phase difference detection.
[0031] In step S303, the image processing unit 105 performs filter processing on each of the A image signal for phase difference detection and the B image signal for phase difference detection. The filter processing is performed using, for example, a high-pass filter configured by FIR (Finite Impulse Response). In the present embodiment, the description will be made using the A image signal for phase difference detection and the B image signal for phase difference detection that have passed through the high-pass filter, but the present invention is not limited to this. Signals passed through a band-pass filter or a low-pass filter with different filter coefficients may be generated respectively. Then, the correlation operation processing described later may be performed using the generated A image signal for phase difference detection and the B image signal for phase difference detection respectively.
[0032] In step S304, the image processing unit 105 divides the A image signal for phase difference detection and the B image signal for phase difference detection that have undergone filter processing in step S303 into minute blocks and performs a correlation operation. There is no limitation on the size or shape of the minute blocks, and the regions of adjacent blocks may overlap.
[0033] Here, the correlation operation between the A image and the B image, which are a pair of images, will be described. The signal sequences of the A image at the pixel position of interest are denoted as E(1) to E(m), and the signal sequences of the B image at the pixel position of interest are denoted as F(1) to F(m). While relatively shifting the signal sequences F(1) to F(m) of the B image with respect to the signal sequences E(1) to E(m) of the A image, the correlation amount C(k) at the shift amount k between the two signal sequences is calculated using the following formula (1).
[0034]
Equation
[0035] (1) In Equation (1), the Σ operation means an operation for calculating the sum with respect to n. In the Σ operation, the ranges that n and n + k take are limited to the range from 1 to m. The shift amount k is an integer value and is a relative pixel shift amount in units of the detection pitch of a pair of data. Hereinafter, k when the discrete correlation amount C(k) becomes minimum is denoted as kj. In an ideal state where there is no noise, the operation result of Equation (1) when the correlation of a pair of image signal sequences is high is shown in FIG. 4. As shown in FIG. 4, at the shift amount (k = kj = 0) where the correlation of a pair of image signal sequences is high, the correlation amount C(k) becomes minimum. By the in-point interpolation processing shown in Equations (2) to (4), x that gives the minimum value C(x) for the continuous correlation amount is calculated. Note that the pixel shift amount x is a real value and the unit is pixel.
[0036] [Number]
[0037] (2)
[0038] [Number]
[0039] (3)
[0040] [Number]
[0041] (4) The SLOP in Equation (4) represents the inclination of the change in the minimum and extremely small correlation amount and the adjacent correlation amount. In FIG. 4, as a specific example, let C(kj) = C(0) = 1000, C(kj - 1) = C(-1) = 1700, C(kj + 1) = C(1) = 1830. Also, kj = 0. From Equations (2) to (4), SLOP = 830 and x = -0.078 [pixel]. In the case of the in-focus state, the ideal value of the pixel shift amount x with respect to the signal sequence of image A and the signal sequence of image B is 0.00.
[0042] On the other hand, FIG. 5 shows the operation result when Equation (1) is applied to a minute block where noise exists. As shown in FIG. 5, due to the influence of randomly distributed noise, the correlation between the signal sequence of the A image and the signal sequence of the B image decreases. The minimum value of the correlation quantity C(k) becomes larger compared to the minimum value in FIG. 4, and the curve of the correlation quantity becomes flat overall (the absolute value of the difference between the maximum value and the minimum value is small).
[0043] In FIG. 5, as a specific example, let C(kj)=C(0)=1300, C(kj - 1)=C(-1)=1480, and C(kj + 1)=C(1)=1800. Also, kj = 0. From Equations (2) to (4), SLOP = 500 and x = -0.32 [pixel]. Compared with the operation result in the state where there is no noise shown in FIG. 4, the pixel shift amount x deviates from the ideal value.
[0044] When the correlation between a pair of image signal sequences is low, the change amount of the correlation quantity C(k) becomes small, and the curve of the correlation quantity becomes flat overall, so the value of SLOP becomes small. Also, when the subject image has low contrast, similarly, the correlation between a pair of image signal sequences becomes low, and the curve of the correlation quantity becomes flat. Based on this property, the reliability of the calculated pixel shift amount x can be determined by the value of SLOP. That is, when the value of SLOP is large, the correlation between a pair of image signal sequences is high, and when the value of SLOP is small, it can be determined that no significant correlation was obtained between the pair of image signal sequences. In this embodiment, since Equation (1) is used for the correlation operation, the correlation quantity C(k) is minimum and extremely small at the shift amount where the correlation between a pair of image signal sequences is the highest. As another method, a correlation algorithm in which the correlation quantity C(k) is maximum and extremely large at the shift amount where the correlation between a pair of image signal sequences is the highest may be used.
[0045] In step S305, the image processing unit 105 calculates the reliability. The reliability can be defined by C(kj) indicating the degree of coincidence of the two images calculated in step S304 and the value of SLOP as described above.
[0046] In step S306, the image processing unit 105 performs interpolation processing. In some cases, the pixel displacement amount calculated in step S304 cannot be adopted because the reliability calculated in step S305 is low. In that case, it is necessary to interpolate from the pixel displacement amounts calculated around. As an interpolation method, a median filter may be applied, or the pixel displacement amount data may be reduced and then enlarged again. Also, color data may be extracted from the (A-image signal + B-image signal) for imaging, and the pixel displacement amount may be interpolated using the color data.
[0047] In step S307, the image processing unit 105 calculates the defocus amount with reference to the pixel displacement amount x calculated in step S304. Specifically, the defocus amount (denoted as DEF) can be obtained by the following formula (5).
[0048] DEF = P × x (5) In formula (5), P is a conversion coefficient determined by the detection pitch (pixel arrangement pitch) and the distance between the projection centers of the two viewpoints on the left and right in a pair of parallax images, and the unit is mm / pixel.
[0049] In step S308, the image processing unit 105 calculates the distance from the defocus amount calculated in step S307. When the distance to the subject is Da, the focal position is Db, and the focal length is F, the following approximate formula (6) holds.
[0050]
Equation
[0051] (6) Therefore, the distance Da to the subject is expressed by the following formula (7).
[0052]
Equation
[0053] (7) When Db when DEF = 0 is set as Db0, the absolute distance Da' to the subject is expressed by the following formula (8).
[0054]
Equation
[0055] (8) The relative distance Da - Da' is expressed by the following formula (9) from formula (7) and formula (8).
[0056]
Equation
[0057] (9) As described above, by performing processing along the flow of FIG. 3, the pixel displacement amount, the defocus amount, and the distance information can be calculated from the A image signal for phase difference detection and the B image signal for phase difference detection.
[0058] FIG. 6 is a diagram showing the imaging system of the present embodiment. The imaging system includes an imaging device 100, a video signal processing device (processing device) 700, and a display device 300. The imaging device 100, the video signal processing device 700, and the display device 300 are connected by wire or wirelessly. The display device 300 displays a video signal (first video) input via a video input terminal (not shown). The coordinate detection device 601 is attached to the imaging device 100, and detects the position and orientation of the imaging device 100 by referring to markings applied to the ceiling, floor, etc. or by emitting infrared rays or the like and detecting the reflected light, and transmits the detected information to the video signal processing device 700. The imaging device 100 captures an image (second video) including the display area in which the first video of the display device 300 is displayed.
[0059] FIG. 7 is a block diagram of the video signal processing device 700. The video signal processing device 700 can input, output, and further record an image. Each part connected to the internal bus 701 can exchange data with each other via the internal bus 701.
[0060] The CPU 702 controls each part of the video signal processing device 700 in accordance with the programs stored in the ROM 703, using the RAM 704 as a work memory. The ROM 703 is a non-volatile recording element and stores programs for operating the CPU 702, various adjustment parameters, and the like. The RAM 704 is a volatile memory using semiconductor elements, and generally, a low-speed and small-capacity one is used compared to the frame memory 709. The frame memory 709 is an element capable of temporarily storing an image signal and reading it out when necessary. Since the image signal has a huge amount of data, a high-bandwidth and large-capacity one is required. In recent years, DDR4-SDRAM or the like has been used as the frame memory 709. By using the frame memory 709, for example, it becomes possible to perform processes such as synthesizing images that are different in time or cutting out only a necessary area.
[0061] The image processing unit 705 performs various image processes on the image data stored in the frame memory 709 and the recording medium 712 based on the control of the CPU 702. Note that the image processing unit 705 may be composed of dedicated circuit blocks for performing specific image processes. Depending on the type of image process, it is also possible for the CPU 702 to perform image processes according to a program without using the image processing unit 705.
[0062] The operation unit 710 is an interface with the outside of the device that receives the user's operations. The operation unit 710 is composed of a mouse, a keyboard, a touch panel, or the like.
[0063] The display unit 711 is a display device that can be visually recognized by the user. By displaying, for example, an image processed by the image processing unit 705 or a setting menu on the display unit 711, the user can check the operation status of the video signal processing device 700. In recent years, as the display unit 711, small and low-power consumption devices such as LCDs and organic ELs have been used. Also, the display unit 711 may be equipped with a thin film element such as a resistive film type or a capacitive type called a touch panel, and may be used as an alternative to the operation unit 710. The CPU 702 generates a character string for notifying the user of the setting state etc. of the video signal processing device 700 and a menu for setting the video signal processing device 700, superimposes them on the image processed by the image processing unit 705, and displays them on the display unit 711.
[0064] The video terminal unit 707 is composed of a plurality of video terminals. The video terminals are, for example, SDI, HDMI (registered trademark), and DisplayPort (registered trademark), etc. By outputting a video signal via the video terminal unit 707, it becomes possible to display an image in real time on an external monitor (not shown). Also, the CPU 702 can process the image input via the video terminal unit 707 in the image processing unit 705 and display it on the display unit 711.
[0065] The network module 706 is an interface for inputting and outputting image signals and audio signals. The network module 706 can communicate with external devices via the Internet or the like and can transmit and receive various data such as files, commands, video signals, and metadata. The transmission method by the network module 706 may be wireless or wired.
[0066] The recording medium 712 can record image data and various setting data, and a large-capacity storage element is used. As the recording medium 712, for example, an HDD, an SSD, etc. are used and are mounted on the recording medium I / F 708. When the user presses a predetermined button on the operation unit 710 (hereinafter referred to as the recording button), the CPU 702 starts recording the image data processed by the image processing unit 705 or the image data input from the video terminal unit 707 to the recording medium 712 via the recording medium I / F 708. When the user presses the recording button again, the recording is stopped. Also, the CPU 702 can record the image data in response to a signal indicating that recording is in progress received via the network module 706 or the video terminal unit 707.
[0067] Also, the video signal processing device 700 acquires lens information such as the focal length of the imaging device 100 and information such as exposure via the network module 706 or the video terminal unit 707. The CPU 702 reads out the 3D model data recorded in the recording medium I / F 708. Next, the CPU 702 re-renders the 3D model data in the image processing unit 705 based on the lens information and exposure information of the imaging device 100 and the information on the position and orientation of the imaging device 100 obtained from the coordinate detection device 601. Then, the CPU 702 generates a CG image suitable for the angle of view of the imaging device 100 and outputs it to the display device 300 as a background video.
[0068] Hereinafter, with reference to FIG. 8, the abnormal detection / removal process of the background video of the present embodiment will be described. FIG. 8 is a flowchart showing the abnormal detection / removal process of the background video of the imaging system of the present embodiment.
[0069] In step S801, the CPU 702 first reads out the three-dimensional model data recorded in the recording medium I / F 708. Next, the CPU 702 re-renders the three-dimensional model data in the image processing unit 705 based on the lens information of the imaging device 100, information such as exposure, and the position and orientation information of the imaging device 100 obtained from the coordinate detection device 601. Then, the CPU 702 generates a CG image adjusted to the angle of view of the imaging device 100 and outputs it to the display device 300 as a background video (first video). The background video is displayed on the display device 300.
[0070] In step S802, the CPU 102 acquires the image data (second video) including the subject and the background video captured by the imaging unit 107 and processed by the image processing unit 105, and the distance information, and stores them in the frame memory 111.
[0071] In step S803, the CPU 102 first reads out the image data and the distance information from the frame memory 111. Next, the CPU 102 determines the subject area in the image data where the subject is reflected and the background area (display area) where the background video is reflected as shown in FIG. 9 according to the distance information. Next, the CPU 102 adds flags indicating the subject area and the background area to the image data and stores them in the frame memory 111. The method for determining the subject area and the background area is to set a threshold value for the distance information, and if it is greater than or equal to the threshold value, it is determined as the background area, and if it is less than the threshold value, it is determined as the subject area. The threshold value may be set manually by the user, or may be automatically set according to the aperture and focus position of the lens unit 106. Then, the CPU 102 outputs the flags indicating the subject area and the background area and the image data to the video signal processing device 700 via the video terminal unit 109 or the network module 108.
[0072] In this embodiment, the CPU 102 discriminates the subject area and the background area. However, the CPU 702 may acquire the image data and the distance information from the imaging device 100 and discriminate the subject area and the background area.
[0073] In step S804, the CPU 702 acquires the flags indicating the subject area and the background area and the image data output by the imaging device 100 via the video terminal unit 707 or the network module 706. That is, the CPU 702 functions as acquisition means for acquiring flags indicating the subject area and the background area, respectively, as discrimination results of discriminating the subject area and the background area. When the resolution and shooting angle of the background video output in step S801 and the image data output by the imaging device in step S803 are different, the CPU 702 performs resizing processing and angle correction processing on the background video to match their resolution and shooting angle. Then, the CPU 702 compares the image data of the portions corresponding to those background areas. In this embodiment, the CPU 702 functions as detection means for comparing the image data of the portions corresponding to the background area of each video and detecting the difference (first difference).
[0074] In step S805, the CPU 702 determines whether the difference (first difference) of each pixel in the area compared in step S804 is equal to or greater than a predetermined value (equal to or greater than a first predetermined value). In the present embodiment, the CPU 702 determines whether there are pixels whose difference is equal to or greater than the predetermined value. When the CPU 702 determines that the difference of each pixel is equal to or greater than the predetermined value, it executes the process of step S806. When the CPU 702 determines that the difference of each pixel is not equal to or greater than the predetermined value, it ends this flow. Note that in the present embodiment, it is determined whether there are pixels whose difference of each pixel in the compared area is equal to or greater than the predetermined value, but the present invention is not limited to this. For example, the process of step S806 may be configured to be executed when the average value of the differences of each pixel is equal to or greater than the predetermined value, or when the number of pixels whose difference of each pixel is equal to or greater than the predetermined value is counted and is equal to or greater than a predetermined number of pixels.
[0075] In step S806, the CPU 702 determines whether the image data is being recorded. When the CPU 702 determines that it is being recorded, it executes the process of step S807, and when it determines that it is not being recorded, it executes the process of step S808.
[0076] In step S807, the image processing unit 705 generates a corrected image of the image data input from the imaging device 100 based on the control of the CPU 702. Specifically, first, the CPU 702 performs FFT (Fast Fourier Transform) on the background video output in step S801 and the portion corresponding to the background area of the image data output by the imaging device in step S803, respectively. Thereby, the signal on the spatial axis can be converted into the signal on the frequency axis. Then, the respective frequency components are compared, and the signal of the frequency component that exists only in the image data input from the imaging device 100 is determined as moire, and the signal processing for removing the frequency component (the frequency component is smaller than the first predetermined amount) is performed. As a means for removing a specific frequency component, for example, there is a filter process using a notch filter. Then, by performing inverse FFT on the signal on the frequency axis from which the moire frequency component has been removed and converting it into the signal on the spatial axis, image data with moire removed can be obtained.
[0077] In step S808, the CPU 702 generates a warning display for notifying that moire or the like has occurred and an unintended background video has been captured, and outputs it to the imaging device 100 and the display device 300 via the video terminal unit 707. Then, the warning is displayed on the display unit 114 of the imaging device 100, the display device 300, and the display unit 711 of the video signal processing device 700, and is conveyed to the user. Further, the warning display input to the imaging device 100 may be output from the video terminal unit 109, and the warning may be displayed on an external monitor (not shown). In the present embodiment, the CPU 702 functions as a control means for performing control to notify that the difference in the image data of the portion corresponding to the background area is large and an unintended background video with moire or the like has been captured.
[0078] Note that in the present embodiment, the CPU 702 displays a warning display on the display unit, but is not limited thereto as long as it can notify that moire or the like has occurred and an unintended background video has been captured. For example, control for notifying by vibration, sound, or the like may be performed.
[0079] In step S809, when the resolution or shooting angle of the background video output in step S801 is different from that of the corrected image generated in step S807, the CPU 702 performs resizing processing or angle correction processing on the background video to match the resolution or shooting angle. Then, the image data of the portions corresponding to those background areas are compared. The CPU 702 determines whether correction is possible according to the difference of each pixel in the compared areas. When the CPU 702 determines that the difference of each pixel in the compared areas is equal to or greater than a predetermined value (equal to or greater than a second predetermined value), that is, when it determines that correction is impossible, it executes the process of step S810. When the CPU 702 determines that the difference of each pixel in the compared areas is less than the predetermined value (less than the second predetermined value), that is, when it determines that correction is possible, it executes the process of step S811. In this embodiment, it is determined whether the difference of each pixel in the compared areas is equal to or greater than the predetermined value, but the present invention is not limited to this. For example, it may be configured to execute the process of step S810 when the average value of the differences of each pixel is equal to or greater than the predetermined value, or when the number of pixels whose differences of each pixel are equal to or greater than the predetermined value is counted and is equal to or greater than a predetermined number of pixels.
[0080] In step S810, the CPU 702 generates a warning display indicating that moire or the like has occurred and an unintended background video has been shot, and outputs it to the imaging device 100 via the video terminal unit 707. Then, the warning is displayed on the display unit 114 of the imaging device 100 and the display unit 711 of the video signal processing device 700, and the user is notified so as not to affect the recorded image, and the image data input from the imaging device 100 is recorded on the recording medium I / F 708. Note that the warning display input to the imaging device 100 may be output from the video terminal unit 109 and the warning may be displayed on an external monitor (not shown).
[0081] In step S811, the CPU 702 performs a process of recording the corrected image generated by the image processing unit 705 in step S807 on the recording medium I / F 708 and a process of outputting it via the video terminal unit 707.
[0082] In step S812, the CPU 702 performs a process of recording, as metadata, data indicating the area corrected in step S807 and a time code indicating the time and frame at which the correction was performed on the recording medium I / F 708, or a process of outputting via the video terminal unit 707.
[0083] Note that in this embodiment, a configuration is adopted in which it is determined whether correction is possible in step S809, and the corrected image is recorded and output in step S811. However, the present invention is not limited to this. The process of step S810 may be performed without determining whether correction is possible in step S809, and the background area may be corrected in a post-processing step after shooting using the recorded captured image, the metadata recorded in step S812, and the background video output in step S801. At this time, the CPU 702 records, in the recording medium 712, the captured image after correction and the metadata indicating that moire has been corrected in association with each other. Further, the CPU 702 may output and record, in the recording medium 712 together with the captured image, metadata indicating that moire has occurred in the display area of the display device 300 without performing correction for moire removal in step S807 in order to perform correction later.
[0084] Also, in this embodiment, as a method for removing moire in step S807, an example is shown in which image data is converted into a signal on the frequency axis to identify and remove the frequency components of moire. However, the present invention is not limited to this. Image data obtained by converting the background video output in step S801 in accordance with the resolution and shooting angle of the image data output by the imaging device 100 may be generated, and the background area of the image data input from the imaging device 100 may be replaced to remove moire.
[0085] Also, in this embodiment, as a means for separating the subject area and the background area, distance information calculated using the A image signal and the B image signal obtained from the imaging unit 107 is used. However, the present invention is not limited to this. A configuration may be adopted in which the distance to the subject is obtained using another means such as a distance sensor, or a configuration may be adopted in which the subject area and the background area are separated using image segmentation technology without using distance information.
[0086] By performing the processes described above, it is possible to detect and remove moiré that occurs when shooting with a background image displayed on the display. Further, if there is a difference equal to or greater than a predetermined value between the background image output from the video signal processing device 700 to the display device 300 and the background image displayed on the display device 300 and captured by the imaging device 100, it is possible to detect and remove not only moiré but also changes such as luminance, color, distortion, and pixel / area defects. Further, as a method for detecting such changes, it is not limited to the method of taking the difference described above, and for example, any calculation method that can compare the background image output to the display device 300 and the background image captured by the imaging device 100 for each region or each pixel, such as a ratio, can be applied.
[0087] Further, not limited to the moiré detection method in the present embodiment, the occurrence of moiré may be estimated from the optical design information of the imaging device 100 (such as the pixel pitch of the sensor, focus, zoom, and subject distance) and the information on the pixel pitch of the display device 300.
[0088] <Second Embodiment> In the first embodiment, an example of removing moiré by generating a corrected image of the image data input from the imaging device 100 was described.
[0089] Moiré may occur during imaging due to interference with a repeating pattern included in the video itself being displayed as the background video. In that case, the occurrence of moiré can be suppressed by correcting the background video to be displayed. Therefore, in the present embodiment, an example of removing moiré by correcting and outputting the background video output from the video signal processing device 700 will be described.
[0090] FIG. 10 is a flowchart showing the abnormal detection / removal process of the background video of the imaging system of the present embodiment. In the present embodiment, the same reference numerals are given to the same or similar configurations and steps as in the first embodiment, and redundant explanations are omitted.
[0091] In step S1001, based on the control of the CPU 702, the image processing unit 705 generates a corrected image of the background video output in step S801 and outputs it to the display device 300. Specifically, for example, an image is generated by applying a low-pass filter to the background video output in step S801 to reduce high-frequency components (the high-frequency components are smaller than a second predetermined amount), and the image is output to the display device 300.
[0092] In step S1002, when the resolution and shooting angle of the background video output in step S1001 and the image data captured in step S802 are different, the CPU 702 performs resizing processing and angle correction processing to match the resolution and shooting angle on the background video. Then, the CPU 102 compares the image data of the portions corresponding to those background areas. When the CPU 702 determines that the difference between the pixels of the compared areas is equal to or greater than a predetermined value (equal to or greater than a third predetermined value), it executes the process of step S810. When the CPU 702 determines that the difference between the pixels of the compared areas is less than the predetermined value (less than the third predetermined value), it executes the process of step S1003. Note that in this embodiment, it is determined whether the difference between the pixels of the compared areas is equal to or greater than the predetermined value, but the present invention is not limited to this. For example, the process of step S810 may be configured to be executed when the average value of the differences between the pixels is equal to or greater than the predetermined value, or when the number of pixels whose differences between the pixels are equal to or greater than the predetermined value is counted and is equal to or greater than a predetermined number of pixels.
[0093] In step S1003, the CPU 102 captures the background video and the subject output to the display device 300 in step S1001 and records them on the recording medium I / F 110. Further, the CPU 102 may output the image data captured via the video terminal unit 109 and record it on the recording medium 712 of the video signal processing device 700.
[0094] By performing the processes described above, it is possible to detect and remove abnormalities such as moiré that occur when shooting with the background video displayed on the display. [Other Embodiments] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and causing one or more processors in a computer of the system or apparatus to read and execute the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.
[0095] The disclosure of the present embodiment includes the following configurations and methods. (Configuration 1) A processing device that causes a display device to display a first video and acquires a second video captured by an imaging device that images an angle of view including a display area of the display device, an acquisition unit that acquires a determination result obtained by determining an area in the second video that images a subject existing between the display area and the imaging device and the display area; a detection unit that detects a change in the second video by comparing the first video and the video of the display area in the second video; and a control unit that performs first control for notifying information regarding a detection result by the detection unit, the processing device being characterized by having the above. (Configuration 2) The processing device according to Configuration 1, wherein the determination result is obtained according to a distance from the imaging device to a shooting target. (Configuration 3) The processing device according to Configuration 1 or 2, wherein the detection unit performs the comparison by detecting a difference between a third video obtained by correcting at least one of a resolution and a shooting angle of the first video and the display area. (Configuration 4) The detection unit detects a first difference between the first video and the display area, The control unit performs the first control assuming that the change has been detected when the processing device is not recording a video and the first difference is equal to or greater than a first predetermined value, the processing device according to any one of Configurations 1 to 3. (Configuration 5) The detection means detects a first difference between the first video and the display area, The detection means detects a second difference between the display area in a fourth video obtained by correcting the second video so that the first difference becomes smaller than a first predetermined amount, and a fifth video obtained by correcting at least one of the resolution and the shooting angle of the first video, When the second difference is equal to or greater than a second predetermined value, the control means performs second control to notify information regarding the second difference, records the second video, and when the second difference is less than the second predetermined value, the control means records the fourth video, and the processing device according to any one of Configurations 1 to 4. (Configuration 6) The control means of the processing device according to Configuration 5 records data indicating at least one of information on an area, a time, and a frame in which correction is performed in the second video. (Configuration 7) The detection means detects a first difference between the first video and the display area, The detection means detects a third difference between a seventh video obtained by correcting at least one of the resolution and the shooting angle of a sixth video obtained by correcting the first video so that the first difference becomes smaller than a second predetermined amount, and the display area in the second video, When the third difference is equal to or greater than a third predetermined value, the control means performs third control to notify information regarding the third difference, records the second video, and when the third difference is less than the third predetermined value, the control means does not perform the third control and records the second video, and the processing device according to any one of Configurations 1 to 4. (Configuration 8) The control means of the processing device according to Configuration 7 records data indicating at least one of information on an area, a time, and a frame in which correction is performed in the first video. (Configuration 9) A processing device that causes a display device to display a first video and acquires a second video captured by an imaging device that captures an angle of view including a display area of the display device, detection means for detecting whether moire occurs in the display area in the second video; a processing apparatus, comprising: control means for recording, on a recording medium, a video obtained by correcting the moire and information on the area where the moire is detected when the moire is detected by the detection means. (Configuration 10) an imaging system, comprising: the processing apparatus according to any one of Configurations 1 to 9; a display device for displaying a first video; an imaging device for imaging a second video including the display area of the display device. (Method 1) a processing method for causing a display device to display a first video and obtaining, from an imaging device, a second video obtained by imaging an angle of view including the display area of the display device, the method comprising: a first step of obtaining a discrimination result for discriminating between an area of an object imaged between the display area in the second video and the imaging device and the display area; a second step of detecting a change in the second video by comparing the first video and the video of the display area in the second video; a third step of performing first control for notifying information regarding the detection result in the second step. (Configuration 11) a program for causing a computer to execute the processing method according to Method 1.
[0096] As described above, the preferred embodiments of the present invention have been described. However, the present invention is not limited to these embodiments, and various modifications and changes can be made within the scope of the gist thereof.
Description of Reference Numerals
[0097] 100 Imaging device 300 Display device 700 Video signal processing device
Claims
1. A processing device that causes a display device to display a first video and acquires, from an imaging device, a second video obtained by imaging an angle of view including the display area of the display device, the processing device comprising: acquisition means for acquiring a determination result obtained by determining an area of a subject imaged between the display area and the imaging device in the second video and the display area; detection means for detecting a change in the second video by comparing the first video and the video of the display area in the second video; and control means for performing first control for notifying information regarding a detection result by the detection means, the processing device being characterized by comprising the acquisition means, the detection means, and the control means.
2. The processing device according to claim 1, wherein the determination result is acquired according to a distance from the imaging device to a shooting target.
3. The processing device according to claim 1 or 2, wherein the detection means performs the comparison by detecting a difference between a third video obtained by correcting at least one of a resolution and a shooting angle of the first video and the display area.
4. The detection means detects a first difference between the first video and the display area, and the control means performs the first control assuming that the change has been detected when the processing device is not recording a video and the first difference is equal to or greater than a first predetermined value, the processing device according to claim 1 or 2 being characterized by this.
5. The detection means detects a first difference between the first video and the display area, the detection means detects a second difference between the display area in a fourth video obtained by correcting the second video so that the first difference becomes smaller than a first predetermined amount and a fifth video obtained by correcting at least one of a resolution and a shooting angle of the first video, and the control means performs second control for notifying information regarding the second difference, records the second video when the second difference is equal to or greater than a second predetermined value, and records the fourth video when the second difference is less than the second predetermined value, the processing device according to claim 1 or 2 being characterized by this.
6. The processing device according to claim 5, wherein the control means records data indicating at least one of information on an area corrected by the second video, a time, and a frame.
7. The detection means detects a first difference between the first video and the display area, The detection means detects a third difference between a sixth video obtained by correcting the first video so that the first difference is smaller than a second predetermined amount, a seventh video obtained by correcting at least one of the resolution and the shooting angle, and the display area in the second video. When the third difference is equal to or greater than a third predetermined value, the control means performs third control to notify information regarding the third difference and records the second video. When the third difference is less than the third predetermined value, the control means does not perform the third control and records the second video. The processing apparatus according to claim 1 or 2, characterized in that.
8. The control means according to claim 7, wherein the control means records data indicating at least one of information on a region, a time, and a frame in which correction is performed in the first video.
9. A processing apparatus that causes a display device to display a first video and acquires a second video captured by an imaging device of an imaging angle including a display area of the display device, detection means for detecting whether moire is generated in the display area in the second video; A control means for recording on a recording medium a video obtained by correcting the moire and information on a region where the moire is detected when the moire is detected by the detection means.
10. The processing apparatus according to claim 1 or 2, a display device that displays a first video, An imaging system comprising: an imaging device that images a second video including a display area of the display device.
11. A processing method for causing a display device to display a first video and acquiring a second video captured by an imaging device of an imaging angle including a display area of the display device, A first step of acquiring a determination result of determining between a region where a subject existing between the display area in the second video and the imaging device is imaged and the display area; A second step of detecting a change in the second video by comparing the first video and the video of the display area in the second video; A third step of performing first control for notifying information regarding a detection result in the second step.
12. A program characterized by causing a computer to execute the processing method according to claim 11.
Citation Information
Patent Citations
Moire removing method and manufacturing method of display
JP2008011334A