Display video generation device and program
The display image generating device addresses the challenge of unifying resolution and frame rate by using a shared pixel structure and conversion models to reduce data transmission and enhance image quality perception.
Patent Information
- Application Number
- JP2024072393
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-26
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies face challenges in unifying resolution and frame rate for video signals, leading to increased data output and transmission requirements, despite improving image quality perception.
A display image generating device that includes an acquisition unit, generation unit, and output unit, utilizing a shared pixel structure to calculate total pixel values and convert pixel information using conversion models based on scene analysis, reducing data transmission by generating video signals with higher frame rates and resolutions.
The solution effectively reduces data transmission by converting pixel information into higher frame rate and resolution formats, optimizing image quality perception while minimizing data output.
Smart Images

Figure 2025167603000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a display image generating device and a program. [Background technology]
[0002] Generally, the quality of image quality (subjective image quality) perceived by viewers (viewers) is influenced by the resolution, frame rate, and the magnitude of the subject's movement. For example, moving objects can experience motion blur due to the charge accumulation time of a camera during image capture and the hold-type display of a display. Therefore, even when captured at a high resolution, viewers are less likely to perceive the difference between images captured at a low resolution and images captured at a low frame rate. On the other hand, stationary objects exhibit small movements. Therefore, even when captured at a low frame rate, viewers are less likely to perceive the difference between images captured at a high frame rate and images captured at a high frame rate. Patent Document 1 discloses a technology that uses an image sensor capable of binning to change imaging conditions, such as resolution and frame rate, for each area within an imaging area. This technology reduces the signal data volume of captured images by selectively using different imaging conditions: low resolution and high frame rate, and high resolution and low frame rate. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2022-123539 Summary of the Invention [Problem to be solved by the invention]
[0004] Since the resolution and frame rate of video signals captured using the technology described in Patent Document 1 vary from area to area, they must be unified when displaying the video. However, if the resolution and frame rate are unified to high image quality and high frame rate so that viewers do not perceive the image quality as poor, the amount of data output from the imaging device will be larger than the amount of data captured at the time of capture. Therefore, even if the technology described in Patent Document 1 is used, there is a problem in that the amount of data to be transmitted will end up being large.
[0005] The present invention has been made in view of the above-mentioned points, and has an object to provide a technique for reducing the amount of data to be transmitted. [Means for solving the problem]
[0006] One aspect of the present invention is a display image generating device that includes: an acquisition unit that acquires, from an imaging device in which a pixel region having a plurality of photoelectric conversion elements is divided into a plurality of areas, thereby enabling imaging at different resolutions and frame rates for each of the areas; a first video signal that includes first pixel information based on charges photoelectrically converted by the photoelectric conversion elements and scene information generated by performing motion analysis or illuminance analysis from the captured video signal at coordinates corresponding to the coordinates of the area within the pixel region; a generation unit that generates, for each of the areas, a second video signal that has more information than the first video signal and has images of multiple consecutive frames by converting the first pixel information into second pixel information using either a first conversion model or a second conversion model determined in accordance with the scene information at coordinates corresponding to the coordinates of the area; and an output unit that outputs the second video signal generated by the generation unit.
[0007] In addition, one aspect of the present invention is an imaging device as described above, which includes n or more (n is a natural number greater than or equal to 2) photoelectric conversion elements, has a shared pixel structure capable of calculating a total pixel value based on the combined charges photoelectrically converted by the n photoelectric conversion elements, calculates the total pixel value for each subframe obtained by dividing a frame period into n, sets the total pixel value for each subframe to the pixel value of a pixel at each coordinate corresponding to the time of each subframe, and outputs first pixel information indicating the pixel value of the pixel at the coordinate and the scene information indicating whether the pixel value of each pixel in the area including the pixel at the coordinate is based on the total pixel value.
[0008] Furthermore, in one aspect of the present invention, in the above-mentioned display image generating device, a subframe period is each period obtained by dividing a frame period into n periods, where n is the number of photoelectric conversion elements included in the shared pixel structure, and the first conversion model converts, for the n subframes included in the second pixel information, the pixel values of each of the n pixels in the subframe into pixel values included in the first pixel information and at coordinates corresponding to the coordinates within the pixel region of each of the n pixels in the subframe, and the second conversion model converts, for any of the n subframes included in the second pixel information, the pixel values of the n pixels in the subframe into the pixel value of the pixel included in the first pixel information, for that one subframe out of the n subframes included in the second pixel information, which is determined according to the coordinates within the pixel region of that one pixel out of the n pixels included in the first pixel information, for each of the n pixels included in the first pixel information.
[0009] In addition, according to one aspect of the present invention, in the display image generating device described above, the scene information is stored in an ancillary area of the first image signal.
[0010] In one aspect of the present invention, in the display image generating device described above, the ancillary area further stores information indicating whether the information stored in the ancillary area is scene information.
[0011] In one aspect of the present invention, in the display image generating device described above, the first image signal stores the scene information in a number equal to or greater than the number of areas including pixels indicated by the first pixel information.
[0012] Another aspect of the present invention is a program that causes a computer to execute the following steps for an imaging device in which a pixel region having a plurality of photoelectric conversion elements is divided into a plurality of areas, thereby enabling imaging with different resolutions and frame rates for each of the areas: an acquisition step of acquiring a first video signal including first pixel information based on charges photoelectrically converted by the photoelectric conversion elements and scene information generated by performing motion analysis or illuminance analysis from the captured video signal at coordinates corresponding to the coordinates of the area within the pixel region; a generation step of generating a second video signal having more information than the first video signal and including images of multiple consecutive frames by converting the first pixel information into second pixel information for each of the areas using either a first conversion model or a second conversion model determined in accordance with the scene information at the coordinates corresponding to the coordinates of the area; and an output step of outputting the second video signal generated by the generation step. [Effects of the Invention]
[0013] According to the present invention, the amount of data to be transmitted can be reduced. [Brief explanation of the drawings]
[0014] [Figure 1] 1A and 1B are diagrams for explaining aspects of an imaging system according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of an area control image sensor. [Figure 3] FIG. 10 is a diagram illustrating an example of a shared pixel structure. [Figure 4] 10 is a timing chart showing an example of timing for reading pixel values in a shared pixel structure according to the embodiment. [Figure 5] 10A and 10B are diagrams for explaining an example of mapping of pixel values performed by a video signal processing circuit. [Figure 6] 4A and 4B are diagrams for explaining an example of a first video signal according to the embodiment. [Figure 7] FIG. 10 is a diagram illustrating an example of a data structure of an ancillary area. [Figure 8] FIG. 2 is a diagram illustrating an example of a functional configuration of a display image generating device. [Figure 9] 10 is a timing chart showing how a second video signal is generated. [Figure 10] 10 is a flowchart illustrating an example of a flow of processing performed by the display image generating device. [Figure 11] 10 is a flowchart illustrating an example of a flow of processing performed by the imaging system. [Figure 12] FIG. 10 is a diagram illustrating a first video signal according to Modification 1 of the embodiment. [Figure 13] FIG. 10 is a diagram for explaining an example of a data structure of an ancillary region according to Modification 2 of the present embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0015] [Embodiment] The imaging system and display image generating device according to the present embodiment will be described in detail below with reference to the accompanying drawings, showing preferred embodiments. Note that the present embodiment is not limited to these embodiments and includes various modifications and improvements. In other words, the components described below include those that would be easily conceivable to a person skilled in the art and those that are substantially identical, and the components described below can be combined as appropriate. Furthermore, various omissions, substitutions, or modifications of the components of the present embodiment can be made without departing from the spirit of the present invention.
[0016] First, an embodiment of an imaging system will be described with reference to FIG.
[0017] 1 is a diagram illustrating an aspect of an imaging system 1 according to an embodiment. The imaging system 1 includes an imaging device 10, a display image generation device 20, and a display device 30. The imaging system 1 may include multiple imaging devices 10 and one display image generation device 20, or may include one imaging device 10 and multiple display image generation devices 20. An example in which the imaging system 1 includes one imaging device 10 and one display image generation device 20 will be described below.
[0018] The imaging device 10 and the display image generation device 20 are often installed independently at separate locations. The imaging device 10 and the display image generation device 20 are connected to each other via wireless communication or wired communication. In wired communication, the imaging device 10 transmits information to the display image generation device 20 using, for example, an SDI (Serial Digital Interface) signal. In this embodiment, an example is described in which a video signal is transmitted using a 12G-SDI signal, which is a standard capable of transmitting a 4K, 60 fps video signal. Note that FIG. 1 illustrates an example in which the imaging device 10 directly transmits information to the display image generation device 20. However, this embodiment is not limited to this example, and the imaging device 10 may temporarily transmit information to a storage unit (not shown) located on a server or cloud. The information transmitted to the storage unit is then transmitted to the display image generation device 20.
[0019] The display image generation device 20 may be installed independently of the display device 30 in a position close to each other. The display image generation device 20 may also be included in the housing of the display device 30, that is, may be provided as a component of the display device 30. The display image generation device 20 and the display device 30 are connected to each other by a wired connection, for example, a signal line that transmits a video signal.
[0020] The imaging device 10 includes a lens 11, an imaging unit 12, a scene information processing circuit 13, a drive circuit 14, a video signal processing circuit 15, a video output processing unit 16, and a display unit 17. The video output processing unit 16 is realized, for example, using a computer and software. The video output processing unit 16 may also be realized using an electronic circuit, if necessary. Furthermore, each functional unit does not have to be included in a single device, and the imaging device 10 may be configured from multiple devices.
[0021] The lens 11 collects incident light and forms an image of the subject on the imaging surfaces of area control image sensors 122 (area control image sensors 122R, 122G, and 122B) and event sensor 123, which will be described later.
[0022] The imaging unit 12 includes a spectrometer 121, an area control image sensor 122, and an event sensor 123. FIG. 1 shows area control image sensor 122R, area control image sensor 122G, and area control image sensor 122B as examples of the area control image sensor 122. Note that in this embodiment, a four-chip system using four RGB-infrared sensors is shown. However, this embodiment is not limited to this example, and a three-chip system or a single-chip system that does not use infrared may also be used. Note that in the case of a single-chip system, the area control image sensor 122 may have red, green, and blue photoelectric conversion elements PD arranged in a quad Bayer array.
[0023] The spectroscope 121 separates the light incident through the lens 11. The spectroscope 121 separates the incident light into red (R), green (G), blue (B), and infrared (IR). The spectroscope 121 then directs each of the separated light components to the area control image sensor 122R, the area control image sensor 122G, the area control image sensor 122B, and the event sensor 123. The spectroscope 121 may be, for example, a four-plate prism optical system. The red, green, and blue light form images on the imaging surfaces of the area control image sensors 122 for each of the RGB colors. The infrared light forms an image on the imaging surface of the event sensor 123.
[0024] Each area control image sensor 122 (area control image sensor 122R, area control image sensor 122G, area control image sensor 122B) receives visible light incident from the spectroscope 121 with a photoelectric conversion element. Each area control image sensor 122 outputs an electrical signal based on charges photoelectrically converted by the photoelectric conversion element to the scene information processing circuit 13 and the video signal processing circuit 15. Alternatively, the signal may be transferred from the video signal processing circuit 15 to the scene information processing circuit 13, or vice versa. The area control image sensor 122 is driven by a drive circuit 14, which will be described later.
[0025] The event sensor 123 receives infrared light from the spectroscope 121 and outputs an electrical signal based on charges photoelectrically converted by the multiple photoelectric conversion elements PD to the scene information processing circuit 13. Specifically, the event sensor 123 may detect a change in the position of the subject and output information on the coordinates (x coordinate and y coordinate) of the pixel where the change occurred and the time.
[0026] The area control image sensor 122 will be specifically described below with reference to FIG.
[0027] 2 is a diagram illustrating an example of the configuration of the area control image sensor 122. The configuration of the area control image sensor 122 will be described in detail with reference to the diagram. The area control image sensor 122 includes a pixel unit 1221 and a control unit 1222. The area control image sensor 122 may have a layered structure in which the pixel unit 1221 and the control unit 1222 overlap each other. This allows the area control image sensor 122 to arrange more photoelectric conversion elements PD in the same area.
[0028] The pixel section 1221 has a plurality of photoelectric conversion elements PD on the imaging surface. The pixel section 1221 is sometimes referred to as a pixel region. The pixel section 1221 is divided into a plurality of control areas CA, each including a plurality of photoelectric conversion elements PD. FIG. 2 shows an example in which 48×48 photoelectric conversion elements PD are arranged in the pixel section 1221. Also shown is an example in which the pixel section 1221 is divided into 6×6 control areas CA. That is, the control area CA includes 8×8 photoelectric conversion elements PD.
[0029] FIG. 3 is a diagram illustrating an example of a shared pixel structure SPS. The control area CA has one or more shared pixel structures SPS in which multiple photoelectric conversion elements PD share one floating diffusion (FD), amplification transistor (SF), and selection transistor (SL). The number of photoelectric conversion elements PD in a shared pixel structure SPS may be referred to as n (a natural number greater than or equal to 2). The following description assumes that the control area CA has 2 × 2 = 4 shared pixel structures SPS. The area control image sensor 122 can acquire pixel values by treating multiple photoelectric conversion elements PD in the shared pixel structure SPS as a single photoelectric conversion element PD. Hereinafter, this pixel acquisition method may be referred to as "binning." Specifically, the area control image sensor 122 calculates pixel values based on the sum of the charges of multiple photoelectric conversion elements PD, and defines the pixel value of the single photoelectric conversion element PD obtained by combining the multiple photoelectric conversion elements PD. When pixel values are calculated by binning, the amount of charge photoelectrically converted over time increases, allowing appropriate imaging even with a short exposure time. In other words, when binning is performed, imaging can be performed by lowering the resolution and increasing the frame rate while keeping the data rate the same. Note that the area control image sensor 122 can also acquire pixel values by regarding each of the multiple photoelectric conversion elements PD of the shared pixel structure SPS as an independent photoelectric conversion element PD.
[0030] In the following description of this embodiment, the number of photoelectric conversion elements PD included in the shared pixel structure SPS is assumed to be 2 × 2 = 4. The four photoelectric conversion elements PD may be referred to as the first photoelectric conversion element PD1, the second photoelectric conversion element PD2, the third photoelectric conversion element PD3, and the fourth photoelectric conversion element PD4. That is, when binning is performed, the horizontal resolution is halved, the vertical resolution is halved, and the frame rate is quadrupled. Note that the arrangement of the photoelectric conversion elements PD included in the shared pixel structure SPS is not limited to a 2 × 2 square. The arrangement of the photoelectric conversion elements PD included in the shared pixel structure SPS may be, for example, an arrangement in which the number of pixels in the horizontal direction (H) is different from the number of pixels in the vertical direction (V), or may be a square arrangement such as a 3 × 3 or 4 × 4.
[0031] Returning to FIG. 2 , the control unit 1222 will be described. The control unit 1222 includes a mode control circuit 1223, a pixel driving circuit 1224, an A / D conversion circuit 1225, and an output circuit 1226 for each control area CA. The control unit 1222 controls the imaging mode of the pixel unit 1221 for each control area CA. The imaging mode is the imaging conditions during imaging. Factors for determining the imaging mode include resolution, frame rate, exposure time, etc. For example, any of the resolution, frame rate, exposure time, etc. may be different for each imaging mode. Imaging modes according to the embodiment include, for example, a basic mode, a high-speed mode that performs binning, a high-brightness mode that shortens the exposure time of the basic mode, and a low-illuminance mode that lengthens the exposure time of the basic mode. As an example, the basic mode may be set to 4K resolution and 60 [fps], and the high-speed mode may be set to 2K resolution and 240 [fps].
[0032] The mode control circuit 1223 acquires drive information from the drive circuit 14, which will be described later. The mode control circuit 1223 outputs control information indicating an imaging mode corresponding to the acquired drive information to the pixel drive circuit 1224. The pixel drive circuit 1224 drives in the imaging mode indicated by the control information acquired from the mode control circuit 1223. For example, the pixel drive circuit 1224 does not perform binning when it acquires control information indicating the basic mode, but performs binning when it acquires control information indicating the high-speed mode.
[0033] The A / D conversion circuit 1225 acquires the charges photoelectrically converted by each photoelectric conversion element PD and converts them into an electrical signal. The output circuit 1226 outputs the electrical signal converted by the A / D conversion circuit 1225 to the scene information processing circuit 13 and the video signal processing circuit 15.
[0034] 4 is a timing chart showing an example of the timing for reading out pixel values in the shared pixel structure SPS according to the embodiment. One frame period T is composed of subframe (SF) periods equal to the number of photoelectric conversion elements PD in the shared pixel structure. That is, in this embodiment, one frame period T is composed of four subframe periods.
[0035] FIG. 4A shows an example of pixel value readout timing in high-speed mode (with binning). In high-speed mode, pixel values are read out from all photoelectric conversion elements PD in the shared pixel structure SPS for each subframe, i.e., at intervals of one frame period T / 4. Specifically, within one frame period T, pixel values based on the sum of the charges photoelectrically converted by all photoelectric conversion elements PD (first photoelectric conversion element PD1 to fourth photoelectric conversion element PD4) in the shared pixel structure SPS are read out at the following times: T / 4, 2T / 4, 3T / 4, and T. Therefore, the exposure time of each photoelectric conversion element PD is the subframe period.
[0036] FIG. 4B shows an example of pixel value readout timing in basic mode (without binning). In basic mode, pixel values are read out one by one from the multiple photoelectric conversion elements PD in the shared pixel structure SPS for each subframe. Specifically, within one frame period T, a pixel value is read out from the first photoelectric conversion element PD1 after T / 4 has elapsed, a pixel value is read out from the second photoelectric conversion element PD2 after 2T / 4 has elapsed, a pixel value is read out from the fourth photoelectric conversion element PD4 after 3T / 4 has elapsed, and a pixel value is read out from the fourth photoelectric conversion element PD4 after T has elapsed. Therefore, the exposure period of each photoelectric conversion element PD is one frame period T.
[0037] Returning to Fig. 2, the scene information processing circuit 13 included in the imaging device 10 will be described. First, the scene information processing circuit (scene information processing section) 13 acquires an electrical signal (hereinafter, sometimes referred to as first pixel information) indicating the pixel value of each photoelectric conversion element PD from the imaging section 12. The scene information processing circuit 13 also acquires an electrical signal from the imaging section 12 indicating information on the coordinates (x coordinate and y coordinate) and time of a pixel that has changed.
[0038] Next, the scene information processing circuit 13 performs motion analysis, illuminance analysis, etc. based on the first pixel information output from the imaging unit 12 to determine the situation at the time of imaging of the subject for each control area CA. The scene information processing circuit 13 may also determine the situation at the time of imaging of the imaging device 10, such as whether the imaging device 10 is moving or zooming, using a sensor (not shown) provided in the imaging device 10. Specifically, the scene information processing circuit 13 may calculate a difference between frames for each control area CA, and determine that the subject is moving quickly if the difference is greater than an arbitrarily set threshold, or determine that the subject is moving slowly if the difference is smaller than the arbitrarily set threshold. The situation at the time of imaging (imaging situation) may be determined by any method suitable for this embodiment, such as a method of calculating a motion vector.
[0039] Furthermore, the scene information processing circuit 13 determines imaging conditions suitable for each control area CA according to the determined imaging conditions, such as setting a high frame rate for a control area CA for imaging a fast-moving subject, a short exposure time for a control area CA for imaging a bright subject, and a low frame rate for a control area CA for imaging a dark area. The scene information processing circuit 13 may also determine an imaging mode, such as basic mode or high-speed mode, according to the imaging conditions. The scene information processing circuit 13 determines the imaging conditions and generates scene information indicating the determined imaging conditions (imaging mode). The scene information processing circuit 13 outputs the generated scene information to the drive circuit 14, the video signal processing circuit 15, and the video output processing unit 16.
[0040] The drive circuit 14 drives the imaging unit 12 based on the scene information acquired from the scene information processing circuit 13. Specifically, the drive circuit 14 outputs the drive information to drive each control area CA of each area control image sensor 122 in an imaging mode corresponding to the imaging situation indicated by the scene information. A drive circuit 14 is provided for each of one or more area control image sensors 122. Specifically, one area control image sensor 122 may be driven by one drive circuit 14, or multiple area control image sensors 122 may be driven by one drive circuit 14.
[0041] In the above description, an example is shown in which, when the imaging device 10 continuously captures frame images, the scene information processing circuit 13 acquires one of the frame images (first pixel information) captured by the imaging unit 12 and determines the imaging mode after capturing the frame image (for example, the image one frame after the frame image). That is, an example is shown in which the scene information processing circuit 13 determines the imaging mode for the next image capture depending on the determined imaging conditions. This embodiment is not limited to this example, and the scene information processing circuit 13 may estimate the current imaging conditions based on multiple past imaging conditions. This makes it possible to capture an image in an imaging mode that corresponds to the current imaging conditions.
[0042] The video signal processing circuit 15 acquires first pixel information from the imaging unit 12. The video signal processing circuit 15 also acquires scene information from the scene information processing circuit 13. The video signal processing circuit 15 converts the first pixel information for each control area CA according to the scene information. For example, if an image is captured with an exposure time set to one-fourth of a reference exposure time, the video signal processing circuit 15 may convert the signal gain of the captured image by four times. Furthermore, if the frame rate is higher than the reference rate due to binning (capturing in high-speed mode), the video signal processing circuit 15 may convert the frame rate to the reference frame rate by mapping pixel values acquired for each subframe within one frame. The video signal processing circuit 15 outputs the converted first pixel information to the video output processing unit 16 and the display unit 17.
[0043] The conversion of the video signal performed by the video signal processing circuit 15 will be described in detail with reference to FIG. 5. FIG. 5 is a diagram illustrating an example of mapping of the video signal performed by the video signal processing circuit 15. The video signal processing circuit 15 stores first pixel information captured in high-speed mode in a buffer memory for each subframe. The stored first pixel information for each subframe is designated IN(x, y, a). x and y indicate coordinates, and a indicates the subframe number. In this embodiment, the resolution of the video captured in high-speed mode is 2K (1920 × 1080), and the number of subframes is four. The video output when captured in high-speed mode is a signal equivalent to 2K and 240 [fps]. However, 2K and 240 [fps] video signals are not standardized in SDI (Serial Digital Interface), a transmission standard mainly used in broadcasting systems. Therefore, this signal needs to be converted into a 4K and 60 [fps] 12G-SDI signal and output. If the converted first pixel information equivalent to 4K, 60 [fps] is OUT(x, y), an example of an equation for converting (mapping) the first pixel information of 2K, 240 [fps] to the first pixel information equivalent to 4K, 60 [fps] is shown below.
[0044] OUT(2x-1, 2y-1)=IN(x, y, 1) OUT(2x, 2y-1)=IN(x, y, 2) OUT(2x-1, 2y)=IN(x, y, 3) OUT(2x, 2y)=IN(x, y, 4)
[0045] The video output processing unit 16 acquires the converted first pixel information from the video signal processing circuit 15. The video signal processing circuit 15 also acquires scene information from the scene information processing circuit 13. The video signal processing circuit 15 outputs the first pixel information and scene information as a first video signal. For example, the video output processing unit 16 may transmit the first video signal according to a transmission standard such as a 4K, 60 fps 12G-SDI signal. The 12G-SDI signal is divided into four 2K-sized sub-images and then further divided into two, resulting in eight DataStreams for transmission. FIG. 6 is a diagram illustrating an example of a first video signal according to an embodiment. The first pixel information (PI) is stored in an active area, and the scene information (SI) is stored in an ancillary area. It is desirable that the scene information be collectively stored in one of the eight DataStreams to facilitate processing on the receiving side. Additionally, the scene information is preferably stored on the same horizontal line as the pixel in question or on a nearby horizontal line to minimize delay.
[0046] The ancillary area in which scene information is stored will be specifically described with reference to Fig. 7. Fig. 7 is a diagram illustrating an example of the data structure of the ancillary area. The ancillary area stores information such as "ADF1", "ADF2", "ADF3", "DID", "DBN", "DC", "UDW1" to "UDW240", and "CS" in that order.
[0047] A fixed value is stored in "ADF1" through "ADF3" and "DC." "DID" stores information indicating the type of auxiliary information stored in the ancillary area. "DID" stores information indicating, for example, scene information, text information, audio information, etc. In this embodiment, "0x2CF" stored in "DID" indicates that scene information is stored in the ancillary area. "DBN" stores information indicating the continuity of the information stored in the ancillary area. "UDW" stores auxiliary information, i.e., scene information. In this embodiment, four imaging modes are shown as examples: basic mode, high-speed mode, low-illumination mode, and high-brightness mode. Therefore, a 2-bit area is required to store scene information for one control area CA. There are 960 control areas CA per horizontal line in the pixel area. If scene information for four control areas CA is stored in one word of "UDW", the total number of words is 960 / 4 = 240, and scene information can be stored in the ancillary area ("UDW1" to "UDW240") of one horizontal line of video signal (first video signal). "CS" stores a checksum of the information from "DID" to "UDW240".
[0048] Display unit 17 displays an image corresponding to the first pixel information output from video signal processing circuit 15. The image displayed on display unit 17 includes an image captured in basic mode and an image captured in high-speed mode and converted (mapped) by video signal processing circuit 15. Display unit 17 may be included in imaging device 10, such as a rear monitor of imaging device 10, or may be a display or the like that is connected to imaging device 10 via a signal line or network that transmits the first pixel information and exists independently of imaging device 10.
[0049] FIG. 8 is a diagram illustrating an example of the functional configuration of the display image generation device 20. The display image generation device 20 includes an acquisition unit 21, a generation unit 22, and an output unit 23. These functional units are implemented using, for example, a computer and software. Alternatively, each functional unit may be implemented using an electronic circuit, as necessary. Furthermore, each functional unit does not have to be included in a single device, and the display image generation device 20 may be configured from multiple devices. Regarding the first pixel information included in the first video signal, the first pixel information acquired in the basic mode has correct pixel coordinates corresponding to the coordinates of the photoelectric conversion element PD. On the other hand, the video signal acquired in the high-speed mode is converted (mapped) by the video signal processing circuit 15 after binning. Therefore, the pixel coordinates do not correspond to the coordinates of the photoelectric conversion element PD. Therefore, to display the video captured by the imaging device 10 on the display device 30, the video signal must be reconstructed into a video signal suitable for viewing. The display image generation device 20 is sometimes referred to as a video signal conversion device.
[0050] The acquisition unit 21 acquires a first video signal including first pixel information (PI) and scene information (SI) from the imaging device 10. The acquisition unit 21 outputs the acquired first video signal to the generation unit 22.
[0051] The generation unit 22 generates a second video signal (4K, 240 [fps]) by converting the first pixel information (4K, 60 [fps]) acquired by the acquisition unit 21 into second pixel information using a conversion model that differs for each control area CA. An example of generation performed by the generation unit 22 will be described with reference to FIG. 9. FIG. 9 is a timing chart showing the manner in which the second video signal is generated. First, a conversion model (first conversion model) that converts the first pixel information captured in the basic mode into second pixel information will be described. The first conversion model converts the 4K, 60 [fps] first pixel information into second pixel information of 4K, 240 [fps] by repeatedly outputting the first pixel information for each of four subframes that make up one frame period T without changing the coordinates of the pixels indicated by the first pixel information. Here, the number of subframes corresponds to the number of photoelectric conversion elements PD included in the shared pixel structure SPS. In other words, the first conversion model converts the first pixel information into second pixel information with the same resolution as that used during image capture but with a frame rate four times faster.
[0052] Next, a conversion model (second conversion model) for converting first pixel information captured in high-speed mode and converted by the video signal processing circuit 15 into second pixel information will be described. The second conversion model converts the pixel values of pixels in the shared pixel structure SPS indicated by the first pixel information of 4K, 60 [fps] into second pixel information of 4K, 240 [fps] as pixel values of multiple pixels in the shared pixel structure SPS of a subframe corresponding to the coordinates of the pixel. Specifically, the pixel value of the pixel located in the upper left corner of the shared pixel structure SPS in the first pixel information is converted into the second pixel information as the pixel values of four pixels in the shared pixel structure SPS in the first subframe (SF1), the pixel located in the upper right corner is converted into the pixel values of four pixels in the shared pixel structure SPS in the second subframe (SF2), the pixel located in the lower left corner is converted into the pixel values of four pixels in the shared pixel structure SPS in the third subframe (SF3), and the pixel located in the lower right is converted into the pixel values of four pixels in the shared pixel structure SPS in the fourth subframe (SF4). That is, the second conversion model generates second pixel information having twice the horizontal resolution and twice the vertical resolution at the same frame rate as during imaging. Note that the specific numerical values used to explain the conversion model are not limited to those exemplified here. Furthermore, the conversion model may have a different configuration from the first conversion model and the second conversion model depending on the scene information, i.e., the imaging mode. Furthermore, the conversion model is not necessarily limited to conversion to second pixel information having the highest resolution (4K) and the highest frame rate (240 [fps]) in the basic mode and high-speed mode.
[0053] The output unit 23 acquires the generated second pixel information from the generation unit 22. The output unit 23 outputs a second video signal including the second pixel information to the display device 30, which is connected by a signal line. That is, the display image generation device 20 may be configured to be included in the display device 30, or may exist independently from the display device 30 and be connected to the display device 30 via a signal line capable of transmitting the second video signal. The display device 30 may also store a program that causes the display image generation device 20 to execute the functions of the display image generation device 20.
[0054] 10 is a flowchart illustrating an example of the flow of processing performed by display image generation device 20. Display image generation device 20 acquires a first image signal including first pixel information and scene information from imaging device 10 (step S101). For each control area CA, display image generation device 20 converts the first pixel information of 4K, 60 [fps] into second pixel information of 4K, 240 [fps] using a conversion model according to the scene information (step S102). Display image generation device 20 outputs the generated second image signal including the second pixel information to display device 30 (step S103).
[0055] Display device 30 displays 4K, 240 [fps] video in accordance with the second video signal output from display video generation device 20. Display device 30 may be, for example, a television, a mobile terminal, a monitor, or the like.
[0056] FIG. 11 is a flowchart illustrating an example of the processing flow performed by the imaging system 1. The imaging device 10 captures an image of a subject using different imaging modes for each control area CA (step S201). The imaging device 10 converts the resolution and frame rate of the images (first pixel information) captured in the different imaging modes to conform to a certain transmission standard (12G-SDI) (step S202). The imaging device 10 transmits a first video signal including the converted first pixel information and scene information indicating the imaging mode for each control area CA to the display image generation device 20 (step S203). The display image generation device 20 generates second pixel information from the first pixel information using a conversion model corresponding to the scene information (step S204). The display image generation device 20 transmits a second video signal including the second pixel information to the display device 30 via a wired signal line (step S205). The display device 30 displays an image corresponding to the second video signal (step S206).
[0057] Note that the above description shows an example in which the display image generation device 20 is included in the display device 30. However, the present embodiment is not limited to this example, and the display image generation device 20 may be included in the imaging device 10. The imaging device 10 and the display image generation device 20 may be connected by a signal line capable of transmitting the first video signal. Furthermore, the imaging device 10 may store a program for executing the functions of the display image generation device 20. In this case, the display unit 17 according to the embodiment may display the second video signal, i.e., an image converted to be suitable for display. That is, the display device 30 may be configured to be included in the imaging device 10 as the display unit 17.
[0058] In this embodiment, an example is shown in which the imaging device 10 transmits the first video signal to the display image generation device 20. However, this embodiment is not limited to this example. The imaging device 10 may transmit the first video signal to a recording device (not shown). In this case, the recording device temporarily stores the first video signal and then outputs the first video signal to the display image generation device 20. The recording device may be, for example, an SSD (Solid State Drive), HDD (Hard Disk Drive), etc.
[0059] In addition, in the present embodiment, an example is shown in which the video signal processing circuit 15 converts the gain of the video signal. However, the present embodiment is not limited to this example, and the gain may be converted when the display image generating device 20 generates the second pixel information. In this case, the configuration of the video signal processing circuit 15 can be simplified.
[0060] [Modification 1 of this embodiment] In the above-described embodiment, an example is shown in which 2-bit scene information is stored in one first video signal (horizontal line). In contrast, in Modification 1 of this embodiment, a case will be described in which there are many variations in imaging modes by changing the exposure time in stages, or by changing the resolution and frame rate in stages. In Modification 1 of this embodiment, the scene information is assumed to be 4 bits.
[0061] FIG. 12 is a diagram illustrating a first video signal according to Modification 1 of this embodiment. FIG. 12(A) shows an example of the first video signal according to Modification 1 of this embodiment. FIG. 12(B) shows an example of the data structure of the ancillary area according to Modification 1 of this embodiment. Each word in "UDW1" to "UDW240" is 10 bits. Therefore, scene information for two control areas CA is stored in one word. In this case, 960 / 2 = 480 words are required to store scene information for all control areas CA on one horizontal line, so two horizontal blanking periods are used. When one control area CA is 4 × 4 pixels, the number of control areas CA is 540 × 960, so even if two lines are used, it fits within 1080 lines. Therefore, scene information for all control areas CA can be stored in one DataStream. [Modification 2 of this embodiment] In the above embodiment, an example has been described in which the control area CA includes four sharing pixel structures SPS (4 × 4 pixels). In contrast, in Modification 2 of this embodiment, a case will be described in which the control area CA includes one sharing pixel structure SPS (2 × 2 pixels). In Modification 2 of this embodiment, the scene information is assumed to be 4 bits.
[0062] If scene information for two control areas CA is stored in one word, 1920 / 2 = 960 words are required. FIG. 13 is a diagram illustrating an example of the data structure of an ancillary area according to Modification 2 of this embodiment. Because scene information is determined by the object being captured, similar to general video signals, there is a high correlation between adjacent pixels, and nearby pixels and control areas CA are likely to have the same scene information. Therefore, the ancillary area can be used efficiently by storing scene information and information indicating the length (number) of control areas CA in which that scene information is consecutive. For example, as shown in FIG. 13, if the lower 4 bits are used as scene information and the upper 4 bits are used as the number of control areas CA in which that scene information is consecutive, scene information for up to 16 control areas CA can be described together. If the average number of consecutive control areas CA with the same scene information is eight, the required number of words is 1920 / 8 = 240 words, allowing scene information to be stored in the ancillary area of one horizontal line. Note that the variation of scene information (4 bits) and the number of consecutive control areas CA with the same scene information (continuous length) are not limited to these examples. Also, by using a set of two words, the maximum length of consecutive entries that can be described may be extended to 255. Furthermore, without using the above-described description method, the scene information may be divided into four DataStreams and stored in the same manner as in the above-described embodiment, or the scene information may be compressed and stored in one DataStream.
[0063] [Summary of this embodiment] According to the above-described embodiment, the imaging system 1 includes an imaging device 10, a display image generating device 20, and a display device 30. The imaging device 10 includes an area control image sensor 122 that divides a pixel region having a plurality of photoelectric conversion elements PD into a plurality of control areas CA, thereby enabling imaging with different resolutions and frame rates for each control area CA. The imaging device 10 also includes a scene information processing circuit 13 that generates scene information by performing motion analysis or illuminance analysis on an image signal captured at coordinates corresponding to the coordinates of the control areas CA within the pixel region. The imaging device 10 outputs a first video signal including pixel information and scene information based on charges photoelectrically converted by the photoelectric conversion elements PD. The display image generating device 20 generates a second video signal having a larger amount of information than the first video signal and including frame images of multiple consecutive frames, based on the first pixel information and scene information acquired from the imaging device 10. The imaging device 10 according to this embodiment has a shared pixel structure SPS. Therefore, the imaging device 10 can change imaging conditions (imaging modes) such as resolution, frame rate, and exposure time for each control area CA while maintaining the same data rate during imaging. Furthermore, based on the scene information, the display image generation device 20 can reconstruct (convert) images captured under different imaging conditions to generate images suitable for display. Furthermore, when displaying captured images on the display device 30, the imaging system 1 according to this embodiment transmits a first image signal, which is an image before processing into an image suitable for display and has a relatively small amount of data, to the display image generation device 20. The imaging system 1 generates a second image signal from the transmitted first image signal, which has a high resolution and high image quality suitable for display, i.e., a relatively large amount of data, and then displays the second image signal on the display device 30. Therefore, by transmitting the first pixel information, which has a small amount of data because it is not processed yet, and the scene information required for conversion into the second pixel information, rather than generating and transmitting the second image signal in the imaging device 10, the amount of data required for transmitting the image to be displayed on the display device 30 can be reduced. Therefore, the imaging system 1 according to this embodiment can change the shooting conditions for each subject while minimizing the amount of data required for transmitting the image. Furthermore, according to this embodiment, data before processing, which has a relatively small amount of data, is transmitted, eliminating the need for data compression.This prevents degradation of the image quality of the displayed video and prevents delays.
[0064] Furthermore, according to the above-described embodiment, the imaging device 10 and the display image generation device 20 are connected to each other via wireless communication or wired communication. The imaging device 10 and the display image generation device 20 may be installed at a distance from each other. When the imaging device 10 and the display image generation device 20 are located at a distance from each other, there has been a demand for reducing the amount of data transmitted. According to the imaging system 1 of the embodiment, first pixel information, which is a relatively small amount of data before being processed into an image suitable for display, and scene information indicating the imaging mode are transmitted, and second pixel information, which is suitable for display and has a relatively large amount of data, is generated at the destination display image generation device 20. Therefore, the amount of data transmitted between the imaging device 10 and the display image generation device 20, which are connected via a wireless communication network, can be reduced.
[0065] Furthermore, according to the above-described embodiment, the display image generation device 20 and the display device 30 are wired to each other by a signal line that transmits a second video signal. The second video signal generated by the display image generation device 20 has a high resolution, a high frame rate, and a large amount of data. Therefore, by wiredly connecting the display image generation device 20 that generates the second video signal and the display device 30 by a signal line that transmits the second video signal, the burden of transmitting the second video signal can be reduced.
[0066] According to the above-described embodiment, the imaging device 10 and the display image generation device 20 are wired to each other by a signal line that transmits the first video signal. That is, the display image generation device 20 may be included in the imaging device 10. By including the display image generation device 20 in the imaging device 10, the imaging device 10 can capture 4K, 60 [fps] video, which has a relatively small amount of data, and then generate 4K, 240 [fps] video using the display image generation device 20, without capturing 4K, 240 [fps] video. This allows the amount of data exchanged within the imaging device 10 to be reduced.
[0067] Furthermore, according to the above-described embodiment, the imaging device 10 further includes a display unit 17 that displays an image based on the first video signal. This allows a person capturing an image using the imaging device 10 to roughly grasp the state of the captured subject. Furthermore, the image displayed on the display unit 17 has a smaller data volume than the image displayed on the display device 30. This reduces the burden on the imaging device 10 of displaying the image captured.
[0068] According to the above-described embodiment, the display image generating device 20 includes an acquisition unit 21, a generation unit 22, and an output unit 23. The acquisition unit 21 acquires a first video signal from an imaging device 10, which divides a pixel region having a plurality of photoelectric conversion elements PD into a plurality of control areas CA and can capture images at different resolutions and frame rates for each control area CA. The first video signal includes first pixel information based on charges photoelectrically converted by the photoelectric conversion elements PD and scene information obtained by performing motion analysis or illuminance analysis on the captured video signal at coordinates corresponding to the coordinates of the control areas CA within the pixel region. The generation unit 22 converts the first pixel information into second pixel information for each control area CA using either a first conversion model or a second conversion model determined according to the scene information at the coordinates corresponding to the coordinates of the control area CA, thereby generating a second video signal having a larger amount of information than the first video signal and including images of multiple consecutive frames. The output unit 23 outputs the second video signal generated by the generation unit 22. When displaying an image captured by the imaging device 10 on the display device 30, the display image generating device 20 generates a second video signal suitable for display and having a relatively large amount of data based on an unprocessed first video signal having a relatively small amount of data. That is, the display image generating device 20 according to the embodiment can transmit a video signal with a relatively small amount of data and generate a second video signal with a relatively large amount of data when displayed. Therefore, by transmitting the first video signal including the unprocessed first pixel information and the scene information required for generation, the amount of data to be transmitted can be reduced compared to when the second video signal is transmitted from the imaging device 10 to the display device 30.
[0069] Furthermore, according to the above-described embodiment, the imaging device 10 includes n or more photoelectric conversion elements PD (n is a natural number equal to or greater than 2) and has a shared pixel structure SPS capable of calculating a total pixel value based on the combined charges photoelectrically converted by the n photoelectric conversion elements PD. The imaging device 10 calculates the total pixel value for each subframe obtained by dividing one frame period into n subframes, sets the total pixel value for each subframe as the pixel value of a pixel at a corresponding coordinate corresponding to the time of the subframe, and outputs first pixel information indicating the pixel value of the pixel at the corresponding coordinate and the scene information indicating whether the pixel value of each pixel in the area including the pixel at the corresponding coordinate is based on the total pixel value. When imaging is performed in imaging modes with different resolutions and frame rates, such as basic mode and high-speed mode, for each control area CA, the captured video includes a mixture of areas captured in basic mode and areas captured in high-speed mode, resulting in a lack of a unified standard. Therefore, it is necessary to conform to a single transmission standard. The video signal processing circuit 15 converts video captured in high-speed mode to the standard for video captured in basic mode, enabling transmission according to a single transmission standard.
[0070] Furthermore, according to the above-described embodiment, a subframe period refers to each period obtained by dividing one frame period into n periods, where n is the number of photoelectric conversion elements included in the shared pixel structure. The first conversion model converts, for n subframes included in the second pixel information, pixel values of each of the n pixels in the subframe into pixel values of coordinates included in the first pixel information and corresponding to the coordinates in the pixel region of each of the n pixels in the subframe. The second conversion model converts, for each of the n pixels included in the first pixel information, pixel values of each of the n subframes included in the second pixel information, determined according to the coordinates in the pixel region of any one of the n pixels included in the first pixel information. The above-described conversion model enables a large amount of data-sized video with a uniform resolution and frame rate to be generated from video captured in different imaging modes. Because the display image generation device 20 converts the first pixel information into second pixel information using a conversion model corresponding to scene information, the imaging device 10 can transmit a video signal with a relatively small amount of data. Therefore, according to the display image generation device 20 according to the embodiment, the amount of data transmitted by the imaging device 10 can be reduced.
[0071] Furthermore, according to the above-described embodiment, the scene information is stored in the ancillary area of the first video signal. Conventionally, in a video signal, pixel information is stored in the active area, and auxiliary information such as audio information and text information is stored in the ancillary area. In the embodiment, the scene information is stored in the ancillary area. Therefore, the first pixel information and scene information can be transmitted using an existing system.
[0072] Furthermore, according to the above-described embodiment, the ancillary area further stores information indicating whether the information stored in the ancillary area is scene information. That is, the ancillary area stores "DID" information. This allows the ancillary area to store information other than scene information, such as audio information and text information.
[0073] Furthermore, according to the above-described embodiment, the first video signal stores scene information for more than the number of control areas CA containing pixels indicated by the first pixel information. The number of scene information pieces depends on the number of control areas CA. For an area containing four shared pixel structures SPS, i.e., 4 × 4 pixels, one control area CA must superimpose scene information for a 960 × 540 area in the ancillary area. When pixels on a horizontal line in the imaging area are read, the first pixel information indicating the pixel information and 960 areas of scene information are stored in the first video signal containing the read pixels or in a first video signal adjacent to the read pixels, thereby minimizing display delay. The 960 pieces of scene information can be stored in groups of four, "UDW1" to "UDW240." Therefore, display delay can be minimized by transmitting the pixels in the control area CA and the scene information for the control area CA in the same first video signal.
[0074] Note that all or part of the functions of the units of the imaging device 10, display image generating device 20, and display device 30 in the above-described embodiments may be realized by recording a program for realizing these functions on a computer-readable recording medium, and reading and executing the program recorded on the recording medium into a computer system. Note that the term "computer system" here includes hardware such as an OS and peripheral devices.
[0075] Furthermore, "computer-readable recording media" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage units such as hard disks built into computer systems. Furthermore, "computer-readable recording media" may also include devices that dynamically store programs for a short period of time, such as communication lines used when transmitting programs over networks like the Internet or communication lines like telephone lines, or devices that store programs for a fixed period of time, such as volatile memory within a computer system that serves as a server or client in such cases. Furthermore, the program may be one that realizes part of the aforementioned functions, or may be one that can realize the aforementioned functions in combination with a program already stored in the computer system.
[0076] Although one embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to the above, and various design modifications can be made without departing from the scope of the present invention. Furthermore, the configurations described in the above-described embodiments and examples can be combined. [Explanation of symbols]
[0077] 1...imaging system, 10...imaging device, 11...lens, 12...imaging section, 13...scene information processing circuit, 14...drive circuit, 15...video signal processing circuit, 16...video output processing section, 17...display section, 121...spectroscope, 122...area control image sensor, 123...event sensor, 1221...pixel section, 1222...control section, 1223...mode control circuit, 1224...pixel drive circuit, 1225...A / D conversion circuit, 1226...output circuit, 20...display image generating device, 21...acquisition section, 22...generation section, 23...output section, 30...display device, PD...photoelectric conversion element, SF...subframe, PI...first pixel information, SI...scene information, SPS...shared pixel structure
Claims
1. an acquisition unit that acquires, from an imaging device in which a pixel region having a plurality of photoelectric conversion elements is divided into a plurality of areas and imaging can be performed with different resolutions and frame rates for each of the areas, a first video signal including first pixel information based on charges photoelectrically converted by the photoelectric conversion elements and scene information generated by performing motion analysis or illuminance analysis on the captured video signal at coordinates corresponding to the coordinates of the areas within the pixel region; a generation unit that converts the first pixel information into second pixel information for each of the areas using either a first conversion model or a second conversion model that is determined according to the scene information of coordinates corresponding to the coordinates of the area, thereby generating a second video signal having a larger amount of information than the first video signal and including images of a plurality of consecutive frames; an output unit that outputs the second video signal generated by the generation unit; A display image generating device comprising:
2. The imaging device includes n or more (n is a natural number of 2 or more) photoelectric conversion elements, and has a shared pixel structure capable of calculating a total pixel value based on charges obtained by combining charges photoelectrically converted by the n photoelectric conversion elements, calculates the total pixel value for each subframe obtained by dividing a period of one frame into n subframes, sets the total pixel value for each subframe to a pixel value of a pixel at each coordinate corresponding to a time of each of the subframes, and outputs first pixel information indicating the pixel value of the pixel at the coordinates, and the scene information indicating whether the pixel value of each pixel in the area including the pixel at the coordinates is based on the total pixel value. The display image generating device according to claim 1 .
3. The subframe period is a period obtained by dividing one frame period into n periods, the number of which is the number of the photoelectric conversion elements included in the shared pixel structure, the first conversion model converts, for the n subframes included in the second pixel information, pixel values of n pixels in the subframes into pixel values that are included in the first pixel information and have coordinates corresponding to coordinates within the pixel region of the n pixels in the subframes; The second conversion model is determined according to the coordinates in the pixel area of any one of the n pixels included in the first pixel information, and for any one of the n subframes included in the second pixel information, conversion is performed for each of the n pixels included in the first pixel information, such that pixel values of the n pixels in the subframe are set to the pixel value of the one pixel included in the first pixel information. The display image generating device according to claim 2 .
4. The scene information is stored in an ancillary area of the first video signal.
3. The display image generating device according to claim 1 or 2.
5. The ancillary area further stores information indicating whether the information stored in the ancillary area is scene information. The display image generating device according to claim 4 .
6. The first video signal stores the scene information equal to or greater than the number of the areas including the pixel indicated by the first pixel information.
3. The display image generating device according to claim 1 or 2.
7. On the computer, an acquisition step for acquiring a first video signal including first pixel information based on charges photoelectrically converted by the photoelectric conversion elements and scene information generated by performing a motion analysis or an illuminance analysis on the captured video signal at coordinates corresponding to the coordinates of the areas in the pixel region, for an imaging device in which a pixel region having a plurality of photoelectric conversion elements is divided into a plurality of areas and imaging can be performed with different resolutions and frame rates for the areas; a generating step of converting, for each of the areas, the first pixel information into second pixel information using either a first conversion model or a second conversion model determined according to the scene information of coordinates corresponding to the coordinates of the area, thereby generating a second video signal having a larger amount of information than the first video signal and including images of a plurality of consecutive frames; an output step of outputting the second video signal generated in the generating step; A program that executes the following.
Citation Information
Patent Citations
Imaging element
JP2022123539A