Control device, imaging device, control method, and program

The imaging system addresses the issue of decreased visibility in vibrating environments by extracting frame images based on the sky-sea boundary, generating clear digest videos with improved stability and coherence.

JP2026066862APending Publication Date: 2026-04-17CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2024-10-07
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing image capture technologies on vibrating platforms, such as ships, suffer from decreased visibility in digest images due to vibrations caused by waves or wind, making it difficult to generate clear and coherent digest videos.

Method used

An imaging system that extracts frame images based on the boundary between the sky and sea, using methods like edge detection and machine learning, to generate a digest video with high visibility by selecting images where the boundary line is horizontal and centered, thereby suppressing the effects of vibrations.

Benefits of technology

The system effectively generates highly visible digest videos by selectively extracting images with a stable boundary line, ensuring the subject remains centered and clear despite vibrations, enhancing image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026066862000001_ABST
    Figure 2026066862000001_ABST
Patent Text Reader

Abstract

When images are taken in vibrating locations such as on a ship, the visibility of the digest video is reduced due to vibrations caused by waves. [Solution] The control unit is a control device that extracts frame images from multiple images to generate a digest video, and comprises: acquisition means for acquiring multiple images in a time series; and generation means for detecting the boundary line between a first region and a second region contained in each of the multiple images, and based on the boundary line, extracting frame images from the multiple images to generate the digest video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0006] , , , , ,

[0001] The present invention relates to a control device, an imaging device, a control method, and a program.

Background Art

[0002] For example, when transmitting an image captured by an external device from a ship or the like, due to communication restrictions or the like, it may be necessary to extract frame images from a plurality of frame images of the captured image and transmit a summarized image as a digest image. Regarding the generation of such a digest image, Patent Document 1 discloses a method for generating a fast-forward image from an image captured by a wearable camera.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the above-described technology, when an image is captured at a vibrating location such as on a ship, there arises a problem that the visibility of the digest image decreases due to vibrations caused by waves or the like.

[0005] Therefore, the present invention provides a technology for a digest image with high visibility.

Means for Solving the Problems

[0006] To solve this problem, for example, the control device of the present invention has the following configuration. That is, In a control device that extracts frame images from a plurality of images to generate a digest image, acquisition means for acquiring a plurality of images in time series; A generation means that detects the boundary line between a first region and a second region contained in each of the plurality of images, and extracts frame images from the plurality of images based on the boundary line to generate the digest video, It is equipped with. [Effects of the Invention]

[0007] According to the present invention, it is possible to provide a technology for highly visible digest videos. [Brief explanation of the drawing]

[0008] [Figure 1] A block diagram showing the overall configuration of the imaging system according to the first embodiment. [Figure 2] A diagram showing the vibration state when the ship is rocking. [Figure 3] This figure shows an example of an image obtained when an imaging device photographs a subject from the ship itself. [Figure 4] A diagram illustrating the calculation of the approximate straight line L of the boundary line. [Figure 5] A flowchart illustrating the process of generating frame images of a digest video according to the first embodiment. [Figure 6] A block diagram showing the overall configuration of the imaging system according to the second embodiment. [Figure 7] A diagram illustrating how periodic vibrations caused by waves affect the ship. [Figure 8] A flowchart illustrating the process of generating frame images for a digest video according to the second embodiment. [Figure 9] A flowchart illustrating the process of generating frame images for a digest video according to the third embodiment. [Figure 10] A flowchart illustrating the process of generating frame images for a digest video according to the fourth embodiment. [Modes for carrying out the invention]

[0009] The embodiments will be described in detail below with reference to the attached drawings. Note that the following embodiments do not limit the invention as defined in the claims. While the embodiments describe multiple features, not all of these features are essential to the invention, and the features may be combined in any way. Furthermore, in the attached drawings, identical or similar configurations are given the same reference numerals, and redundant descriptions are omitted.

[0010] (First Embodiment) The configuration of the imaging system 10 of the first embodiment will be described below with reference to Figure 1. Figure 1 is a block diagram showing the overall configuration of the imaging system 10 of the first embodiment.

[0011] The imaging system 10 generates and transmits video for monitoring at another location, such as another ship or on land, from multiple images taken on a ship, for example, to assist in ship navigation and enable autonomous navigation. Here, since the imaging system 10 transmits video using communication methods with limited bandwidth, such as satellite communication and wireless communication, it generates and transmits a digest video containing frame images extracted from the images, rather than a video containing all the images taken. Furthermore, in order to suppress the effects of vibrations such as waves, the imaging system 10 extracts frame images based on the boundary between the sky and the sea to generate a digest video with high visibility.

[0012] As shown in Figure 1, the imaging system 10 includes an imaging device 100 and an external device 111. The imaging device 100 and the external device 111 are connected via a network 110. This allows the imaging device 100 and the external device 111 to send and receive signals from each other.

[0013] The imaging device 100 may be, for example, a digital camera installed on a vibrating object such as a ship. The imaging device 100 captures the outside from the installed ship to generate a video having a plurality of images or a plurality of frame images, generates a digest video from the frame images extracted from the images, and transmits it to the external device 111. The imaging device 100 includes an imaging unit 101, an image processing unit 102, a communication unit 103, a control unit 104, a storage unit 105, and a bus 107. The imaging unit 101, the image processing unit 102, the communication unit 103, the control unit 104, and the storage unit 105 are connected to each other via the bus 107 so as to be able to transmit and receive electrical signals. The combination of the image processing unit 102, the communication unit 103, the control unit 104, the storage unit 105, and the bus 107 may be a computer or an information processing device.

[0014] The imaging unit 101 captures a subject and generates an image. The term "image" may include a still image, a moving image, a video, a frame image included in a moving image and a video, and their data. The imaging unit 101 includes a lens and an imaging element.

[0015] The lens is arranged to transmit light from the subject and guide it to the imaging element. The lens may be configured to be able to adjust the zoom magnification, the focus position, etc. using a lens drive motor in addition to a single-focus lens. Further, the lens may have a mechanism for adjusting the amount of light incident on the imaging element by aperture control and ND filter insertion / removal. The lens may be in a state where a filter for transmitting / attenuating a specific wavelength can be inserted / removed.

[0016] The imaging element outputs an electrical signal (also referred to as a pixel signal) corresponding to the image of the subject to the image processing unit 102 by photoelectrically converting the subject image (optical image) formed through the lens. The imaging element may be any of, for example, a CMOS (Complementary Metal-Oxide-Semiconductor) sensor and a CCD (Charge Coupled Device) sensor. Here, the imaging element may have sensitivity to light in the visible light region, for example, and may also have sensitivity to light in the infrared light region.

[0017] The imaging unit 101 may be configured to include a plurality of imaging elements and optically convert a subject image formed through different lenses or the same lens. In this embodiment, the imaging unit 101 transfers the pixel signal to the image processing unit 102 as it is, but the transfer method is not limited to this. The imaging unit 101 may perform signal processing inside the imaging element before transferring the pixel signal to the image processing unit 102.

[0018] Based on the control signal from the control unit 104, the image processing unit 102 performs predetermined signal processing such as correction processing and development processing on the pixel signal from the imaging unit 101 and outputs the processed image signal. The image processing unit 102 may be either a processor such as a GPU or one or more circuits such as an ASIC. The correction processing includes white balance processing, gamma processing, noise reduction processing, and the like.

[0019] The communication unit 103 outputs the image signal from the image processing unit 102 to the external device 111 as a video signal. By converting it into data compliant with the communication protocol at that time, efficient transmission becomes possible. When transmitting, it may be transmitted as still image data or compressed and transmitted as moving image data. Also, the communication unit 103 can receive a control signal for controlling the imaging device 100 from the external device 111.

[0020] The control unit 104 comprehensively controls each part within the imaging device 100 and controls the imaging device 100 based on various parameter settings. The control unit 104 may change the control of the imaging device 100 by changing various parameter settings according to the control signal received by the communication unit 103.

[0021] The control unit 104 may be, for example, a CPU (Central Processing Unit). The control unit 104 may have other processors such as an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), an NPU (Neural Processing Unit), and a QPU (Quantum Processing Unit) instead of a CPU, or in addition to a CPU.

[0022] Some or all of the functions of the control unit 104 and the image processing unit 102 are realized by one or more processors, including a CPU, reading programs stored in the storage of the memory unit 105, expanding them into memory, and executing them. Alternatively, some or all of the functions of the control unit 104 and the image processing unit 102 may be realized by one or more circuits, such as an ASIC (Application Specific Integrated Circuit) and a PLD (Programmable Logic Device) including an FPGA (Field Programmable Gate Array).

[0023] The control unit 104 may perform the digest video generation process by executing a computer program for digest video generation processing stored in the memory unit 105. In this case, the control unit 104 is an example of an acquisition means and a generation means. For example, the control unit 104 acquires a set of N consecutive images from video captured by the imaging unit 101 and processed by the image processing unit 102. The consecutive images may be multiple frame images included in multiple still images or videos, and may be consecutive images in chronological order. The control unit 104 detects subjects such as other ships, the sky, and the sea from the images and calculates the position of the subjects within the screen. The control unit 104 detects the boundary line between the sky and the sea and calculates the slope and intercept of an approximate straight line that approximates the boundary line. Here, the sky and the sea are examples of a first region and a second region. Based on the calculated approximate straight line, the control unit 104 extracts images (frame images) from the set of N consecutive images. The control unit 104 generates digest video data based on the multiple images extracted by repeating this process multiple times. Here, the boundary line is assumed to be the horizontal line, but it may also be the horizon or the boundary of other structures. Details of the digest video generation process will be described later.

[0024] The memory unit 105 stores data such as various parameter settings and operation programs, as well as computer programs. The memory unit 105 has memory such as RAM (Random Access Memory) and ROM (Read Only Memory), and storage such as HDD (Hard Disk Drive) and SSD (Solid State Drive). The memory unit 105 outputs data such as images based on instructions from the control unit 104. The memory unit 105 can temporarily store image signals captured by the imaging unit 101, or image signals that have undergone predetermined signal processing by the control unit 104 and the image processing unit 102, as image data. The memory unit 105 stores subject detection information and boundary detection information performed by the control unit 104, linked to the original image.

[0025] Network 110 may be a WAN (Wide Area Network) and may include routers, switches, cables, and even satellites that satisfy communication standards such as Ethernet (registered trademark). The imaging device 100 is connected to the external device 111 via network 110. Network 110 can be any network that enables communication between the imaging device 100 and the external device 111, regardless of its communication standard, size, configuration, or whether it is wired or wireless, and may even be configured via a cloud.

[0026] External devices 111 include computers and mobile terminals. External devices 111 are connected to the imaging device 100 via the network 110 in a manner that allows them to communicate with each other.

[0027] Next, with reference to Figures 2 to 4, we will explain the method for generating a digest video based on footage taken from the ship. Figure 2 shows the vibration state when the ship 200 is rocking due to vibrations caused by waves and wind. The x and y axes are directions on the horizontal plane that are orthogonal to each other. The z axis is vertical. The bow is facing in the direction of the y axis. The vibrations occurring in the ship 200 correspond to rotations around the x, y, and z axes, which are the pitch direction (the direction in which the bow moves up and down), the roll direction (the direction in which the ship rotates around the bow direction as an axis), and the yaw direction (the direction in which the ship rotates parallel to the horizontal plane), respectively. The vibrations in each direction can be considered independently. Also, the shooting direction of the imaging device 100 is assumed to be horizontal (on the xy plane) due to the nature of monitoring from on the ship 200. Due to the structure of the ship, the vibration in the yaw direction is relatively small, so vibrations that move only in the horizontal direction of the screen do not need to be considered, and therefore, in the following, we will only consider the change in the field of view due to vibrations in the pitch direction and the roll direction.

[0028] Figure 3 shows examples of images obtained when the imaging device 100 photographs a subject 310 (in this case, a ship) floating on the sea 320 with the sky 330 as the background, from the ship itself. Figure 3(a) shows an example of an image in a waveless state. Figure 3(b) shows an example of an image in a wavey state with tilting from side to side. Figure 3(c) shows an example of an image in a wavey state with up-and-down rocking. Figure 3(d) shows an example of an image in a state where waves obscure the subject 310.

[0029] The imaging device 100 photographs the ship as the subject 310 with a fixed field of view, and by continuously photographing the subject 310 over time, it generates multiple images including the subject 310, as shown in Figure 3. The control unit 104 may extract only specific images from these images, or images at fixed intervals, and arrange them again in chronological order to generate a digest video. If both the imaging device 100 and the subject 310 are stationary, the entire subject 310 will be contained within the captured field of view, as in image 340 of Figure 3(a), and an image will always be obtained in which the boundary between the sky 330 and the sea 320 is nearly horizontal and passes through the center of the image. In that case, regardless of which image is selected from the captured images to generate the digest video, almost the same digest video (excluding changes due to the movement of the ship or the subject 310) will be obtained. Therefore, the control unit 104 can generate a digest video by periodically sampling the captured images and stitching the sampled images together.

[0030] However, in reality, vibrations caused by waves and wind occur in the vessel and the subject 310. As a result, as shown in image 350 in Figure 3(b) and image 360 ​​in Figure 3(c), consecutive captured images may include images where the subject 310 is tilted, and images where the subject 310 is shifted upwards and not centered. Furthermore, in the case of rough waves, the captured images may include images where the subject 310 is covered by waves, as shown in image 370 in Figure 3(d). In this case, if the control unit 104 generates a digest video by periodic sampling, it will be difficult to extract images in which the subject 310 is clearly visible. Therefore, the control unit 104 needs to selectively extract images from the continuously captured images in which the subject 310 is centered in the field of view and the boundary line is horizontal, as shown in Figure 3(a), and generate a digest video using these images. If the control unit 104 selects images in which the boundary line is horizontal and is centered on the screen, it will be able to select images in which the subject 310 is centered in the vertical direction of the field of view.

[0031] One example of a method for detecting boundaries is to use far-infrared images of the subject 310. Since far-infrared images can capture the boundary between the sea and the air with high contrast, the boundary can be easily detected by performing image processing such as edge extraction on the obtained image. Alternatively, the control unit 104 may match the field of view of the far-infrared image and the visible image and estimate the boundary in the visible image based on the boundary extracted by the far-infrared image. Furthermore, the control unit 104 can efficiently extract the boundaries between the sky and the sea, the sky and the ground, and the sky and buildings by using a trained model that has been trained on the visible image using machine learning.

[0032] The boundary lines obtained by the above method generally have fine irregularities, so it is desirable to approximate them to a straight line using the least squares method or similar. Figure 4 is a diagram illustrating the calculation of the approximate straight line of the boundary line. Figure 4(a) is a schematic diagram showing the boundary line between the sea and the sky in the image. Figure 4(b) is a schematic diagram showing the approximate straight line L of the boundary line. As shown in image 410 of Figure 4(a), the boundary line 420 is the boundary between the sky 430 and the sea 440.

[0033] The approximate line L can be represented as a line Y = aX + b with parameters of slope a and Y-intercept b (hereinafter referred to as intercept b) by least squares in a Cartesian coordinate system with the center of image 410 as the origin and an X-axis parallel to the horizontal direction and a Y-axis parallel to the vertical direction. If the values ​​of slope a and intercept b are close to 0, or if the absolute values ​​of slope a and intercept b are small, it indicates that the boundary line is more horizontal and that the boundary line is not shifted vertically from the center of the image. Therefore, by extracting an image with appropriate values ​​of slope a and intercept b, the control unit 104 can capture the target subject or boundary line in the center of the screen and generate a highly visible digest video. Appropriate values ​​of slope a and intercept b may be values ​​where at least one of the values ​​of slope a and intercept b is close to 0, or where at least one of the absolute values ​​of slope a and intercept b is small.

[0034] Furthermore, if we extract frame images from the digest video solely by comparing the size of the slope a and intercept b of the approximate line L, there is a possibility of extracting images like Figure 3(d). Therefore, when determining the approximate line L, the control unit 104 may also calculate the variance of the approximate curve L with respect to the original boundary line and exclude images whose variance is greater than a predetermined variance threshold. In addition, to avoid selecting images where either the slope a or intercept b is extremely small, the control unit 104 may select the image where the sum of the slope a and intercept b is closest to 0. The sum here may be the sum of the absolute value of the slope a and the absolute value of the intercept b. Furthermore, the control unit 104 may select images by comparing either the weighted average of the slope a or the weighted average of the intercept b. The weighted average here may be either the weighted average of the absolute value of the slope a or the weighted average of the absolute value of the intercept b. Furthermore, the control unit 104 may select the image where the weighted average of the slope a and intercept b is closest to 0. The weighted average here may be the weighted average of the absolute values ​​of the slope a and the absolute values ​​of the intercept b.

[0035] Furthermore, when generating a digest video, the control unit 104 may generate the digest video using all images extracted from all captured images, or it may select the most appropriate image from images captured at a predetermined time interval T0 and combine the selected images to generate the digest video. The above time interval T0 may be a fixed value or may be variable according to the degree of limitation of the communication bandwidth. In either case, the time interval T0 may be stored in the storage unit 105. In addition, the control unit 104 may superimpose the capture time information of the images used when generating the digest video onto the video.

[0036] Next, the operation of the imaging system in this embodiment will be described with reference to Figure 5. Figure 5 is a flowchart showing the process of generating frame images of the digest video in the first embodiment. Each step in Figure 5 is mainly performed by the control unit 104.

[0037] First, in step S501, the control unit 104 acquires a set of consecutive images captured by the imaging unit 101 and makes them the target of the process for generating a digest video. Here, the set of consecutive images may be a set of multiple images taken between predetermined time intervals T0 from among multiple images or multiple frames of a video stored in the storage unit 105. The time interval T0 is a parameter that determines the degree of summarization of the digest video. The control unit 104 may read out the set of consecutive images again after saving them to the storage unit 105 after capture. Alternatively, the control unit 104 may use multiple images acquired by the imaging unit 101 from a certain time until a time interval T0 has elapsed as the set of consecutive images. Here, the control unit 104 first acquires N images corresponding to the time interval T0 from among the multiple images stored in the storage unit 105 as the set of consecutive images. Then, it initializes M=1 and proceeds to step S502. M is a positive integer from 1 to N for counting.

[0038] In step S502, the control unit 104 selects the Mth image from the sequence of images acquired in step S501 and proceeds to step S503.

[0039] In step S503, the control unit 104 uses the image processing unit 102 to extract boundaries within the image. The method for extracting boundaries may be edge processing or other methods. Once the extraction is complete, the control unit 104 proceeds to step S504. If boundaries cannot be extracted from the image, the control unit 104 proceeds to step S506 without using the image for digest video generation.

[0040] In step S504, the control unit 104 finds the approximate straight line L of the boundary line extracted in step S503 and calculates its slope a and intercept b. The control unit 104 may set up a Cartesian coordinate system with the center of the image as the origin and the horizontal and vertical directions of the image as the X and Y axes, respectively, and calculate the approximate straight line L using the least squares method in this coordinate system. After that, the control unit 104 proceeds to step S505.

[0041] In step S505, the control unit 104 associates the inclination a and intercept b obtained in step S504 with the selected image as parameters indicating the horizontality of the boundary line, and stores them in the storage unit 105. After that, the control unit 104 proceeds to step S506.

[0042] In step S506, if M≠N, that is, if M has not reached N, the control unit 104 proceeds to step S508.

[0043] In step S508, the control unit 104 sets M=M+1 and proceeds to step S502.

[0044] On the other hand, in step S506, if the control unit 104 has completed the boundary line extraction process for all selected images and has determined that M=N, it proceeds to step S507.

[0045] In step S507, the control unit 104 selects the optimal image for the digest video from the N images included in the sequence based on the slope a and intercept b. For example, the control unit 104 may extract the image with the slope a and intercept b closest to 0 as the optimal digest video frame image. The control unit 104 may extract the image with the smallest absolute value of slope a and absolute value of intercept b as the optimal digest video frame image. Alternatively, the control unit 104 may calculate the sum of the absolute values ​​of the slope a and intercept b of the approximation line of the boundary line of each of the N images, and select the image with the sum closest to 0 as the optimal digest video frame image. Furthermore, the control unit 104 may extract the digest video frame images after excluding images where the boundary line, approximation line, and variance value are greater than a predetermined variance threshold.

[0046] The imaging device 100 performs the above operations for the number of frames required for the digest video, and combines multiple frame images into data for a single digest video. In this way, the imaging device 100 calculates an approximate straight line of the boundary between the sky and the sea, extracts frame images based on this approximate straight line, and generates a digest video. Therefore, it is possible to generate a highly visible digest video using a simple method while suppressing the effects of wave vibrations.

[0047] (Second Embodiment) Next, with reference to Figures 6 to 8, the configuration of the imaging system 10 of the second embodiment, which differs from the first embodiment, will be described. Figure 6 is a block diagram showing the overall configuration of the imaging system 10 of the second embodiment. The imaging device 100 of this embodiment has a vibration detection unit 106 in addition to the configuration of the first embodiment.

[0048] The vibration detection unit 106 measures and detects the direction, magnitude, and period of the vibration applied to the imaging device 100 as vibration parameters. The vibration detection unit 106 may be a dedicated device such as a gyro sensor. Also, the vibration detection unit 106 may be an arithmetic processing unit that measures the applied vibration from the change in the image. Further, the vibration detection unit 106 may read the image, execute software, and calculate and detect the vibration. In this case, the vibration detection unit 106 may be realized as a function of the control unit 104. The vibration detection unit 106 stores the detected vibration parameters in the storage unit 105 in association with time information or the video captured simultaneously with the vibration.

[0049] FIG. 7 is a diagram showing a state in which periodic vibration due to waves is applied to the own ship 200. The horizontal axis in FIG. 7 is the time axis. The white arrow in FIG. 7 indicates the direction of the imaging angle of the imaging device 100. The period T1 in FIG. 7 indicates the period of the vibration. That is, the imaging device 100 captures the same direction every period T1. In the figure, only the vibration in the pitch direction is drawn for simplicity, but it can be assumed that the roll direction vibrates in the same manner. It is desirable that the frame images used when generating the digest video are extracted from those captured at the same phase timing due to the periodicity of the waves. Here, consider a process in which the control unit 104 processes N images corresponding to the time interval T0 and extracts one optimal digest video frame image. When the period T1 is longer than the time interval T0 (T0 < T1), the control unit 104 may not be able to extract the frame images captured at the same phase timing from the set of each consecutive image. Therefore, the control unit 104 changes the time interval T0 so that T0 ≥ T1, and includes the images captured at the same phase timing of the waves in the set of consecutive images. Thereby, the control unit 104 can generate a digest video with higher visibility. Although a more appropriate digest video can be generated by setting a longer time interval T0, on the other hand, the frame rate of the digest video will decrease. Therefore, the control unit 104 may set the time interval T0 to be the same as or approximately the same length as the wave period T1.

[0050] Next, referring to FIG. 8, the operation of the imaging system in the second embodiment will be described. FIG. 8 is a diagram of a flowchart showing the generation process of the frame image of the digest video in the second embodiment. Each step in FIG. 8 is mainly executed by the control unit 104.

[0051] In step S801, the control unit 104 reads out the period T1 of the vibration applied to the imaging device 100 detected by the vibration detection unit 106 from the storage unit 105, and compares it with the time interval T0 also stored in the storage unit 105. Here, the period T1 is a vibration parameter associated with the time when the image that is the source of the digest video to be generated is taken. The time interval T0 is a parameter that determines the degree of summarization of the digest video. If T0 < T1, the control unit 104 proceeds to step S802; if T0 ≥ T1, the control unit 104 proceeds to step S803.

[0052] In step S802, the control unit 104 changes the time interval T0 so that T0 = T1, and stores it in the storage unit 105. Then, the control unit 104 proceeds to step S803.

[0053] In step S803, the control unit 104 acquires a set of consecutive images during the time interval T0 from a plurality of images captured by the imaging unit 101 as the target of the process for generating the digest video. The set of consecutive images may be a set of images captured during a predetermined time interval T0 stored in the storage unit 105. The time interval T0 is a parameter that determines the degree of summarization of the digest video. The control unit 104 may read out again the set of consecutive images stored in the storage unit 105 after shooting. Also, the control unit 104 may use, as the set of consecutive images, a plurality of images acquired continuously by the imaging unit 101 until a time interval T0 elapses starting from a certain time. Here, the control unit 104 acquires, as the set of consecutive images, N images corresponding to the time interval T0 among the plurality of images once stored in the storage unit 105. Also, initialize M = 1 and proceed to step S804. M is a positive integer from 1 to N for counting.

[0054] Next, in steps S804 to S810, the control unit 104 performs operations corresponding to steps S502 to S508 in Figure 5.

[0055] The imaging device 100 performs the above operations for the number of frames required for the digest video, and by combining multiple frame images into data for a single digest video, it becomes possible to generate a digest video with high visibility.

[0056] Furthermore, in the second embodiment, the control unit 104 of the imaging device 100 changes the time interval T0 to the period T1 if the period T1 is longer than the time interval T0. As a result, the control unit 104 can extract images of the same phase and generate a digest video from the extracted images, thereby improving the visibility of the digest video.

[0057] (Third embodiment) Next, the configuration of the imaging system 10 of the third embodiment, which differs from the embodiments described above, will be explained. Note that the configuration of the imaging device 100 of the third embodiment is the same as that of the second embodiment, so its explanation will be omitted. The imaging device 100 of the third embodiment has a vibration detection unit 106. The vibration detection unit 106 measures the direction, magnitude, and period of vibration applied to the imaging device 100.

[0058] The control unit 104 can improve the visibility of the digest video by extracting images taken at the same phase timing based on the periodicity of the wave, as frame images to be used when generating the digest video.

[0059] Here, if the control unit 104 extracts only images in which the absolute values ​​of the slope a and intercept b, which are parameters of the approximate line L of the boundary line, are smaller than the slope threshold a0 and intercept threshold b0, respectively, and generates a digest video using all the extracted images, then probabilistically, the longer the wave period T1, the more difficult it becomes to extract frames to be used in the digest video. In other words, a problem arises in which the frame rate of the digest video changes depending on the wave period.

[0060] Therefore, the control unit 104 of this embodiment suppresses changes in the frame rate of the digest video regardless of the wave period T1 by changing the slope threshold a0 and the intercept threshold b0 based on the wave period T1. For example, the control unit 104 can suppress changes in the frame rate of the digest video by setting the slope threshold a0 and the intercept threshold b0 to be larger as the period T1 is longer. In addition, depending on the installation location of the imaging device 100, the vibration periods in the pitch direction and the roll direction may differ. In this case, the control unit 104 may individually change the slope threshold a0 and the intercept threshold b0 based on the direction of the period vibration. Specifically, the control unit 104 may increase the slope threshold a0 as the period T2 in the pitch direction is longer, and increase the intercept threshold b0 as the period T3 in the roll direction is longer.

[0061] Here, we assume that the wave period T1 is nearly constant, but in reality, the period T1 changes. And if the wave period T1 changes frequently, there is a possibility that an image taken at a specific timing may be extracted as a frame image in the digest video. Therefore, the control unit 104 may calculate the average value of the periods of multiple waves over a certain period of time as the value of the wave period T1 and store it in the storage unit 105. In this case, the control unit 104 may further calculate the average value of the wave periods at regular intervals and update the wave period T1.

[0062] Next, with reference to Figure 9, the operation of the imaging system in the third embodiment will be described. Figure 9 is a flowchart showing the process of generating frame images of the digest video in the third embodiment. Each step in Figure 9 is mainly performed by the control unit 104. Here, the control unit 104 processes the images captured by the imaging unit 101 without saving them to the storage unit 105. The control unit 104 executes the flowchart in Figure 9 each time it generates P frame images of the digest video in order to update the value of the wave period T1.

[0063] First, in step S901, the control unit 104 calculates the period T1 of the vibration applied to the imaging device 100 using the vibration detection unit 106. Alternatively, the control unit 104 may acquire the period T1 detected by the vibration detection unit 106. After that, the control unit 104 proceeds to step S902.

[0064] In step S902, the control unit 104 changes the slope threshold a0 and the intercept threshold b0 according to the calculated vibration period T1, and stores the changed values ​​in the storage unit 105. The control unit 104 may also change the slope threshold a0 and the intercept threshold b0 according to the direction of vibration. After that, the control unit 104 proceeds to step S903.

[0065] In step S903, the control unit 104 extracts boundary lines within the image captured by the imaging unit 101. The control unit 104 may also use the image processing unit 102 to extract the boundary lines. The control unit 104 may use edge processing or other methods as the method for extracting boundary lines. If the control unit 104 can extract boundary lines within the image, it proceeds to step S904. If the control unit 104 cannot extract boundary lines within the image, it does not use the image for generating the digest video and repeats step S903.

[0066] In step S904, the control unit 104 finds the approximate line L of the boundary line extracted in step S903 and calculates the slope a and intercept b of the approximate line L. The control unit 104 may calculate the approximate line L of the boundary line using the least squares method. After that, the control unit 104 proceeds to step S905.

[0067] In step S905, the control unit 104 compares the absolute values ​​of the slope a and intercept b of the approximation curve L calculated in step S904 with the slope threshold a0 and intercept threshold b0 stored in the memory unit 105. If the control unit 104 determines that the absolute values ​​of the slope a and intercept b are smaller than the slope threshold a0 and intercept threshold b0, it extracts the selected image as a frame image of the digest video. After that, the control unit 104 proceeds to step S906. On the other hand, if the control unit 104 determines that the absolute values ​​of the slope a and intercept b are greater than or equal to the slope threshold a0 and intercept threshold b0, it returns to step S903.

[0068] In step S906, if the control unit 104 determines that the number of frames in the generated digest video is P or less, it proceeds to step S903. On the other hand, if the control unit 104 determines that the number of frames in the video has reached P, it terminates the process.

[0069] The imaging device 100 performs the above operations for the number of frames required for the digest video, and by combining multiple frame images into data for a single digest video, it becomes possible to generate a digest video with high visibility.

[0070] Furthermore, the control unit 104 of the imaging device 100 in the third embodiment changes the slope threshold a0 and the intercept threshold b0 according to the period T1. As a result, the control unit 104 can suppress changes in the frame rate of the digest video and generate a digest video with high visibility, even when the ship is subjected to vibrations caused by waves.

[0071] (Fourth Embodiment) Next, the configuration of the imaging system 10 of the fourth embodiment, which differs from the embodiments described above, will be explained. In the fourth embodiment, we consider the case where the ship to be the subject is detected by the control unit 104 or the image processing unit 102 of this embodiment. When the ship to be the subject is located near the ship, the speed of the subject moving within the field of view changes depending on the distance. When the speed increases, even if it is desired to capture the subject in detail, the proportion of the subject included in the digest video will decrease. Therefore, when a subject is detected, it may be better to prioritize whether the subject is present in the image over the change in the field of view. In this embodiment, the control unit 104 changes the time interval T0, which is a parameter that determines the degree of summarization of the digest video, based on whether or not a subject has been detected, in order to include more frame images containing the subject in the digest video. Specifically, when the control unit 104 detects a subject, it reduces the time interval T0. As a result, this embodiment can capture the subject in more detail while generating a digest video with high visibility. It should be noted that a similar effect can be obtained by increasing the slope threshold a0 and intercept threshold b0 in the third embodiment.

[0072] Next, with reference to Figure 10, the operation of the imaging system in the fourth embodiment will be described. Figure 10 is a flowchart showing the process of generating frame images of a digest video in the fourth embodiment. Each step in Figure 10 is mainly performed by the control unit 104. Here, as in the first embodiment, the control unit 104 processes a set of N consecutive images contained within a time interval T0 from among multiple frame images of multiple images or videos that have been stored in the storage unit 105. The control unit 104 or the image processing unit 102 may detect a subject in the captured image in advance and store the detection result in the storage unit 105, or it may detect a subject at the same time as generating the digest video.

[0073] In step S1001, the control unit 104 selects a set of consecutive images contained in the time interval T0 from multiple images or multiple frame images of a video captured by the imaging unit 101, and detects a subject from these images. If the control unit 104 detects a subject from any of the consecutive images, it proceeds to step S1002. On the other hand, if the control unit 104 cannot detect a subject, it proceeds to step S1003.

[0074] In step S1002, the control unit 104 modifies the time interval T0 stored in the memory unit 105 to a smaller value, and then stores the modified time interval T0 back in the memory unit 105. After that, the control unit 104 proceeds to step S1003.

[0075] In step S1003, the control unit 104 selects a set of consecutive images acquired in step S1001 for a time interval T0 minutes and makes them the target of the process for generating a digest video. Also, M=1 and proceeds to step S1004.

[0076] In steps S1004 to S1010, the control unit 104 performs the same processing as in steps S502 to S508 in Figure 5.

[0077] By performing the above operations for the number of frames required for the digest video and combining the frame images into a single video data, it becomes possible to generate a digest video while simultaneously capturing the subject in more detail.

[0078] (Other embodiments) In the above-described embodiment, the control unit 104 calculated an approximate straight line as a line approximating the boundary line, but the approximate line is not limited to an approximate straight line. For example, the control unit 104 may calculate a parabola (e.g., a quadratic curve) as the approximate line approximating the boundary line.

[0079] In the above-described embodiment, the control unit 104 performed the digest video generation process, but the main entity performing the execution is not limited to the control unit 104. For example, the image processing unit 102 may perform part or all of the digest video generation process. Alternatively, the control unit 104 and the image processing unit 102 may jointly perform the digest video generation process. In this case, for example, the image processing unit 102 may, based on instructions from the control unit 104, further detect a subject from the image and the boundary line between the sky and the sea, and determine the position of the subject within the screen and the inclination of the boundary line, using the image signal on which predetermined signal processing has been performed. Furthermore, the image processing unit 102 may generate digest video data based on the detected information, etc. Here, the boundary line is assumed to be the horizon, but it may also be the sky horizon or the boundary of other structures. In this case, the control unit 104 and the image processing unit 102 are examples of acquisition means and generation means.

[0080] In the above-described embodiment, the control unit 104 acquired N consecutive images and performed the digest video generation process, but the images to be acquired are not limited to consecutive images. For example, the control unit 104 may acquire N images from the video at intervals of multiple frames (e.g., every two frames) in chronological order and perform the digest video generation process.

[0081] In the above-described embodiment, the control unit 104 extracted images as frame images of the digest video, such as images where the slope and intercept of the approximate line calculated from the boundary line are close to 0, in a Cartesian coordinate system with the center of the image as the origin and the horizontal and vertical directions of the image as the X and Y axes, respectively. However, the method of extracting frame images is not limited to this. For example, the control unit 104 may extract images as frame images of the digest video if the approximate lines are close to each other. Specifically, the control unit 104 may calculate an approximate line from the boundary line extracted from the image and use this approximate line as a reference line. The control unit 104 may then extract images as frame images that have an approximate line of a boundary line with a slope and intercept closest to the slope and intercept of the reference line. After calculating the reference line, the control unit 104 may extract frame images based on the sum or weighted average of the absolute values ​​of the slope and intercept, similar to the above-described embodiment. Alternatively, after extracting the reference line, the control unit 104 may remove some images based on variance before extracting frame images.

[0082] The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. Furthermore, the present invention can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0083] The disclosures herein include the following control devices, imaging devices, control methods, and programs. (Item 1) In a control device that extracts frame images from multiple images to generate a digest video, A means for acquiring multiple images in a time series, A generation means that detects the boundary line between a first region and a second region contained in each of the plurality of images, and extracts frame images from the plurality of images based on the boundary line to generate the digest video, A control device characterized by comprising: (Item 2) The generation means calculates an approximate straight line that approximates the boundary line, and extracts the frame image based on the approximate straight line. The control device according to item 1, characterized in that it is a control device. (Item 3) The generation means extracts the frame image based on at least one of the slope and intercept of the approximate line. The control device according to item 2, characterized in that (Item 4) The generation means extracts the image as the frame image in which at least one of the slope and intercept of the approximate line in a Cartesian coordinate system passing through the center of the image and parallel to the horizontal and vertical directions of the image is closest to 0. The control device according to item 3, characterized in that (Item 5) The generation means extracts the image as the frame image in which at least one of the sum of the absolute value of the slope and the absolute value of the intercept of the approximate line, and the weighted average of the absolute value of the slope and the absolute value of the intercept, is smallest. A control device according to item 3 or item 4, characterized in that it is a control device according to item 4. (Item 6) The generation means extracts the frame image from images excluding those in which the variance value of the approximate line with respect to the boundary line is greater than a predetermined variance threshold. A control device according to any one of items 2 to 5, characterized in that it is a control device. (Item 7) The acquisition means acquires a set of the plurality of images taken at predetermined time intervals. A control device according to any one of items 1 to 6, characterized in that it is a control device. (Item 8) The generation means shortens the time interval when a subject is detected in the image. The control device according to item 7, characterized in that it is a control device. (Item 9) The generating means changes the time interval based on at least one of the direction and period of vibration of the imaging device that captured the plurality of images. A control device according to item 7 or item 8, characterized by the above. (Item 10) The generation means extracts the frame image based on the tilt and a predetermined tilt threshold. A control device according to any one of items 3 to 5, characterized in that it is a control device. (Item 11) The generation means extracts the frame image based on the intercept and a predetermined intercept threshold. A control device according to any one of items 3 to 5, characterized in that it is a control device. (Item 12) The generating means modifies the tilt threshold based on at least one of the direction and period of vibration of the imaging device that captured the plurality of images. The control device according to item 10, characterized in that (Item 13) The generating means modifies the intercept threshold based on at least one of the direction and period of vibration of the imaging device that captured the plurality of images. The control device according to item 11, characterized in that it is a control device. (Item 14) The generation means changes the tilt threshold based on whether or not a subject is detected in the image. A control device according to item 10 or item 12, characterized in that it is a control device according to item 10 or item 12. (Item 15) The generation means changes the intercept threshold based on whether or not a subject is detected in the image. A control device according to item 11 or item 13, characterized in that it is a control device according to item 11 or item 13. (Item 16) The control device described in item 1, A vibration detection means for detecting vibrations, An imaging means for photographing a subject and generating an image, An imaging device characterized by comprising: (Item 17) In a control method for extracting frame images from multiple images to generate a digest video, Acquire multiple images in a time series, The boundary line between the first region and the second region contained in each of the multiple images is detected, and based on the boundary line, frame images are extracted from the multiple images to generate the digest video. A control method characterized by the following: (Item 18) A program for causing a computer to function as one of the control devices described in any one of items 1 through 15.

[0084] The invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention. [Explanation of symbols]

[0085] 10...Imaging system, 100...Imaging device, 111...External device, 101...Imaging unit, 104...Control unit, 420...Boundary line, 106...Vibration detection unit, Approximate straight line...L.

Claims

1. In a control device that extracts frame images from multiple images to generate a digest video, A means for acquiring multiple images in a time series, A generation means that detects the boundary line between a first region and a second region contained in each of the plurality of images, and extracts frame images from the plurality of images based on the boundary line to generate the digest video, A control device characterized by comprising:

2. The generation means calculates an approximate straight line that approximates the boundary line, and extracts the frame image based on the approximate straight line. The control device according to feature 1.

3. The generation means extracts the frame image based on at least one of the slope and intercept of the approximate line. The control device according to claim 2.

4. The generation means extracts the image as the frame image in which at least one of the slope and intercept of the approximate line in a Cartesian coordinate system passing through the center of the image and parallel to the horizontal and vertical directions of the image is closest to zero. The control device according to claim 3.

5. The generation means extracts the image as the frame image in which at least one of the sum of the absolute value of the slope and the absolute value of the intercept of the approximate line, and the weighted average of the absolute value of the slope and the absolute value of the intercept, is smallest. The control device according to claim 3.

6. The generation means extracts the frame image from images excluding those in which the variance value of the approximate line with respect to the boundary line is greater than a predetermined variance threshold. The control device according to claim 2.

7. The acquisition means acquires a set of the plurality of images taken at predetermined time intervals. The control device according to feature 1.

8. The generation means shortens the time interval when a subject is detected in the image. The control device according to feature 7.

9. The generating means changes the time interval based on at least one of the direction and period of vibration of the imaging device that captured the plurality of images. The control device according to feature 7.

10. The generation means extracts the frame image based on the tilt and a predetermined tilt threshold. The control device according to claim 3.

11. The generation means extracts the frame image based on the intercept and a predetermined intercept threshold. The control device according to claim 3.

12. The generating means modifies the tilt threshold based on at least one of the direction and period of vibration of the imaging device that captured the plurality of images. The control device according to claim 10.

13. The generating means modifies the intercept threshold based on at least one of the direction and period of vibration of the imaging device that captured the plurality of images. The control device according to feature 11.

14. The generation means changes the tilt threshold based on whether or not a subject is detected in the image. The control device according to claim 10.

15. The generation means changes the intercept threshold based on whether or not a subject is detected in the image. The control device according to feature 11.

16. The control device according to claim 1, A vibration detection means for detecting vibrations, An imaging means for photographing a subject and generating an image, An imaging device characterized by comprising:

17. In a control method for extracting frame images from multiple images to generate a digest video, Acquire multiple images in a time series, The boundary line between the first region and the second region contained in each of the plurality of images is detected, and based on the boundary line, frame images are extracted from the plurality of images to generate the digest video. A control method characterized by the following:

18. A program for causing a computer to function as one of the means of the control device described in any one of claims 1 to 15.

Citation Information

Patent Citations

  • Method and system for generating adaptive fast forward of egocentric videos

    US9672626B2