Image capturing apparatus, control method, and storage medium
By capturing and utilizing non-recorded images alongside recorded images, the imaging device enhances the accuracy of image caption generation through a neural network, addressing the limitations of existing technologies.
Patent Information
- Application Number
- JP2024120333
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2026-02-05
AI Technical Summary
Existing image caption generation technologies struggle to accurately describe the content of images due to limited information from a single image.
An imaging device captures both recorded and non-recorded images, using a neural network to generate captions based on the recorded images and one or more non-recorded images that satisfy specific conditions, such as being captured within a certain time frame or including a priority subject.
Improves the accuracy of generating linguistic information that describes the content of images by leveraging additional contextual information from non-recorded images.
Smart Images

Figure 2026018963000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an imaging device, a control method, and a program. [Background technology]
[0002] There is a known technology for generating image summaries (captions) using neural networks. Patent Document 1 discloses a technology that extracts overall and partial features from an image, identifies an area of interest from the two features, and assigns weights to the area of interest to improve the accuracy of caption generation. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2020-13427 Summary of the Invention [Problem to be solved by the invention]
[0004] Since the information that can be obtained from a single image is limited, depending on the content of the image, the technology of Patent Document 1 may not necessarily be able to generate a caption that explains the content of the image with high accuracy.
[0005] The present invention has been made in view of the above circumstances, and aims to provide a technique for improving the accuracy of generating linguistic information that describes the content of an image. [Means for solving the problem]
[0006] In order to solve the above problem, the present invention provides an imaging device comprising: an imaging means for capturing at least one first image that is not recorded in non-volatile storage and a second image that is recorded in the non-volatile storage; and a generation means for generating linguistic information that describes the content of the second image based on the second image and one or more first images that satisfy one or more conditions among the at least one first image. [Effects of the Invention]
[0007] According to the present invention, it is possible to improve the accuracy of generating linguistic information that describes the content of an image.
[0008] Other features and advantages of the present invention will become more apparent from the accompanying drawings and the following detailed description of the preferred embodiment of the present invention. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram showing an example of the hardware configuration of an imaging device 100. [Figure 2] FIG. 2 is a diagram showing the configuration of a function for generating captions for recorded images in the imaging device 100. [Figure 3] FIG. 3 is a conceptual diagram showing the relationship between a recorded image and one or more non-recorded images used to generate a caption according to the first embodiment. [Figure 4A] 5 is a flowchart of a photographing process executed by a CPU 102 according to the first embodiment. [Figure 4B] 4 is a flowchart of a caption generation process executed by an input control unit 201 and a caption generation unit 202 according to the first embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing the relationship between a recorded image and one or more non-recorded images used to generate a caption according to the second embodiment. [Figure 6A] 10 is a flowchart of a photographing process executed by a CPU 102 according to a second embodiment. [Figure 6B]10 is a flowchart of a caption generation process executed by an input control unit 201 and a caption generation unit 202 according to the second and third embodiments. [Figure 7] FIG. 11 is a conceptual diagram showing the relationship between a recorded image and one or more non-recorded images used to generate a caption according to a third embodiment. [Figure 8] 10 is a flowchart of a caption generation process executed by an input control unit 201 and a caption generation unit 202 according to the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.
[0011] [First embodiment] 1 is a diagram showing an example of the hardware configuration of an imaging device 100. In the imaging device 100, a CPU 102, a ROM 103, a memory 104, an interface unit 105, a display unit 106, an imaging unit 107, and a storage 108 are connected to a system bus 101. The units connected to the system bus 101 are configured to be able to exchange data with each other via the system bus 101.
[0012] The ROM 103 stores various programs for the operation of the CPU 102. The storage location of the various programs for the operation of the CPU 102 is not limited to the ROM 103, but may be, for example, a hard disk or the like.
[0013] The memory 104 is a volatile memory, and is configured by, for example, a RAM. The CPU 102 operates according to a program stored in the ROM 103, and uses the memory 104 as a work memory.
[0014] The interface unit 105 receives a user operation, generates a control signal corresponding to the operation, and supplies the control signal to the CPU 102. For example, the interface unit 105 has physical operation buttons, a touch panel, or the like as an input device for receiving the user operation. The touch panel is an input device configured to output coordinate information corresponding to a position touched on an input unit configured, for example, as a plane.
[0015] The CPU 102 controls each unit including the display unit 106 and the imaging unit 107 according to a program based on a control signal supplied in response to a user operation via the interface unit 105. This allows the CPU 102 to cause the display unit 106 and the imaging unit 107 to perform operations in response to the user operation.
[0016] The display unit 106 includes, for example, a display. The display unit 106 includes a mechanism for outputting a display signal for displaying an image on the display. If the interface unit 105 includes a touch panel, the touch panel and the display can be configured as an integrated unit. For example, the touch panel is configured so that the light transmittance does not interfere with the display on the display, and is attached to the upper layer of the display surface of the display. Then, by associating input coordinates on the touch panel with display coordinates on the display, a touch panel that functions as the interface unit 105 can be configured.
[0017] The imaging unit 107 includes a lens, a shutter with an aperture function, and an imaging element (CCD, CMOS, etc.) that converts an optical image into an electrical signal. The imaging unit 107 also includes an image processing unit that performs various image processes such as exposure control and distance measurement control based on the signal from the imaging element, and is configured to perform a series of imaging processes. The imaging unit 107 can perform imaging in accordance with user operations via the interface unit 105 under the control of the CPU 102.
[0018] The storage 108 is a non-volatile storage, and is configured, for example, by a memory card. The memory card may be detachable from the imaging device 100.
[0019] The imaging device 100 can capture (acquire) images for recording (hereinafter also referred to as "recorded images") and images for non-recording (hereinafter also referred to as "non-recorded images") using the imaging unit 107. Recorded images are images acquired in response to user instructions obtained via the interface unit 105, for example, and are recorded (saved) in non-volatile storage 108. Recorded images may also be temporarily recorded (saved) in volatile memory 104 before being recorded in storage 108. Non-recorded images are images that are temporarily required for reasons such as displaying on the display unit 106 and using in calculating shooting parameters, and include, for example, live view images (LV images). Non-recorded images are temporarily recorded in the volatile memory 104, but are not recorded in the non-volatile storage 108.
[0020] 2 is a diagram showing the configuration of a function for generating captions for recorded images in the imaging device 100. As shown in FIG.
[0021] The input control unit 201 acquires recorded images captured by the imaging unit 107 from the storage 108 (or memory 104) and inputs them to the caption generation unit 202. The input control unit 201 also acquires non-recorded images captured by the imaging unit 107 from the memory 104 and inputs them to the caption generation unit 202. The functions of the input control unit 201 are realized by the CPU 102 executing a program.
[0022] The caption generation unit 202 generates language information that explains the content of the recorded image based on the recorded image and one or more non-recorded images input from the input control unit 201. In this embodiment, it is assumed that a so-called caption is generated as the language information that explains the content of the recorded image.
[0023] The method for generating captions is not particularly limited, and any method can be used as long as it is based on a recorded image and one or more non-recorded images input from the input control unit 201. For example, the caption generation unit 202 can generate captions by inference processing using a neural network or by rule-based inference processing.
[0024] In the description of this embodiment, the caption generation unit 202 generates captions by inference processing using a neural network. A learning model is stored in advance in the ROM 103. This learning model is a machine learning model that receives a recorded image and one or more non-recorded images as input and is trained using the captions of the corresponding recorded images as training data. The caption generation unit 202 acquires the learning model from the ROM 103 and infers (generates) captions for the recorded images by inputting the recorded image and one or more non-recorded images input from the input control unit 201 into the learning model.
[0025] The functions of the caption generation unit 202 are realized by the CPU 102 executing a program. Alternatively, the imaging device 100 may include a graphics processing unit (GPU), and the functions of the caption generation unit 202 may be realized by the CPU 102 and the GPU working together to execute processing in accordance with the program.
[0026] The specific processing contents of the input control unit 201 and the caption generation unit 202 will be described later with reference to FIG. 4B.
[0027] 3 is a conceptual diagram showing the relationship between a recorded image and one or more non-recorded images used to generate a caption according to the first embodiment. In FIG. 3, time passes from left to right. The imaging device 100 captures multiple non-recorded images to be displayed as LV images on the display unit 106, and one recorded image corresponding to a user's shooting instruction.
[0028] In conventional techniques, captions for recorded images are generated based on the recorded images. However, in the example of Figure 3, it is not easy to determine which of the two people in the recorded image 301 is trying to extinguish the candle, so captions cannot be generated with high accuracy.
[0029] In this embodiment, in addition to the recorded image 301, one or more non-recorded images captured within a predetermined period of time including the capture of the recorded image are used to generate a caption. In the example of FIG. 3, two non-recorded images (non-recorded images 302 and 303) captured before and after the recorded image are used as the one or more non-recorded images. Of the two people included in the recorded image 301, only the left person is captured in the non-recorded image 302. The left person blowing on the candle is more clearly visible in the non-recorded image 303 than in the recorded image 301. Therefore, it can be determined that the left person is a more important subject in the scene of the recorded image 301, and it becomes possible to generate a highly accurate caption that emphasizes the left person.
[0030] 4A is a flowchart of the photographing process executed by the CPU 102 according to the first embodiment. The CPU 102 executes the process of this flowchart in accordance with a program stored in the ROM 103. When the operation mode of the imaging device 100 is set to the photographing mode by a user operation via the interface unit 105, the CPU 102 starts the process of this flowchart.
[0031] In S401, the CPU 102 causes the imaging unit 107 to capture an LV image.
[0032] In S402, the CPU 102 stores (records) the LV image captured in S401 in the memory 104. Furthermore, the CPU 102 may erase old LV images (LV images that are unlikely to be used to generate captions) stored in the memory 104 as necessary (for example, when the remaining capacity of the memory 104 is low).
[0033] In S403, the CPU 102 determines whether or not a shooting instruction has been input from the interface unit 105. If a shooting instruction has been input, the process proceeds to S404. If a shooting instruction has not been input, the process returns to S401. Therefore, LV images are repeatedly shot until a shooting instruction is input.
[0034] In S404, the CPU 102 causes the imaging unit 107 to capture an image to be recorded.
[0035] In S405, the CPU 102 stores the recording image captured in S404 in the storage 108. After that, the process returns to S401. Therefore, after capturing the recording image, LV images are repeatedly captured until a capture instruction is input again.
[0036] Fig. 4B is a flowchart of the caption generation process according to the first embodiment, which is executed by the input control unit 201 and the caption generation unit 202. The caption generation process in Fig. 4B is executed in parallel with the shooting process in Fig. 4A.
[0037] In S451, the input control unit 201 determines whether or not the recorded image has been stored in the storage 108. The input control unit 201 repeats the determination in S451 until the recorded image has been stored in the storage 108. Once the recorded image has been stored in the storage 108 (i.e., once the recorded image has been stored in S405 of FIG. 4A), the process proceeds to S452.
[0038] In S452, the input control unit 201 waits for a predetermined time. During the wait in S452, the shooting process of Fig. 4A is executed in parallel, so that LV images are repeatedly shot and stored in the memory 104. Note that if an LV image shot after the recorded image is not used to generate a caption (if the first period described below does not include a period after the shooting of the recorded image), the process of S452 is not necessary.
[0039] In S453, the input control unit 201 acquires a recorded image (for example, recorded image 301 shown in FIG. 3) from the storage 108 and inputs it to the caption generation unit 202. Furthermore, the input control unit 201 acquires from the memory 104 one or more LV images (for example, non-recorded images 302 and 303 shown in FIG. 3) that were captured within a first period including the time when the recorded image was captured, out of at least one LV image stored in the memory 104, and inputs them to the caption generation unit 202. Examples of the "first period including the time when the recorded image was captured" here include the period from 0.05 seconds before to 0.05 seconds after the capture of the recorded image, and the period from 0.05 seconds before to the time when the recorded image was captured.
[0040] In S454, the caption generation unit 202 generates a caption for the recorded image based on the recorded image and one or more non-recorded images input from the input control unit 201. As described above, the caption generation unit 202 can infer a caption by inputting the recorded image and one or more non-recorded images input from the input control unit 201 into a learning model. Thereafter, the process returns to S451. Therefore, every time a new recorded image is stored in the storage 108, a corresponding caption is generated.
[0041] In the above description, it is assumed that the recorded images are still images. However, the recorded images may be moving images. When the recorded images are moving images, the recorded images are a group of recorded still images (group of frames), and LV images acquired before the start of recording the moving image and after the end of recording are non-recorded images. Therefore, the input control unit 201 inputs the recorded images, which are moving images, and one or more LV images to the caption generation unit 202. The caption generation unit 202 generates captions for the recorded images based on the recorded images, which are moving images, and one or more LV images. In this case, the caption generation unit 202 may generate one caption for the entire moving image, or may generate a caption for each frame of the moving image.
[0042] As described above, according to the first embodiment, the imaging device 100 captures at least one non-recorded image (e.g., an LV image) that is not recorded in the non-volatile storage 108, and a recorded image that is recorded in the non-volatile storage 108. Then, the imaging device 100 generates linguistic information (e.g., a caption) that describes the content of the recorded image, based on one or more non-recorded images that satisfy one or more conditions among the recorded image and the at least one non-recorded image.
[0043] As described above, according to the first embodiment, linguistic information describing the content of a recorded image is generated based not only on the recorded image but also on one or more non-recorded images that satisfy one or more conditions. Therefore, according to this embodiment, it is possible to improve the accuracy of generating linguistic information describing the content of a recorded image.
[0044] The "one or more conditions" referred to here serve as criteria for selecting one or more non-recorded images to be used in generating linguistic information. The content of the "one or more conditions" is not particularly limited, but using conditions that increase the likelihood that non-recorded images highly related to the content of the recorded image will be used can be expected to further improve the accuracy of generating linguistic information. In the example described with reference to FIG. 4B , a condition is used in which one or more non-recorded images are captured within a first period including the time of capturing the recorded image (e.g., within a period from 0.05 seconds before to 0.05 seconds after capturing the recorded image, or within a period from 0.05 seconds before to the time of capturing the recorded image). Non-recorded images that satisfy such conditions are expected to have a relatively high degree of relatedness to the content of the recorded image, and therefore can be expected to further improve the accuracy of generating linguistic information.
[0045] [Second embodiment] In the first embodiment, a condition that one or more non-recorded images are captured within a first period including the time when the recorded images are captured (hereinafter, sometimes referred to as "condition 1") was described as an example of the "one or more conditions" that serve as selection criteria for one or more non-recorded images to be used in generating captions. In the second embodiment, other examples of the "one or more conditions" will be described. In the second embodiment, the basic configuration of the imaging device 100 is the same as in the first embodiment. Below, differences from the first embodiment will be mainly described.
[0046] 5 is a conceptual diagram showing the relationship between a recording image and one or more non-recording images used to generate a caption according to the second embodiment. In FIG. 5, time passes from left to right. The imaging device 100 captures multiple non-recording images to be displayed as LV images on the display unit 106, and one recording image corresponding to a user's shooting instruction.
[0047] In this embodiment, one or more non-recorded images to be used for generating a caption are selected so as to satisfy both the condition that each of the one or more non-recorded images includes a priority subject (a predetermined subject) (hereinafter sometimes referred to as "condition 2") and the condition that the one or more non-recorded images were captured within a second period (hereinafter sometimes referred to as "condition 3") The second period is a period that includes the capture of the recorded image and in which the priority subject is continuously detected.
[0048] A priority subject is a subject that is given priority consideration when generating a caption. The method for selecting a priority subject is not particularly limited. For example, a user can select a priority subject in advance from recorded images that have been previously captured and stored in storage 108. In this case, the user operates interface unit 105 to select a desired recorded image, and then selects a desired subject from the selected recorded image as a priority subject. CPU 102 stores priority subject information representing the priority subject selected by the user in ROM 103. The specific method for detecting a priority subject from an LV image is not particularly limited. For example, a method based on any known technology, such as pattern matching, can be used.
[0049] In the example of Figure 5, it is assumed that the person on the left in the previously captured recorded image 301 (Figure 3) has been selected in advance as the priority subject. In this case, the priority subject is detected in the non-recorded images 503 to 506. Therefore, the non-recorded images 503 to 506 satisfy condition 2. Note that although the non-recorded image 507 actually includes the priority subject, detection of the priority subject has failed due to low brightness, and it is determined that the non-recorded image does not include the priority subject.
[0050] Furthermore, the period during which non-recorded images 503 to 506 were captured includes the time when recorded image 501 was captured, and the priority subject was continuously detected during this period. Therefore, non-recorded images 503 to 506 satisfy condition 3. Even if a non-recorded image captured before non-recorded image 502 contains the priority subject, this non-recorded image does not satisfy condition 3 because non-recorded image 502, in which the priority subject is not detected, exists between this non-recorded image and non-recorded image 503.
[0051] In this way, in the example of FIG. 5, non-recorded images 503 to 506 that satisfy "one or more conditions" including conditions 2 and 3 are used to generate captions.
[0052] 5, consider a case where non-recorded images 504 and 505 satisfy condition 1 described in the first embodiment. In this case, because the change between recorded image 501 and non-recorded images 504 and 505 is small, using non-recorded images 504 and 505 may not significantly improve the accuracy of the caption. In contrast, in the second embodiment, by using "one or more conditions" including conditions 2 and 3, non-recorded image 506, which has a relatively large change from recorded image 501, is used, which is expected to improve the accuracy of the caption.
[0053] It is not essential to use both Condition 2 and Condition 3. For example, a configuration may be adopted in which one or more LV images that satisfy Condition 2 are selected as one or more LV images used to generate a caption.
[0054] 6A is a flowchart of the photographing process executed by the CPU 102 according to the second embodiment. The CPU 102 executes the process of this flowchart in accordance with a program stored in the ROM 103. When the operation mode of the imaging device 100 is set to the photographing mode by a user operation via the interface unit 105, the CPU 102 starts the process of this flowchart.
[0055] In S601, the CPU 102 determines whether or not the priority subject is included in the LV image captured in S401.
[0056] In S602, the CPU 102 associates the result of the priority subject determination performed in S601 (information indicating whether or not the priority subject is included in the LV image) with the LV image.
[0057] Fig. 6B is a flowchart of a caption generation process according to the second embodiment, which is executed by the input control unit 201 and the caption generation unit 202. The caption generation process in Fig. 6B is executed in parallel with the shooting process in Fig. 6A.
[0058] In S653, the input control unit 201 acquires a recorded image (for example, recorded image 501 shown in FIG. 5) from the storage 108 and inputs it to the caption generation unit 202. Furthermore, the input control unit 201 acquires from the memory 104 one or more LV images (that is, one or more LV images that satisfy conditions 2 and 3) (for example, non-recorded images 502 to 506 shown in FIG. 5) that continuously include the priority subject before and after capturing the recorded image, out of at least one LV image stored in the memory 104, and inputs them to the caption generation unit 202. The input control unit 201 can identify one or more LV images that satisfy conditions 2 and 3 based on the determination results associated with each LV image in S602 of FIG. 6A.
[0059] If there is no LV image that satisfies the conditions 2 and 3, the input control unit 201 may input one or more LV images that satisfy the condition 1 to the caption generation unit 202, as in the first embodiment.
[0060] Furthermore, if an LV image that does not include a priority subject is captured while the input control unit 201 is waiting in S452, the input control unit 201 may end the waiting state and proceed to S651. This is because if an LV image that does not include a priority subject is captured, LV images captured thereafter will not satisfy both Condition 2 and Condition 3.
[0061] Note that, above, conditions 2 and 3 have been described as examples of "one or more conditions" that serve as criteria for selecting one or more non-recorded images to be used to generate captions, but other conditions may also be used.
[0062] For example, the imaging device 100 may include a gaze sensor (not shown), and the CPU 102 may calculate the user's gaze level from information from the gaze sensor. The gaze level here is a numerical value calculated from the user's gaze information and indicates the degree to which the user directed their gaze at each subject. For example, in S601 of FIG. 6A, the CPU 102 acquires the user's gaze information from a gaze sensor provided on the display unit 106 or the like, performs segmentation processing and recognition processing of people, objects, etc. on the LV image, and identifies the subjects appearing in the LV image. Then, the CPU 102 calculates the gaze level as the time the user directed their gaze at each subject using the acquired gaze information, and determines whether the gaze level is equal to or greater than a first threshold. In S602, the CPU 102 associates the gaze level determination result with the LV image. 6B, the input control unit 201 selects one or more LV images to be used for generating a caption so as to satisfy a condition that the gaze degree of each of the one or more LV images is equal to or greater than a first threshold (hereinafter, this condition may be referred to as "condition 4"). This makes it possible to predict the start of shooting based on the user's line of sight even before a shooting instruction is input, and improves the accuracy of caption generation while limiting the number of frames of LV images to be used.
[0063] Here, the time the user directed their gaze was used as the degree of attention, but it is also possible to set coefficients in advance for segmentation attributes such as people and animals, and use the product of the time the user directed their gaze and the coefficient as the degree of attention.
[0064] As another example, a condition may be used that considers whether the imaging device 100 has transitioned to a recording image capture preparation state. Specifically, the imaging device 100 includes a capture button (not shown) as an operating member. The CPU 102 transitions the imaging device 100 to the recording image capture preparation state in response to a predetermined user operation on the capture button. If the capture button has a half-pressed state and a full-pressed state, a half-pressed state corresponds to the predetermined user operation, and a full-pressed state corresponds to a capture instruction. In S601 of FIG. 6A, the CPU 102 determines whether the imaging device 100 is in the recording image capture preparation state. In S602, the CPU 102 associates the determination result of the recording image capture preparation state with the LV image. In S653 of FIG. 6B, the input control unit 201 selects one or more LV images to be used for generating a caption so as to satisfy a condition that one or more LV images have been captured within a third period (hereinafter, also referred to as "Condition 5"). The third period is the period following the transition to the most recent recording image capture preparation state before the recording image was captured. In the example of FIG. 5, if the imaging device was in a shooting preparation state when non-recorded images 503 to 506 were captured, non-recorded images 503 to 506 would be selected as one or more LV images that satisfy the fifth condition. Alternatively, the third period may be the period from the transition to the most recent shooting preparation state before capturing the recorded images until the capturing of the recorded images. In this case, even if the imaging device was in a shooting preparation state when non-recorded images 503 to 506 were captured in the example of FIG. 5, non-recorded images 505 and 506 would not satisfy the fifth condition, and non-recorded images 503 and 504 would be selected as one or more LV images that satisfy the fifth condition. As with the case of using condition 4 described above, when condition 5 is used, it is possible to predict the start of shooting before a shooting instruction is input, thereby improving the accuracy of caption generation while limiting the number of LV image frames used.
[0065] Note that condition 1 described in the first embodiment and conditions 2 to 5 described in the second embodiment can be combined as appropriate as long as there is no technical contradiction. As an example, it is possible to adopt a configuration in which one or more LV images that satisfy "one or more conditions" including condition 1 and condition 5 are selected as one or more LV images to be used for generating a caption.
[0066] In this way, by appropriately using various conditions as the "one or more conditions" that serve as selection criteria for one or more non-recorded images to be used in generating captions, it is possible to improve the accuracy of generating linguistic information that describes the content of the recorded images.
[0067] [Third embodiment] In the second embodiment, a configuration was described in which one or more LV images that satisfy "one or more conditions" are used to generate a caption. In the third embodiment, a configuration is described in which, when multiple LV images satisfy "one or more conditions," some of the multiple LV images that satisfy the one or more conditions are thinned out, and the remaining one or more LV images are used to generate a caption. In the third embodiment, the basic configuration of the imaging device 100 is the same as in the second embodiment. Below, differences from the second embodiment will mainly be described.
[0068] Fig. 7 is a conceptual diagram showing the relationship between a recorded image and one or more non-recorded images used to generate a caption according to the third embodiment. Fig. 7 is almost the same as Fig. 5 described in the second embodiment, but differs from Fig. 5 in that non-recorded images 504 and 505 are not used to generate a caption.
[0069] As described above, generating captions for recorded images based on non-recorded images in addition to recorded images can improve the accuracy of caption generation. However, if the change between images is small due to factors such as the slow speed of the subject, the amount of additional information obtained from each non-recorded image is small. In the example of FIG. 7, non-recorded images 504 and 505, which are located before and after recorded image 501, are only slightly different from recorded image 501, so the amount of additional information obtained from non-recorded images 504 and 505 (information that cannot be obtained from recorded image 501 alone) is small. In such a case, using non-recorded images 504 and 505 cannot be expected to significantly improve the accuracy of caption generation, and instead unnecessarily increases the processing load.
[0070] Therefore, in the third embodiment, a process is performed in which some of the LV images that satisfy one or more conditions are thinned out (in the example of FIG. 7, non-recorded images 504 and 505 are thinned out from non-recorded images 503 to 506). This makes it possible to improve the accuracy of caption generation while suppressing an unnecessary increase in processing load.
[0071] Fig. 8 is a flowchart of a caption generation process executed by the input control unit 201 and the caption generation unit 202 according to the third embodiment. The caption generation process in Fig. 8 is executed in parallel with the shooting process in Fig. 6A. That is, the shooting process according to the third embodiment is the same as that in the second embodiment.
[0072] In S851, the input control unit 201 detects (calculates) the magnitude of change between two or more of the multiple LV images and recorded images that satisfy one or more conditions. In the following description, the one or more conditions include condition 2 and condition 3 described in the second embodiment, and it is assumed that the person on the left in the previously captured recorded image 301 (FIG. 3) has been selected in advance as the priority subject. Therefore, in the example of FIG. 7, non-recorded images 503 to 506 correspond to "multiple LV images that satisfy one or more conditions."
[0073] The "magnitude of change" detected (calculated) in S851 is not particularly limited as long as it is an index of the possibility that a LV image that is unlikely to contribute to improving the accuracy of caption generation is included among the multiple LV images that satisfy one or more conditions. Here, the input control unit 201 calculates the velocity of the priority subject as the "magnitude of change." The velocity calculated here is, for example, the velocity of the priority subject at the time of capturing the recorded image. In this case, the input control unit 201 can use the recorded image and the LV image captured immediately before the recorded image (in the example of FIG. 7, the recorded image 501 and the non-recorded image 504) as "two or more images among the multiple LV images and the recorded images that satisfy one or more conditions." Alternatively, the velocity calculated here may be the average velocity of the priority subject over the entire period during which the multiple LV images that satisfy one or more conditions are captured. In this case, the input control unit 201 can use the recorded image 501 and the non-recorded images 503 to 506 as "two or more images among the multiple LV images and the recorded images that satisfy one or more conditions." The velocity of the priority subject can be calculated, for example, by detecting the motion vector of the priority subject between images.
[0074] If there is only one LV image that satisfies one or more conditions, the input control unit 201 skips S851 and S852 and proceeds to the process from S452 to S855.
[0075] In S852, the input control unit 201 determines whether the change detected in S851 (here, the speed of the priority subject) is smaller than the second threshold value. If the detected change is smaller than the second threshold value, the process proceeds to S853; if not, the process proceeds to S855.
[0076] In S853, the input control unit 201 thins out some of the multiple LV images (non-recorded images 503 to 506) that satisfy one or more conditions. The thinning method is not particularly limited, but for example, the input control unit 201 may simply thin out the LV images at regular intervals, or may thin out the LV images based on the magnitude of the pixel difference between the recorded image and each LV image.
[0077] When simply thinning out LV images at a fixed interval, for example, the input control unit 201 thins out LV images at a rate of one out of every two images.
[0078] When LV images are thinned out based on the magnitude of the pixel difference between the recorded image and each LV image, the input control unit 201 detects the pixel difference between the recorded image and each LV image. Then, when the pixel difference is smaller than a predetermined difference (third threshold), the input control unit 201 thins out the corresponding LV image. In this case, in the example of FIG. 7, non-recorded images 504 and 505 are thinned out.
[0079] As another example, the input control unit 201 may adjust the amount of LV image thinning depending on the magnitude of the change calculated in S851 (the speed of the priority subject). More specifically, the input control unit 201 may increase the number of "portions of multiple LV images" to be thinned out as the change calculated in S851 decreases. For example, the input control unit 201 may thin out one LV image out of every two when the speed is equal to or greater than a predetermined speed, and may thin out two LV images out of every three when the speed is less than the predetermined speed. Note that the "predetermined speed" used here is a speed slower than the "second threshold" used in S852. The speed used here is the average speed of the priority subject over the entire period during which multiple LV images that satisfy one or more conditions are captured.
[0080] Furthermore, the input control unit 201 may change the range in which LV images are thinned out depending on the degree of the speed calculated in S851. For example, the input control unit 201 may thin out one LV image before and one LV image after the recorded image when the speed exceeds a first speed, thin out two LV images before and two LV images after the recorded image when the speed does not exceed the first speed but exceeds a second speed, and thin out three LV images before and three LV images before and three LV images before the recorded image when the speed does not exceed the second speed.
[0081] In S854, the input control unit 201 inputs the recorded image and one or more remaining LV images (one or more LV images that were not thinned out in S853 among the multiple LV images that satisfy one or more conditions) to the caption generation unit 202.
[0082] 8, in S852, a process is performed to determine whether the change detected in S851 (for example, the speed of the priority subject) is smaller than a second threshold. However, S852 can be omitted. In this case, in S853 after S851, the input control unit 201 can appropriately thin out some of the multiple LV images according to the change detected in S851 (for example, the speed of the priority subject).
[0083] As described above, according to the third embodiment, the imaging device 100 detects the magnitude of change between two or more of the multiple LV images and the recorded image when the multiple LV images satisfy one or more conditions. If the change is smaller than a second threshold, the imaging device 100 generates a caption for the recorded image based on the recorded image and one or more LV images remaining after thinning (excluding) some of the multiple LV images. Therefore, according to this embodiment, it is possible to improve the accuracy of caption generation while suppressing an unnecessary increase in processing load.
[0084] In addition, in S851 of FIG. 8, the input control unit 201 may detect the magnitude of change (e.g., pixel difference) between each of the multiple LV images that satisfy one or more conditions and the recorded image. In this case, the input control unit 201 may proceed from S851 to S853 and thin out some of the multiple LV images based on the magnitude of change between each LV image and the recorded image. For example, the magnitude of change in each of the portion of the multiple LV images thinned out here is smaller than a third threshold. In other words, the input control unit 201 may thin out LV images whose change (e.g., pixel difference) from the recorded image is smaller than the third threshold.
[0085] [Other embodiments] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0086] [summary] The above-described embodiment discloses at least the inventions shown in the following items, but is not limited to these inventions. (Item 1) an image capturing means for capturing at least one first image that is not recorded in a non-volatile storage and a second image that is recorded in the non-volatile storage; a generating means for generating linguistic information that describes the content of the second image based on one or more first images that satisfy one or more conditions among the second image and the at least one first image; An imaging device comprising: (Item 2) the one or more conditions include a condition that the one or more first images are captured within a first period that includes the time when the second image is captured; 2. The imaging device according to item 1, (Item 3) a first determination means for determining whether or not a predetermined subject is included in each of the at least one first image; the one or more conditions include a condition that each of the one or more first images includes the predetermined subject; 3. The imaging device according to item 1 or 2, characterized in that: (Item 4) the one or more conditions include a condition that the one or more first images are captured within a second time period; the second period includes a time when the second image is captured and is a period during which the predetermined subject is continuously detected; 4. The imaging device according to item 3, (Item 5) a second determination means for determining whether or not a degree of gaze of the user is equal to or greater than a first threshold for each of the at least one first image; the one or more conditions include a condition that the gaze degree of each of the one or more first images is equal to or greater than the first threshold; 5. The imaging device according to any one of items 1 to 4, characterized in that: (Item 6) the photographing means is configured to transition the imaging device to a photographing preparation state for the second image in response to a predetermined user operation on an operation member; the one or more conditions include a condition that the one or more first images are captured within a third time period; the third period is a period from a transition to the latest shooting preparation state before shooting the second image; 6. The imaging device according to any one of items 1 to 5, characterized in that: (Item 7) the third period is a period from a transition to the latest imaging preparation state before imaging the second image to imaging the second image. 7. The imaging device according to item 6, (Item 8) the operating member is a button that has a half-pressed state and a fully-pressed state, the predetermined user operation is a half-press of the button, The photographing means is configured to photograph the second image in response to a full press of the button. 8. The imaging device according to item 6 or 7, characterized in that: (Item 9) further comprising a detection means for detecting a magnitude of change between two or more images of the plurality of first images and the second image when a plurality of first images of the at least one first image satisfy the one or more conditions; When the change is smaller than a second threshold, the generation means generates the linguistic information describing the content of the second image based on the second image and one or more remaining first images excluding some of the plurality of first images. 9. The imaging device according to any one of items 1 to 8, characterized in that: (Item 10) The smaller the change, the greater the number of the portions of the plurality of first images. 10. The imaging device according to item 9, (Item 11) the detection means detects the magnitude of the movement of the subject in the two or more images as the magnitude of the change between the two or more images. 11. The imaging device according to item 9 or 10, (Item 12) a detection means for detecting a magnitude of change between each of the plurality of first images and the second image when a plurality of first images among the at least one first image satisfy the one or more conditions; the generating means generates the linguistic information describing the content of the second image based on the second image and one or more first images remaining after excluding some of the plurality of first images; the variance of each of the portions of the plurality of first images is less than a third threshold; 9. The imaging device according to any one of items 1 to 8, characterized in that: (Item 13) an imaging means for capturing at least one live view image and an image for recording; a generating means for generating language information that explains the content of the image for recording based on one or more live preview images that satisfy one or more conditions among the image for recording and the at least one live preview image; An imaging device comprising: (Item 14) A control method executed by an imaging device, comprising: an imaging step of capturing at least one first image that is not recorded in a non-volatile storage and a second image that is recorded in the non-volatile storage; a generating step of generating linguistic information describing the content of the second image based on one or more first images that satisfy one or more conditions among the second image and the at least one first image; A control method comprising: (Item 15) A control method executed by an imaging device, comprising: an imaging step of capturing at least one live view image and an image for recording; a generating step of generating language information that explains the content of the image for recording based on one or more live preview images that satisfy one or more conditions among the image for recording and the at least one live preview image; A control method comprising: (Item 16) A program for causing a computer to function as each of the means of the imaging device described in any one of items 1 to 13.
[0087] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0088] 100...imaging device, 101...system bus, 102...CPU, 103...ROM, 104...memory, 105...interface section, 106...display section, 107...imaging section, 108...storage, 201...input control section, 202...caption generation section
Claims
1. an imaging means for capturing at least one first image that is not recorded in a non-volatile storage and a second image that is recorded in the non-volatile storage; a generating means for generating linguistic information that describes the content of the second image based on one or more first images that satisfy one or more conditions among the second image and the at least one first image; An imaging device comprising:
2. the one or more conditions include a condition that the one or more first images are captured within a first period that includes a time when the second image is captured; 2. The imaging device according to claim 1.
3. a first determination means for determining whether or not a predetermined subject is included in each of the at least one first image; the one or more conditions include a condition that each of the one or more first images includes the predetermined subject; 2. The imaging device according to claim 1.
4. the one or more conditions include a condition that the one or more first images are captured within a second time period; the second period includes a time when the second image is captured and is a period during which the predetermined subject is continuously detected; 4. The imaging device according to claim 3.
5. a second determination means for determining whether or not a degree of gaze of the user is equal to or greater than a first threshold for each of the at least one first image; the one or more conditions include a condition that the gaze degree of each of the one or more first images is equal to or greater than the first threshold; 2. The imaging device according to claim 1.
6. the photographing means is configured to transition the imaging device to a photographing preparation state for the second image in response to a predetermined user operation on an operation member; the one or more conditions include a condition that the one or more first images are captured within a third time period; the third period is a period following a transition to the latest imaging preparation state before imaging the second image; 2. The imaging device according to claim 1.
7. the third period is a period from a transition to the latest imaging preparation state before imaging the second image to imaging the second image; 7. The imaging device according to claim 6.
8. the operating member is a button that has a half-pressed state and a fully-pressed state, the predetermined user operation is a half-press of the button, the photographing means is configured to photograph the second image in response to a full press of the button.
7. The imaging device according to claim 6.
9. a detection means for detecting a magnitude of change between two or more images of the plurality of first images and the second image when a plurality of first images of the at least one first image satisfy the one or more conditions; When the change is smaller than a second threshold, the generation means generates the linguistic information describing the content of the second image based on the second image and one or more remaining first images excluding some of the plurality of first images.
2. The imaging device according to claim 1.
10. the smaller the change, the greater the number of the portions of the plurality of first images; 10. The imaging device according to claim 9.
11. the detection means detects the magnitude of the movement of the subject in the two or more images as the magnitude of the change between the two or more images.
10. The imaging device according to claim 9.
12. a detection means for detecting a magnitude of change between each of the plurality of first images and the second image when a plurality of first images among the at least one first image satisfy the one or more conditions; the generating means generates the linguistic information describing the content of the second image based on the second image and one or more first images remaining after excluding some of the plurality of first images; the variance in each of the portions of the plurality of first images is less than a third threshold.
2. The imaging device according to claim 1.
13. an imaging means for capturing at least one live view image and an image for recording; a generating means for generating language information that explains the content of the image for recording based on one or more live preview images that satisfy one or more conditions among the image for recording and the at least one live preview image; An imaging device comprising:
14. A control method executed by an imaging device, comprising: an imaging step of capturing at least one first image that is not recorded in a non-volatile storage and a second image that is recorded in the non-volatile storage; a generating step of generating linguistic information describing the content of the second image based on one or more first images that satisfy one or more conditions among the second image and the at least one first image; A control method comprising:
15. A control method executed by an imaging device, comprising: an imaging step of capturing at least one live view image and an image for recording; a generating step of generating language information that explains the content of the image for recording based on one or more live preview images that satisfy one or more conditions among the image for recording and the at least one live preview image; A control method comprising:
16. A program for causing a computer to function as each of the means of the imaging device according to any one of claims 1 to 13.
Citation Information
Patent Citations
Explanatory sentence creation device, object information representation system, and explanatory sentence creation method
JP2020013427A