Imaging device, imaging support method, and program
The imaging device integrates an AI interface to send and display AI-generated support information, addressing the lack of suitable AI configurations for imaging assistance and improving user support.
Patent Information
- Application Number
- PCT/JP2024/026516
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-01-29
AI Technical Summary
Existing technologies do not adequately provide a suitable configuration for utilizing artificial intelligence to support users in imaging tasks.
An imaging device equipped with an imaging unit, display unit, and control unit that sends instructions to a generating AI for imaging support, receives response information, and displays it on the display unit to assist users.
Provides a more suitable response output technique for imaging support, enhancing user assistance in imaging tasks.
Smart Images

Figure JP2024026516_29012026_PF_FP_ABST
Abstract
Description
Imaging device, imaging support method, and program
[0001] The present invention relates to an imaging device, an imaging support method, and a program.
[0002] A response output technology using artificial intelligence such as a language model is disclosed in, for example, Patent Document 1.
[0003] Special table 2019-528512 publication
[0004] However, the disclosure of Patent Document 1 does not sufficiently consider a configuration for more suitably providing a response output technology using artificial intelligence to a user.
[0005] An object of the present invention is to provide a more suitable response output technique.
[0006] In order to solve the above problem, for example, the configuration described in the claims is adopted. The present application includes a plurality of means for solving the above problem, and one example thereof may be an imaging device including an imaging unit, a display unit, and a control unit, wherein the control unit is configured to execute the following processes: sending instruction information to a generating artificial intelligence (AI) requesting imaging support information for supporting a user in imaging using the imaging unit; receiving response information including the imaging support information from the generating artificial intelligence; and displaying the imaging support information included in the received response information or information based on the imaging support information on the display unit.
[0007] As another example, an imaging support method may be configured such that a control unit in an imaging device executes the following processes: sending instruction information to a generating artificial intelligence requesting imaging support information to support a user in imaging using the imaging unit; receiving response information from the generating artificial intelligence including the imaging support information; and displaying the imaging support information included in the received response information or information based on the imaging support information on a display unit.
[0008] As another example, a program may be configured to cause a processor constituting a control unit in an imaging device to execute the following processes: sending instruction information to a generating artificial intelligence requesting imaging support information to assist a user in imaging using the imaging unit; receiving response information from the generating artificial intelligence including the imaging support information; and displaying the imaging support information included in the received response information or information based on the imaging support information on a display unit.
[0009] According to the present invention, a more suitable response output technique can be provided. Other problems, configurations, and effects will become clear in the following description of the embodiments.
[0010] FIG. 1 is a diagram showing an example of an artificial intelligence response output device and system according to an embodiment of the present invention. FIG. 1 is a diagram showing an example of an artificial intelligence response output device according to an embodiment of the present invention. FIG. 2 is a diagram showing an example of the operation of the artificial intelligence response output device and system according to an embodiment of the present invention. FIG. 2 is a diagram showing an example of a system including an imaging device and a generating artificial intelligence server according to an embodiment of the present invention. FIG. 3 is a diagram showing an example of the configuration of an imaging device according to an embodiment of the present invention. FIG. 4 is a diagram showing an example of a processing flow in an imaging device and a generating artificial intelligence server according to an embodiment of the present invention. FIG. 5 is a diagram showing an example of a processing flow in an imaging device and a generating artificial intelligence server according to an embodiment of the present invention. FIG. 6 is a diagram showing an example of a processing flow in an imaging device and a generating artificial intelligence server according to an embodiment of the present invention. FIG. 7 is a diagram showing an example of a processing flow in an imaging device and a generating artificial intelligence server according to an embodiment of the present invention. FIG. 8 is a diagram showing an example of a display screen in an imaging device according to an embodiment of the present invention. FIG. 9 is a diagram showing an example of a change in composition before and after receiving advice from a generating artificial intelligence according to an embodiment of the present invention. FIG. 10 is a diagram showing an example of how an unnecessary object is extracted in a through image obtained by an imaging device according to an embodiment of the present invention. FIG. 11 is a diagram showing an example of an original image and a temporarily processed image displayed in an imaging device according to an embodiment of the present invention.
[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to the description of the embodiments, and various changes and modifications can be made by those skilled in the art within the scope of the technical ideas disclosed in this specification. Furthermore, in all drawings used to explain the present invention, components having the same functions are given the same reference numerals, and repeated explanations thereof may be omitted.
[0012] Note that if the AI response output device according to each embodiment of the present invention has a display screen, it may be referred to as a display device. If the AI response output device has an audio output function, it may be referred to as an audio output device. The AI response output device may simply be referred to as an information processing device. A system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI response output system. Furthermore, if the AI response output device provides a user with a response service based on a large-scale language model, which is an AI, and is helpful to the user, the AI response output device or the display output of the AI response output device can serve as an AI (AI) assistant for the user. Therefore, in this case, the AI response output device may be referred to as an AI assistant device or an AI assistant display device. Similarly, in this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI assistant system or an AI assistant display system. Furthermore, in this case, the AI response output device serves as an interface between the user and the AI, and therefore may be referred to as an AI interface device. In this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI interface system.
[0013] First Embodiment As a first embodiment of the present invention, an AI response output device and system for outputting a response from a large-scale language model AI will be described.
[0014] 1A, an example of an AI response output device 10010 of the present invention will be described. In addition, in the case where the AI response output device 10010 cooperates with a large-scale language model server 19001 via communication or the like, an example of a system in which the AI response output device 10010 includes the large-scale language model server 19001 and / or a multimodal large-scale language model server 20001 will be described.
[0015] In the example of FIG. 1A, the AI response output device 10010 has a display unit 10011. In the example of FIG. 1A, the display unit 10011 may be a flat display, a screen that projects an image from the rear, or a floating image that forms an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which the pixels are self-luminous. The display unit 10011 may also be provided with a touch operation input sensor and configured as a touch panel.
[0016] 1A, the audio output unit 1140 of the AI response output device 10010 is composed of a speaker. The AI response output device 10010 also has a microphone 1139 that can pick up the user's voice. By audio input from the microphone 1139 or user operation input via an operation input unit (described later), the AI response output device 10010 can acquire user input that serves as the basis for instruction sentences (prompts) for the large-scale language model, which is the AI.
[0017] The AI response output device 10010 may be provided with a local large-scale language model in the AI response output device 10010. In this case, the response of the large-scale language model may be output as a display output from the display unit 10011 and / or an audio output from the audio output unit 1140.
[0018] In addition, the artificial intelligence response output device 10010 may not have a local large-scale language model, but may communicate with an external large-scale language model server 19001, and output the response received from the large-scale language model server 19001 as a display output on the display unit 10011 and / or as an audio output on the audio output unit 1140.
[0019] Alternatively, the AI response output device 10010 may also include a local large-scale language model and may be configured to communicate with an external large-scale language model server 19001 having the large-scale language model or an external large-scale language model server 20001 having a multimodal large-scale language model. In this case, the AI response output device 10010 may switch between a response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001, and output either one as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140. Alternatively, a response generated based on both the response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001 may be output as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140.
[0020] The configuration when the AI response output device 10010 communicates and cooperates with an external large-scale language model server 19001 or large-scale language model server 20001 is as follows. The AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 1A , the communication between the communication unit 1132 and the communication device 19011 is shown as being wireless, but wired communication is also acceptable. The communication path from the communication unit 1132 to the communication device 19011 may include wired and wireless portions, or may go via a router or repeater. Furthermore, the communication path from the communication unit 1132 to the Internet 19000 may include wired and wireless portions, or may go via a router or repeater. The AI response output device 10010 can communicate with the large-scale language model server 19001 via the communication device 19011 and the Internet 19000. Furthermore, the AI response output device 10010 can communicate with the large-scale language model server 19001 or the large-scale language model server 20001, and a second server 19002 different from these servers, via a communication device 19011 and the Internet 19000. A configuration including the AI response output device 10010 and the large-scale language model server 19001 or the large-scale language model server 20001 may be considered as a single system.
[0021] In the following explanation, unless otherwise specified, the term "large-scale language model" may be considered to refer to the local large-scale language model provided by the AI response output device 10010, the large-scale language model provided by the large-scale language model server 19001, and the multimodal large-scale language model provided by the large-scale language model server 20001.
[0022] The example of Figure 1A shows an example in which the display unit 10011 displays elements in two display areas: a prompt display area 10051 in which a user inputs a prompt to a large-scale language model, which is an artificial intelligence; and an artificial intelligence response display area 10061 in which a response from the large-scale language model is displayed. In the example of Figure 1A, the prompt display area 10051 displays an icon 10052 indicating a user, text 10053 such as natural language or software code as a component of the prompt, an image 10054 as a component of the prompt, and a video 10055 as a component of the prompt. In the example of Figure 1A, the artificial intelligence response display area 10061 displays an icon 10062 indicating an artificial intelligence or an artificial intelligence assistant, text 10063 such as natural language or software code as a component of the response from the artificial intelligence, an image 10064 as a component of the response from the artificial intelligence, and a video 10065 as a component of the response from the artificial intelligence. The display example of the display unit 10011 of the AI response output device 10010 shown in Fig. 1A is merely an example. Depending on the implementation example in which the AI response output device 10010 is used, a display different from the example shown in Fig. 1A may be performed.
[0023] Here, large-scale language models will be described. Large-scale language models are also referred to as LLMs (Large Language Models). Specifically, various models have been published, such as GPT-1, GPT-2, GPT-3, InstructGPT, and ChatGPT. These technologies may be used in this embodiment as well. Note that these large-scale language models are artificial intelligence models generated by large-scale pre-training on the natural language contained in numerous documents and texts existing in the human world. The number of parameters in artificial intelligence models exceeds 100 million. In addition to this, there are also models that have undergone reinforcement learning based on feedback from humans. An example of a base model is a model called a Transformer. Reference 1, for example, has been published as an example of learning these models.
[0024] [Reference 1] Long Ouyang, et. al. “Training language models to follow instructions with human feedback”, https: / / arxiv.org / pdf / 2203.02155.pdf
[0025] These large-scale language models are capable of natural language translation, natural language text proofreading, and natural language text summarization. Advanced models are capable of natural language question answering (also known as dialogue or conversation), natural language suggestion generation, and programming code generation. Because these AI models have a very large number of parameters, training requires vast amounts of data and computational resources. Therefore, training this level of AI for a specific application is extremely resource-inefficient. Therefore, models are generated through large-scale pre-training as foundation models applicable to various applications. For example, the large-scale language model server 19001 shown in FIG. 1A may be equipped with such a large-scale language model and configured to be accessible on various terminals via an API (Application Programming Interface). Furthermore, the AI response output device 10010 shown in FIG. 1A may be equipped with a local large-scale language model and configured to be used by the AI response output device 10010 itself. The learning of any large-scale language model itself can be generated by separate large-scale pre-learning, and the generated large-scale language model can be replicated and provided in the large-scale language model server 19001, the AI response output device 10010, etc. In this way, instead of performing pre-learning for each application or each terminal, replicating the large-scale language model that is the base model generated by large-scale pre-learning and using it on individual servers or terminals allows the resources used for learning to be shared, resulting in good resource efficiency.
[0026] Furthermore, even if a large-scale language model is used as a base model generated through large-scale pre-training, it may be configured so that additional learning such as transfer learning is performed on individual servers or devices depending on the application or purpose.
[0027] Furthermore, large-scale language models can be pre-trained on natural languages and perform input / output processing targeting natural languages. Furthermore, multimodal large-scale language model AI capable of processing not only natural language text information but also types of information other than natural language text information can also be applied to embodiments of the present invention. FIG. 1A illustrates a large-scale language model server 20001 having a multimodal large-scale language model. For example, specific examples of multimodal large-scale language model AI include GPT-4 (see Reference 2) and Gato (see Reference 3). These technologies may also be used in this embodiment. These multimodal large-scale language models are AI models generated by large-scale pre-training on the natural language contained in numerous documents and texts present in the human world, as well as types of information other than natural language text information (e.g., images, videos, audio, etc.). In addition to this, there are also models that undergo reinforcement learning based on human feedback. Hereinafter, types of information other than natural language text information, such as images, videos, and audio, may be referred to as non-natural language information sources.
[0028] [Reference 2] Open AI “GPT-4 Technical Report”, https: / / cdn.openai.com / papers / gpt-4.pdf [Reference 3] Scott Reed, et. al. “A Generalist Agent”, https: / / arxiv.org / pdf / 2205.06175.pdf
[0029] Next, using Figure 1B, we will explain an example configuration of an artificial intelligence response output device 10010 that accepts input from a user to artificial intelligence such as these large-scale language models and outputs a response from the artificial intelligence such as a large-scale language model to the input from the user.
[0030] The AI response output device 10010 includes a display unit 10011, a control unit 1110, a memory 1109, a non-volatile memory 1108, an external power input interface 1111, an operation input unit 1107, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, a posture sensor 1113, a communication unit 1132, an audio output unit 1140, a microphone 1139, a video signal input unit 1131, an audio signal input unit 1133, an imaging unit 1180, etc. The AI response output device 10010 may have a large screen, such as a monitor or television.
[0031] The display unit 10011 may be a flat display, a screen that projects an image from the rear, or a device that displays a floating image by forming an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which the pixels are self-luminous. If the display unit 10011 is a panel, it may be referred to as a display panel. The display unit 10011 may be provided with a touch operation input sensor and configured to accept touch operation input by the finger of the user 230. In this case, the display unit 10011 may be configured as a touch panel. By the user's operation input via the touch panel, the artificial intelligence response output device 10010 can acquire user input that serves as the basis for instructions (prompts) to the large-scale language model, which is the artificial intelligence.
[0032] The communication unit 1132 may be configured with a Wi-Fi communication interface, a Bluetooth (registered trademark) communication interface, a mobile communication interface such as 4G or 5G, or the like. Using these communication methods, the communication unit 1132 of the AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000. Note that the communication path between the communication unit 1132 and the communication device 19011 may include wired and wireless portions, or may go via a router or repeater. In the case of a wired connection, the communication unit 1132 may have an Ethernet connection interface as hardware and communicate using a LAN communication method. This allows the AI response output device 10010 to communicate with various servers connected to the Internet 19000.
[0033] The AI response output device 10010 is provided with a control unit 1110 such as a CPU and a memory 1109, and the control unit 1110 controls the display unit 10011, the communication unit 1132, and the like.
[0034] The power supply 1106 converts AC current input from the outside via the external power supply input interface 1111 into DC current and supplies the DC current required by each component of the AI response output device 10010. The secondary battery 1112 stores the power supplied from the power supply 1106. Furthermore, the secondary battery 1112 supplies power to each component requiring power via the external power supply input interface 1111 when power is not supplied from the outside.
[0035] The operation input unit 1107 is, for example, an operation button, a signal receiving unit or an infrared light receiving unit of a remote controller, and inputs a signal for an operation different from a user's touch operation on the touch operation input sensor of the display unit 10011. Separate from a user touching the touch operation input sensor of the display unit 10011, the operation input unit 1107 may be used, for example, by an administrator to operate the AI response output device 10010. By the user's operation input via the operation input unit 1107, the AI response output device 10010 can acquire user input that serves as the basis for instructions (prompts) for the large-scale language model, which is the AI. A modified configuration is also possible in which the touch operation input sensor of the display unit 10011 is also included as part of the operation input unit 1107.
[0036] The video signal input unit 1131 connects to an external video output device and inputs video data. The video signal input unit 1131 may be configured with various digital video input interfaces. For example, the video signal input unit 1131 may be configured with a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 1131 may also be configured with various USB interfaces.
[0037] The audio signal input unit 1133 is connected to an external audio output device and inputs audio data. The audio signal input unit 1133 may be configured as an HDMI-standard audio input interface, an optical digital terminal interface, a coaxial digital terminal interface, or the like. The audio signal input unit 1133 may also be various USB interfaces, etc. In the case of an HDMI-standard interface, the video signal input unit 1131 and the audio signal input unit 1133 may be configured as an interface in which a terminal and a cable are integrated.
[0038] The audio output unit 1140 can output audio based on audio data input to the audio signal input unit 1133. The audio output unit 1140 can also output audio based on audio data stored in the storage unit 1170. The audio output unit 1140 may be configured with a speaker. The audio output unit 1140 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 1140 may be configured to output an audio signal as a digital signal to an external device, such as the Audio Return Channel function defined in the HDMI standard. Alternatively, the audio output unit 1140 may be configured to output an audio signal as an analog signal to an external device such as headphones.
[0039] The microphone 1039 is a microphone that picks up sounds around the AI response output device 10010, converts them into signals, and generates audio signals. The microphone may record a person's voice, such as a user's voice, and the control unit 1110, which will be described later, performs voice recognition processing on the generated audio signal to acquire text information from the audio signal. By using audio input from the microphone 1139, the AI response output device 10010 can acquire user input that serves as the basis for instructions (prompts) to the large-scale language model, which is the AI.
[0040] The imaging unit 1180 is a camera having an image sensor. A camera may be provided on the front side of the display unit 10011 of the AI response output device 10010, or on the back side of the display unit 10011. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 1180 will be described as having both a front camera and a back camera.
[0041] The storage unit 1170 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 1170 may be configured with a magnetic recording medium recording device, such as a hard disk drive (HDD), or a semiconductor device memory, such as a solid-state drive (SSD). For example, various types of information, such as video data, image data, and audio data, may be recorded in the storage unit 1170 before product shipment. The storage unit 1170 may also record various types of information, such as video data, image data, and audio data, acquired from an external device, an external server, or the like, via the communication unit 1132. The video data, image data, and the like recorded in the storage unit 1170 are output to the display unit 10011. The video data, image data, and the like recorded in the storage unit 1170 may also be output to an external device, an external server, or the like via the communication unit 1132.
[0042] The video control unit 1160 performs various controls related to the video signal input to the display unit 10011. The video control unit 1160 may be referred to as a video processing circuit and may be configured with hardware such as an ASIC, an FPGA, or a video processor. The video control unit 1160 may also be referred to as a video processing unit or an image processing unit. The video control unit 1160 performs, for example, video switching control, such as determining which video signal to input to the display unit 10011 between the video signal to be stored in the memory 1109 and the video signal (video data) input to the video signal input unit 1131. The video control unit 1160 may also perform control to perform image processing on the video signal input from the video signal input unit 1131 and the video signal to be stored in the memory 1109. Examples of image processing include scaling processing, which enlarges, reduces, or deforms an image; brightness adjustment processing, which changes the brightness; contrast adjustment processing, which changes the contrast curve of an image; and Retinex processing, which decomposes an image into light components and changes the weighting of each component.
[0043] The attitude sensor 1113 is a sensor configured by a gravity sensor or an acceleration sensor, or a combination of these, and can detect the attitude of the AI response output device 10010. Based on the attitude detection result of the attitude sensor 1113, the control unit 1110 may control the operation of each unit connected thereto.
[0044] The non-volatile memory 1108 stores various data used by the AI response output device 10010. The data stored in the non-volatile memory 1108 includes, for example, data for various operations to be displayed on the display unit 10011 of the AI response output device 10010, display icons, data and layout information for objects to be operated by user operations, etc. The memory 1109 stores video data to be displayed on the display unit 10011, data for controlling the device, etc. The control unit 1110 may read various software from the storage unit 1170 and expand and store it in the memory 1109.
[0045] The local LLM processing unit 10028 includes a memory capable of storing a large-scale language model (LLM) and can execute inference of the large-scale language model under the control of the control unit 1110. The hardware may be configured with a so-called GPU (Graphics Processing Unit) or the like. The local LLM processing unit 10028 may perform not only inference but also learning. Note that the local LLM processing unit 10028 is not necessarily required in cases where it is not necessary to execute inference of a large-scale language model in the local environment of the AI response output device 10010.
[0046] The control unit 1110 controls the operation of each connected unit. The control unit 1110 may also work in cooperation with a program stored in the memory 1109 to perform arithmetic processing based on information acquired from each unit in the AI response output device 10010. The control state of the control unit 1110 includes, for example, a state in which a response from the large-scale language model of the local LLM processing unit 10028 or a response from the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 acquired via the communication unit 1132 is output via the display unit 10011 or the audio output unit 1140, such as a speaker.
[0047] When a user inputs via the touch panel, microphone 1139, or operation input unit 1107, an instruction sentence is generated based on the input, and transmitted to the local large-scale language model of the local LLM processing unit 10028 provided in the AI response output device 10010, the large-scale language model provided in the large-scale language model server 19001, or the multimodal large-scale language model provided in the large-scale language model server 20001. All of these controls for obtaining a response from these large-scale language models can be performed by the control unit 1110.
[0048] The storage unit 1170 may also store a fixed response phrase database (which may also be referred to as a fixed response phrase DB) for outputting fixed phrases in response to instruction statements from the AI response output device 10010. The control unit 1110 may control the generation of responses to be output using data stored in the fixed response phrase database. FIG. 1C shows an example of the fixed response phrase database. In the example of FIG. 1C, fixed responses to be output by the AI response output device 10010 are stored for each condition assigned a condition number. For example, as in condition number 1, when the user inputs "Good morning" via the touch panel, microphone 1139, or operation input unit 1107, a response may be output using a fixed response phrase such as "Good morning" or "Today is ____ day of ____ month, isn't it?" The ____ part of "____ day of ____ month" may be generated using information stored in the memory 1109 or the like of the AI response output device 10010.
[0049] Furthermore, in the example of the standard response phrases in the database shown in FIG. 1C, if multiple standard response phrases separated by / are stored, the control unit 1110 may control the output of a response by randomly selecting one of the standard response phrases using a random number or the like. This can eliminate or improve the situation where responses under the same conditions become monotonous. The explanation for the example of condition number 1 is the same for the examples of condition numbers 2, 3, and 4. The control unit 1110 may control the output of the standard response phrase of each example shown in FIG. 1C for the condition content of each example shown in FIG. 1C.
[0050] Next, an example of condition number 5 shown in FIG. 1C will be described. Condition number 5 is an example of control in which, when the control unit 1110 cannot understand the meaning of a user input acquired via the touch panel, microphone 1139, or operation input unit 1107 as natural language or when the user input contains an obvious grammatical error, the control unit 1110 outputs a response using a standard response phrase such as "I didn't quite catch what you said" or "I might not know about that." By responding in this manner, the user can be prompted to input again, and the system can wait for a corrected user input.
[0051] Next, an example of condition number 6 shown in Figure 1C will be described. Condition number 6 is an example of a case where the control unit 1110 detects an error (abnormal state) in any of the components constituting the AI response output device 10010 shown in Figure 1B, and a user input is made via the touch panel, microphone 1139, or operation input unit 1107. In this case, the control unit 1110 performs control to output a response using the standard response phrase "It seems to be working poorly." By responding in this manner, it is possible to explain to the user that the AI response output device 10010 is malfunctioning, and to prompt the user to take action on the error, etc.
[0052] The AI response output device 10010 may output a response using the fixed response phrase database (fixed response phrase DB) described with reference to Fig. 1C instead of a response from a large-scale language model such as the local large-scale language model provided in the AI response output device 10010, the large-scale language model provided in the large-scale language model server 19001, or the multimodal large-scale language model provided in the large-scale language model server 20001. Alternatively, the AI response output device 10010 may output a response that combines the responses from these large-scale language models with a response using the fixed response phrase database (fixed response phrase DB).
[0053] 1C described above may be stored in the storage unit 1170 and used by the control unit 1110 of the AI response output device 10010. However, the fixed response database (fixed response DB) shown in FIG. 1C may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the fixed response database (fixed response DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may transmit a response generated using the fixed response database (fixed response DB) to the AI response output device 10010, instead of a response generated using a large-scale language model stored in the respective server. In this way, even if the artificial intelligence response output device 10010 is not equipped with a standard response phrase database (standard response phrase DB), it is possible to generate a response using the standard response phrase database (standard response phrase DB).
[0054] In the above description, the AI response output device 10010 has been described as having a display panel with a display screen using fixed pixels. This concept may also include a projection type image display device (projector) in which a projection optical system is provided behind the display panel with a display screen using fixed pixels, and an optical image of the image on the display panel of the display screen is projected onto a screen or wall.
[0055] 1A and 1B, an example has been described in which the AI response output device 10010 includes the display unit 10011. However, the AI response output device 10010 according to an embodiment of the present invention does not necessarily have to include the display unit 10011. For example, even if the display unit 10011 is not included, the AI may be configured to accept input from a user to the AI via the voice signal input unit 1133 or the microphone 1139, and output a response from the AI, such as a large-scale language model, in response to the user input via the voice output unit 1140.
[0056] According to the artificial intelligence response output device and artificial intelligence response output system of the first embodiment of the present invention described above, it is possible to accept input from a user to an artificial intelligence such as a large-scale language model, and output a response to the input from the user that is generated by inference by the artificial intelligence, such as a large-scale language model held by a server device on a network or a local large-scale language model held by the artificial intelligence response output device itself.
[0057] Furthermore, the technology according to this embodiment makes it possible to provide a more suitable AI response output technology. Such AI response output technology is expected to be introduced into higher quality, more reliable infrastructure. The introduction of this technology into infrastructure will contribute to supporting economic development and human welfare, with a focus on affordable and fair access for all. This will contribute to the achievement of the "9 Sustainable Development Goals (SDGs)" advocated by the United Nations, "Build resilient infrastructure, promote inclusive and sustainable industrialization, innovate and innovate."
[0058] Furthermore, the technology according to this embodiment makes it possible to provide a more suitable AI response output technology. Such AI response output technology is expected to be introduced into public transportation facilities to improve access to transportation systems for vulnerable people. The introduction of this technology into public transportation can contribute to improving traffic safety through the expansion of public transportation and realizing access to a safe, affordable, and easily accessible sustainable transportation system for all people. This contributes to "Sustainable cities and communities," one of the Sustainable Development Goals (SDGs) advocated by the United Nations.
[0059] Second Embodiment A second embodiment of the present invention is an example of an imaging device and an imaging support method to which an artificial intelligence response output technology is applied. The imaging device and the imaging support method according to the second embodiment support a user in imaging by an AI assistant function using the artificial intelligence response output technology.
[0060] 2A is a diagram illustrating an example of a system including an imaging device and a generative artificial intelligence (hereinafter, artificial intelligence is also referred to as AI) server according to a second embodiment. As illustrated in FIG. 2A , the system according to the second embodiment includes an imaging device 1200, a large-scale language model (hereinafter, large-scale language model is also referred to as LLM) server 19001, a multimodal LLM server 20001, and a second server 19002 different from these servers. The imaging device 1200, the LLM server 19001, the multimodal LLM server 20001, and the second server 19002 are connected via the Internet 19000, which is a communication network. The LLM server 19001 and the multimodal LLM server 20001 are each an example of a generative AI server on which a generative AI is built.
[0061] The imaging device 1200 has an imaging function unit 1201 having an imaging unit 1180, and an AI response output device 10010. The AI response output device 10010 has almost the same configuration and functions as the AI response output device 10010 in the first embodiment.
[0062] 2B is a diagram illustrating an example of the configuration of an imaging device according to Example 2. The imaging function unit 1201 includes an imaging unit 1180, a captured image input unit 1181, an external image input unit 1182, an image input adjustment unit 1183, an image and audio memory unit 1184, an image encoding unit 1185, an audio encoding unit 1186, a data generation unit 1187, a data recording unit 1188, a recording medium 1189, a current position acquisition unit 1195, a video signal input unit 1131, and an audio signal input unit 1133.
[0063] The imaging unit 1180 is for obtaining image information corresponding to a captured image of a subject. The imaging unit 1180 includes, for example, an optical system such as a lens, an aperture, a shutter, an imaging element, etc. The captured image input unit 1181 accepts input of image information of the captured image obtained by the imaging unit 1180. The external image input unit 1182 accepts input of image information of an external image through the video signal input unit 1131. The image input adjustment unit 1183 adjusts whether image information of the captured image or the external image is to be input. The image audio memory unit 1184 temporarily stores image information input from the image input adjustment unit 1183 and audio information input from the audio signal input unit 1133.
[0064] The image encoding unit 1185 converts and compresses the image information stored in the image and audio memory unit 1184 into data in a predetermined format. The image encoding unit 1186 converts and compresses the audio information stored in the image and audio memory unit 1184 into data in a predetermined format. The data generation unit 1187 generates image data or video data based on data from the image encoding unit 1185, or based on data from the image encoding unit 1185 and data from the image encoding unit 1186. The data recording unit 1188 records the image data or video data generated by the data generation unit 1187 on a recording medium 1189. The recording medium 1189 is, for example, an HDD, SSD, USB ROM, DVD, or CD. The current location acquisition unit 1195 acquires location information indicating the current location of the imaging device 1200, such as a GPS system.
[0065] The AI response output device 10010 includes a control unit 1110, a display unit 10011, an external power input interface 1111, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, an operation input unit 1107, an attitude sensor 1113, a memory 1109, a local LLM processing unit 10028, a non-volatile memory 1108, a communication unit 1132, an audio output unit 1140, and a microphone 1139. The configuration and functions of the AI response output device 10010 are substantially the same as those of the AI response output devices according to Examples 1 to 9. However, the control unit 1110 is communicably connected to each unit, including a captured image input unit 1181 and an image input adjustment unit 1183, which constitute the imaging function unit 1201. The control unit 1110 can perform various processes on image information of the captured image obtained by the imaging unit 1180.
[0066] A program for implementing the AI assistant function is stored in the storage unit 1170 or the nonvolatile memory 1108. The control unit 1110 reads out this program, expands it into the memory 1109, and executes it to implement the AI assistant function and control the on / off of this function.
[0067] The control unit 1110 executes a process of transmitting to the generation AI instruction information requesting imaging support information for supporting the user in imaging using the imaging unit 1180. The control unit 1110 also executes a process of receiving response information including the imaging support information from the generation AI. The control unit 1110 further executes a process of displaying the imaging support information included in the received response information or information based on the imaging support information on the display unit 10011. Through such processing by the control unit 1110, the user is assisted in imaging using the imaging unit 1180.
[0068] The storage unit 1170, the nonvolatile memory 1108, the memory 1109, and the recording medium 1189 are each an example of a "storage unit" in the present application. The current position acquisition unit 1195 is an example of a "position information acquisition unit" in the present application.
[0069] 2C to 2F are diagrams illustrating an example of a processing flow in the imaging device and the generation AI server according to the second embodiment. The processing flows illustrated in FIGS. 2C to 2F chronologically describe the main processing executed by the control unit 1110 in the imaging device 1200 and the main processing executed by the generation AI server including the multimodal LLM server 20001. The processing flows illustrated in FIGS. 2C to 2F are processing flows assuming that the processing is executed in parallel with processing in response to user operations in the imaging device 1200. It is assumed that in the imaging device 1200, imaging is performed in the background by the imaging unit 1180 at approximately regular intervals, and multiple time-series through images obtained over a recent fixed period are stored in the recording medium 1189, the storage unit 1170, the non-volatile memory 1108, or the memory 1109.
[0070] In step S1001, a process is performed to determine whether an operation to activate the AI assistant function has been performed. Specifically, the control unit 1110 determines whether an operation to activate the AI assistant function has been performed. The operation to activate the AI assistant function may be, for example, an operation of pressing a physical button or a button displayed on a touch panel via the operation input unit 1107, an operation of inputting a predetermined command by voice via the microphone 1139, or an operation of inputting a command by gesture using hands or facial expressions. In this determination, if it is determined that an operation to activate the AI assistant function has been performed (S1001: Yes), the process proceeds to step S1004. On the other hand, if it is determined that an operation to activate the AI assistant function has not been performed (S1001: No), the process proceeds to step S1002.
[0071] In step S1002, a process is performed to determine whether a condition for initiating an inquiry about whether to activate the AI assistant function is satisfied. Specifically, the control unit 1110 determines whether a condition for initiating an inquiry about whether to activate the AI assistant function is satisfied. The condition may be, for example, detecting that the user is spending a certain amount of time or more deciding on a composition. A more specific example of the condition may be, for example, that the shutter button has not been pressed for a certain period of time since the imaging device 1200 was set to imaging mode.
[0072] In this determination, if it is determined that the condition for initiating an inquiry as to whether or not the AI assistant function needs to be activated is satisfied (S1002: Yes), the processing step proceeds to step S1003. On the other hand, if it is determined that the condition for initiating an inquiry as to whether or not the AI assistant function needs to be activated is not satisfied (S1002: No), the processing step returns to step S1001.
[0073] In step S1003, a process is performed to inquire of the user as to whether or not to activate the AI assistant function. Specifically, the control unit 1110 displays text information in natural language inquiring as to whether or not to activate the AI assistant function on the display unit 10011. Note that this process can also be said to be a process of inquiring of the user as to whether or not imaging assistance is required.
[0074] 2G is a diagram showing an example of a display screen of the imaging device according to Example 2. FIG. 2G shows an example of display of text information representing an interactive expression between a generated AI and a user using an AI assistant function. In the example of FIG. 2G, within a display screen 1057, words 10571 of the generated AI are displayed on the right side, and words 10572 of the user are displayed on the left side. In addition, in the example of FIG. 2G, each word is displayed in chronological order from top to bottom.
[0075] In step S1003, for example, as shown in FIG. 2G, text information such as "Do you want to start the AI assistant?" is displayed. Furthermore, in addition to or instead of the above display, the control unit 1110 may cause the voice output unit 1140 to output a voice such as "Do you want to start the AI assistant?" In this way, by asking the user whether to start the AI assistant function, the user can start the AI assistant function in situations where the user is unaware of the existence of the AI assistant function, has forgotten about it, does not know how to start it, or finds the start-up operation troublesome.
[0076] The control unit 1110 accepts a user's response to an inquiry about whether or not the AI assistant function needs to be activated, i.e., whether or not imaging assistance is needed. If the control unit 1110 accepts a response indicating that the AI assistant function needs to be activated, such as when the user inputs or selects "Yes" in response to the inquiry (S1003: Yes), the processing proceeds to step S1004. On the other hand, if the control unit 1110 accepts a response indicating that the AI assistant does not need to be activated, such as when the user inputs or selects "No" in response to the inquiry (S1003: No), the processing related to the AI assistant function ends, or the parameters related to the conditions for starting the inquiry are reset, and the processing returns to step S1001. The user may answer by pressing the area marked "Yes" or "No" on the screen of the display unit 10011, which functions as a touch panel, with their finger, by pressing the button corresponding to "Yes" or "No," by saying "Yes" or "No," or by inputting their answer using gestures with their hands or facial expressions.
[0077] In step S1004, a process of acquiring current location information is performed. Specifically, the control unit 1110 receives current location information indicating the current location of the imaging device 1200 from the current location acquisition unit 1195. The current location information is, for example, information indicating coordinates on the Earth, and is acquired by GPS.
[0078] The processing of step 1004 may be the following processing. For example, the control unit 1110 receives from the imaging unit 1180 a through image including a landmark that can identify the current location. The control unit 1110 transmits instruction information including through image information of the current location and instructing the generation AI server to provide the current location. Upon receiving this instruction information, the generation AI server generates response information including current location information based on the through image including the landmark, and transmits the response information to the imaging device 1200. The control unit 1110 receives this response information and acquires current location information based on the received response information. The control unit 1110 may also include information indicating the name of the landmark input by the user in the instruction information, thereby increasing the accuracy of determining the current location or the imaging point described below.
[0079] Incidentally, if the generation AI includes a large-scale language model (LLM), the generation AI can handle text information including natural language, colloquial expressions, etc. If the generation AI includes a multimodal LLM rather than just an LLM, the generation AI can handle information other than text information, such as image information and audio information, in addition to text information including natural language, colloquial expressions, etc. The instruction information sent to the generation AI server is also called an instruction sentence, prompt, etc. The generation AI server generates and returns response information in a form that responds to the instruction represented by the received instruction information. The user can receive assistance with imaging in natural language, easily understand the content of the assistance, and easily imagine the image to be captured.
[0080] In step S1005, a process of issuing an instruction to provide an imaging point is performed. Specifically, the control unit 1110 generates instruction information including current location information and instructing the generation AI server to provide an imaging point where an image of a landmark corresponding to the current location represented by the current location information can be captured, and transmits the instruction information to the generation AI server. In other words, this instruction information can also be said to be information requesting, as imaging support information, imaging point information representing an imaging point where an image of a landmark corresponding to the current location can be captured. Note that the imaging point is an example of an "imaging position" in this application, and the imaging point information is an example of "imaging position information" in this application.
[0081] The imaging point may be, for example, a standing or sitting position of the photographer that is considered suitable for capturing an image of a famous landmark near the current location. The imaging point may be, for example, a standing or crouching position of the user that allows the entire landmark to fit within the imaging area. The imaging point may be, for example, a standing or crouching position of the user that allows a composition in which the landmark is placed at a specific position or area within the imaging area. In these cases, the instruction information may include the focal length of the lens attached to the imaging device, the variable range of the focal length of the zoom lens, etc. The generation AI may determine the imaging point based on information regarding the focal lengths of these lenses.
[0082] In step S1006, a process of generating response information including an imaging point is performed. Specifically, upon receiving this instruction information, the generation AI server generates response information including an imaging point corresponding to the current position and transmits the response information to the imaging device.
[0083] In step S1007, a process of determining whether the current position is an imaging point is performed. Specifically, the control unit 1110 compares the current position indicated by the current position information with the imaging point corresponding to the current position, and determines whether the current position is an imaging point based on whether these positions substantially overlap.
[0084] Note that instead of the processing of steps S1004 to S1007, the following processing may be performed. For example, the control unit 1110 transmits to the generation AI server instruction information that includes current position information or through-image information at the current position and instructs the generation AI server to determine whether or not the current position is an imaging point corresponding to the current position. Upon receiving this instruction information, the generation AI server determines whether or not the current position is an imaging point, generates response information that includes determination result information that indicates the determination result, and transmits this response information to the imaging device 1200. The control unit 1110 receives this response information and determines whether or not the current position is an imaging point based on the received response information.
[0085] If it is determined that the current position is an imaging point (S1007: Yes), the process proceeds to step S1011. On the other hand, if it is determined that the current position is not an imaging point (S1007: No), the process proceeds to step S1008.
[0086] In step S1008, a process of instructing the user to be guided to the imaging point is performed. Specifically, the control unit 1110 generates instruction information instructing the user to be guided to the imaging point corresponding to the current position indicated by the current position information, and transmits the generated instruction information to the generation AI server. Note that this instruction information can also be said to be information requesting guidance information for guiding the user to the imaging point as imaging support information.
[0087] In step S1009, a process of generating response information including guidance information for guiding the user to the imaging point is performed. Specifically, upon receiving the instruction information, the generation AI server generates response information including guidance information for guiding the user to the imaging point and transmits the response information to the imaging device 1200.
[0088] 2G, the guidance information is text information in natural language such as, "This is Cologne Cathedral, right? If you turn 1 meter to the right and 3 meters back, you can capture the whole picture." The generation AI server transmits the generated response information to the imaging device 1200.
[0089] In step S1010, a process of outputting guidance information for guiding the user is performed. Specifically, the control unit 1110 receives response information from the generation AI server, including information for guiding the user to the imaging point. Furthermore, the control unit 1110 causes the display unit 10011 to display guidance information for guiding the user based on the received response information. The control unit 1110 may cause the audio output unit 1140 to output the content of the guidance information as audio, in addition to or instead of displaying the guidance information. The guidance information may be information received from the generation AI server as is, or information generated based on the received information may be output. In this way, when the guidance information is output, the user can move to a suitable imaging point in a short time without having to search for the imaging point. Thereafter, the process returns to step S1004.
[0090] Note that instead of the processes of steps S1008 to S1010, the following process may be performed: The control unit 1110 may calculate the difference between the image capture point and the current position, and may generate guidance information for guiding the user to the image capture point based on the difference.
[0091] In step S1011, a process of acquiring at least one of orientation information and through-image information is performed. Specifically, the control unit 1110 receives the acquired orientation information of the image capturing device 1200 from the orientation sensor 1113. Alternatively, the control unit 1110 receives the acquired through-image information from the image capturing unit 1180. The control unit 1110 may receive both the orientation information and the through-image information. As described above, a through-image is an image captured by the image capturing unit 1180 in the background when the image capturing mode is set in the image capturing device 1200 and the shutter button is not pressed. The through-image is acquired, for example, at a predetermined interval.
[0092] In step S1012, a process is performed to instruct the user to give advice on how to capture an image. Specifically, the control unit 1110 generates instruction information that includes at least one of posture information and through-image information and instructs the user to give advice on how to capture an image, and transmits the generated instruction information to the generation AI server. Note that this instruction information can also be said to be information that requests advice information for advising the user on how to capture an image as imaging support information.
[0093] The instruction information may include information indicating a landmark in the through image. For example, the control unit 1110 specifies an arbitrary position or area in the image area of the through image in response to an input operation by the user, and determines an object corresponding to the specified position or area as a landmark. Note that the instruction information may include, for example, an instruction to send a model image to serve as a model.
[0094] In step S1013, a process is performed to generate response information including advice information for advising the user on a suitable imaging method. Specifically, upon receiving the instruction information, the generation AI server generates response information including advice information for advising the user on a suitable imaging method based on at least one of the posture information and the through-image information. Note that the advice information may include, for example, information for instructing the user on the position, orientation, etc. of the imaging device 1200 so that a suitable composition can be obtained.
[0095] As shown in FIG. 2G , the advice information may be text information in natural language, such as, "Hold the camera from the waist down," "Arranging people diagonally will create a sense of depth," or "Pressing the shutter button slowly will prevent blur." Furthermore, for example, the advice information may be text information including natural language, such as, "Placing landmarks on one of the left and right halves and people on the other will result in a well-balanced composition." Furthermore, for example, the advice information may be an image in which an image, symbol, mark, etc. indicating an area in the through image where a landmark should be placed, an area in which the person being photographed should be placed, or both of these areas is superimposed. Furthermore, the advice information may include a model image, which is an exemplary image.
[0096] The content of the advice may be varied depending on the level of the user's imaging skill. For example, the imaging device 1200 may preset the user's imaging skill level from among multiple levels. These multiple levels may be, for example, three levels: advanced, intermediate, and beginner. For example, the imaging device 1200 may record how the imaging device 1200 has been handled in the past in the storage unit 1170 or the data recording unit 1188, and transmit information representing the handling to the generation AI, causing the generation AI to determine the user's imaging skill level. The control unit 1110 then instructs the generation AI to provide advice according to the user's imaging skill level. In this case, the user can receive advice that matches their own imaging skill level.
[0097] As a specific example of advice according to the user's level, the control unit 1110, in cooperation with the generation AI server, presents the user with hints for realizing a suitable position for the imaging device 1200, a suitable orientation for the imaging device 1200, a suitable composition, suitable imaging conditions, and preventing (suppressing) camera shake. If the user is a beginner, the control unit 1110 presents hints on settings related to the aperture, shutter speed, focus, and zoom lens, for example. If the user is an intermediate user, the control unit 1110 presents hints on settings related to ISO sensitivity, white balance (WB), macro lens, and how to deal with fast subject movement (subject shake suppression), for example.
[0098] In addition, the settings of the imaging device 1200, specifically the aperture, shutter speed, ISO sensitivity, on / off of image stabilization, zoom lens magnification, WB, etc., may be automatically set by the control unit 1110 based on the response information received from the generation AI server.
[0099] In step S1014, a process is performed to output advice information for advising the user on how to take an image. Specifically, the control unit 1110 receives response information from the generation AI server, including advice information for advising the user on a suitable way to take an image. Furthermore, the control unit 1110 causes the display unit 10011 to display advice information for advising the user on a suitable way to take an image, based on the response information. The control unit 1110 may cause the audio output unit 1140 to output the advice information for advising the user on the suitable way to take an image by voice, in addition to or instead of the display. The control unit 1110 may output the information received from the generation AI server as is, or may output information generated (edited) based on the received information.
[0100] In addition, if the advice information for advising the user on a suitable imaging method is to be specific, as described above, it is necessary to use posture information or through-image information. On the other hand, if the advice information is to be limited to general content, it is not necessary to use posture information or through-image information. In this case, general advice information may be stored in advance in the storage unit 1170, non-volatile memory 1108, or recording medium 1189 of the imaging device 1200, and the control unit 1110 may read and output the stored advice information. In other words, the control unit 1110 does not need to send instruction information to the generation AI server. In this case, the transmission and reception of information to the generation AI server and the processing in the generation AI server can be reduced, thereby shortening the time required to output advice information and reducing the load on the generation AI server.
[0101] FIG. 2H illustrates an example of a change in composition before and after receiving advice from the generation AI according to Example 2. In FIG. 2H , the composition on the left is a pre-advice composition 251 before the user receives advice from the generation AI, and the composition on the right is a post-advice composition 252 after the user receives advice from the generation AI. In the pre-advice composition 251, as shown in FIG. 2H , a first landmark 10501, a second landmark 10502, and two people 10505 are partially outside the imaging area (angle of view range) 1050. Furthermore, the imaging area 1050 of the pre-advice composition 251 includes a first car 1504, two people 10506 riding bicycles, and a second car 10508 as moving objects moving from left to right. Furthermore, in addition to the person 10505, two other people 10507 are included in the pre-advice composition 251.
[0102] On the other hand, in the post-advice composition 252, as shown in FIG. 2H , the imaging position where the user stands or the magnification of the zoom lens is adjusted so that the first landmark 10501 and the second landmark 10502 do not protrude outside the imaging area 1050. Furthermore, the user guides the imaging target person 10505 or adjusts the magnification of the zoom lens so that the imaging target person 10505 fits within the imaging area 1050. Furthermore, the user guides the two imaging target people 10505 so that they are positioned diagonally so that the three-dimensional effect of the subjects is created in the captured image. Note that no particular guidance is given to moving objects such as the first car 10504, the two people on bicycles 10506, and the second car 10508, or to another person 10507.
[0103] In step S1015, a process is performed to determine whether an operation to obtain a temporary processed image has been performed. Specifically, the control unit 1110 determines whether an operation to obtain a temporary processed image has been performed. The operation to obtain a temporary processed image may be, for example, an operation of half-pressing the shutter button. Generally, when the shutter button is a two-stage switch, the operation of pressing the shutter button to the first stage is called a "half-press," and the operation of pressing the shutter button further to the second stage is called a "full press." If it is determined in this determination that an operation to obtain a temporary processed image has been performed (S1015: Yes), the process proceeds to step S1016. On the other hand, if it is determined that an operation to obtain a temporary processed image has not been performed (S1015: No), the process returns to step S1004.
[0104] In step S1016, a process of issuing an instruction to extract unnecessary objects is performed. Specifically, the control unit 1110 generates instruction information instructing the extraction of unnecessary objects that are considered unsuitable as subjects in the through image, and transmits the generated instruction information to the generation AI server. Note that this instruction information can also be considered information requesting, as imaging support information, extraction result information that represents the result of extracting unnecessary objects that are considered unsuitable as subjects in the through image, which is a captured image. This instruction information may be, for example, information instructing the extraction of at least one of moving objects and non-imaging subjects as unnecessary objects in the through image, excluding landmarks and imaging subjects.
[0105] 2I is a diagram showing an example of how an unnecessary object is extracted from a through image obtained by the imaging device according to Example 2. In Fig. 2I, the upper part shows a plurality of through images 1053A, 1053B, 1053C, and 1053D in time series obtained by processing executed in the background.
[0106] A moving object can be defined as an object whose position changes relative to a landmark or other stationary subject in a plurality of time-series through images 1053A to 1053D, as shown in Fig. 2I. Alternatively, a moving object can be defined as an object that moves across a plurality of grid-like reference lines extending horizontally and vertically in the image area of the through images when image stabilization is in operation.
[0107] 2I, a person who is not a target of imaging can be defined as a person included in any of through images 1053A to 1053D whose facial orientation or line of sight deviates by more than a certain level from the direction facing image capture device 1200. Alternatively, a person who is not a target of imaging can be defined as a person whose proportion of the image area of any of through images 1053A to 1053D is less than a certain level.
[0108] As shown in FIG. 2I , within the imaging area 1050 of the through image 1054, subjects marked with black arrows are extracted as moving objects or people not to be imaged. Specifically, a first automobile 10504, a second automobile 10508, two bicycles, and their drivers, people 10506, are extracted as moving objects whose positions change based on a first landmark 10501 or a second landmark 10502. In addition, two other people 10507, small and facing sideways, are extracted as people not to be imaged, who occupy an area of the imaging area 1050 less than a certain percentage, or whose face or line of sight deviates by more than a certain level from the direction facing the imaging device.
[0109] A moving object is likely to be a passerby, a vehicle, or the like. A person not to be captured is likely to be a stranger from the user's perspective. Therefore, a moving object or a person not to be captured is likely not an object or person that the user originally wants to capture as a subject. When a user wants to remove such objects or people that the user does not want as subjects from a captured image by processing the image, the user can have these objects or people selected automatically to some extent, thereby reducing the need for complicated operations.
[0110] In step S1017, a process is performed to generate response information including extraction result information that indicates the result of the extraction of the unnecessary object. Specifically, upon receiving the instruction information, the generation AI server generates response information including the extraction result information and transmits the generated response information to the imaging device 1200. The extraction result information is, for example, an extraction result image in which a mark indicating the extracted unnecessary object is added to a through image. Furthermore, for example, the extraction result information is information that includes, in addition to the extraction result image, text information including natural language such as "Moving objects and other people in the camera image have been automatically extracted."
[0111] In step S1018, a process of outputting the extraction result information is performed. Specifically, the control unit receives response information including the extraction result information from the generation AI server via the interface, and displays the extraction result information on the display unit 10011. Note that the control unit 1110 may also cause the audio output unit 1140 to output information related to the extraction result as audio, along with displaying the extraction result information. The extraction result information may be the information received from the generation AI server as is, or may be information generated (edited) based on the received information. The extraction result information may also include text information in natural language, such as, for example, "Moving objects and other people in the camera image have been automatically extracted."
[0112] In step S1019, a process for selecting a processing target is performed. Specifically, the control unit 1110 sets all extracted unnecessary objects as processing target candidates as an initial setting. The user performs an operation to remove objects that the user does not want to process from the processing target candidates, or performs an operation to designate objects other than the processing target candidates as processing target candidates. The control unit 1110 selects a processing target in the through image in response to these user operations. Note that the control unit 1110 may select the extracted unnecessary object as the processing target as is, or may select an object desired by the user in the through image as the processing target in response to a user operation in a situation where extraction of an unnecessary object is not performed.
[0113] In step S1020, a process is performed to instruct the processing target to be processed in the through image. Specifically, the control unit 1110 generates instruction information that includes the through image and instructs the processing target to be processed in the through image, and transmits the generated instruction information to the generation AI server. Note that this instruction information can also be considered information requesting, as imaging support information, a processed image in which the processing target, for example, a moving object or a person not to be captured, has been removed from the through image.
[0114] This instruction information may be, for example, information instructing the system to remove an image representing the object to be processed from the through image and restore the area from which the image was removed using image data surrounding the area. In this way, a good-looking image can be generated using AI while using a realistic background included in the composition.
[0115] Furthermore, for example, the instruction information may be information instructing the user to remove an image representing the object to be processed from the through image and restore the removed area using desired image data designated by the user. In this way, a good-looking image that reflects the user's preferences can be generated.
[0116] In step S1021, a process is performed to generate response information including processing result information. Specifically, when the generation AI server receives the instruction information, it processes the processing target in the through image to obtain a processed image, generates response information including the processing result information, and transmits the generated response information to the imaging device. The processing result information includes the processed image. Note that the processing result information may include, in addition to the processed image, text information in natural language such as, for example, "The specified processing target has been excluded and processing has been performed using surrounding image data."
[0117] In step S1022, processing is performed to output the processing result information. Specifically, the control unit 1110 receives response information including the processing result information from the generation AI server, and causes the display unit 10011 to display the processed image included in the processing result information as a temporary processed image. The control unit 1110 also causes the display unit to display text information included in the processing result information, or causes the audio output unit 1140 to output it as audio. The control unit 1110 may also cause the display unit 10011 to display the temporary processed image and the original image, which is a through image before processing, side by side on the display screen of the display unit 10011. This allows the user to compare the temporary processed image with the original image, and to design the processed image to suit their own preferences.
[0118] 2J is a diagram showing an example in which an original image and a temporarily processed image are displayed in the imaging device according to Example 2. For example, as shown in FIG. 2J , an original image 1055 and a temporarily processed image are displayed side by side on the display screen of the display unit 10011. In the original image 1055, in addition to a first landmark 10501, a second landmark 10502, and a person 10505 to be imaged, moving objects including a first automobile 10504 and a second automobile 10508, two bicycles and a person 10506 who is a driver of the bicycle, and a stranger 10507 who is not a person to be imaged are included within the imaging area 1050. On the other hand, in the provisionally processed image 1056, only the first landmark 10501, the second landmark 10502, and the person to be imaged 10505 are included within the image capture area 1050, and the moving objects, the first automobile 10504, the second automobile 10508, the two bicycles and their drivers 10506, and the other person 10507, who is not to be imaged, have been erased.
[0119] In step S1023, a process is performed to determine whether or not the operation to obtain a temporary processed image is being continuously performed. Specifically, the control unit 1110 determines whether or not the operation to obtain a temporary processed image is being continuously performed. If it is determined that the operation to obtain a temporary processed image is being continuously performed (S1023: Yes), the process returns to step S1023. On the other hand, if it is determined that the operation to obtain a temporary processed image is not being continuously performed (S1023: No), the process proceeds to step S1024.
[0120] In step S1024, a process is performed to determine whether an operation to save an image has been performed. Specifically, the control unit 1110 determines whether an operation to save an image has been performed by the user. An example of an operation to save an image is an operation to fully press the shutter button. If it is determined that an operation to save an image has been performed (S1024: Yes), the process proceeds to step S1025. On the other hand, if it is determined that an operation to save an image has not been performed (S1024: No), the process returns to step S1008.
[0121] In step S1025, a process is performed in which the original image and the processed image are stored in association with each other. Specifically, the control unit 1110 stores the processed image as a temporary processed image and the original image as a through image before processing in the storage unit 1170 or the recording medium 1189 in association with each other. For example, the file name of the processed image includes the image capture number and characters, symbols, numbers, extension, etc. indicating that the image has been processed, and the file name of the original image includes the same image capture number and characters, symbols, numbers, extension, etc. indicating that the image is the original image.
[0122] Furthermore, the control unit 1110 causes the display unit to display image storage report information indicating that the image has been stored. In addition to or instead of displaying the image storage report information, the control unit 1110 may cause the audio output unit 1140 to output the image storage report information as audio, or may cause the audio output unit 1140 to output a predetermined sound. The image storage report information may be text information in natural language, such as, for example, "The original image data, which has not been processed, and the processed image data have been recorded in a format that can be distinguished by a human." Once the image has been stored, the processing proceeds to step S1026.
[0123] In this way, by storing the original image and the processed image in association with each other, the user can compare the original image and the processed image when checking the captured image later, and can select the image to use according to the user's convenience or preference.
[0124] In step S1026, a process is performed to determine whether or not to terminate the AI assistant function. Specifically, the control unit 1110 determines whether or not to terminate the AI assistant function based on whether an operation to terminate the AI assistant function has been performed, whether or not a certain amount of time has elapsed without the user using the AI assistant function, and the like. If it is determined in this determination that the AI assistant function should be terminated, the control unit 1110 terminates the AI assistant function. On the other hand, if it is determined that the AI assistant function should not be terminated, the process returns to step S1004.
[0125] As described above, according to the imaging device of Example 2 of the present invention, a user can obtain hints for taking suitable images when taking images, and can take suitable images without relying on his / her own geographical knowledge (local knowledge), imaging skills, etc. For example, in a place visited for the first time, a user can achieve a suitable composition, a suitable angle of view, and suitable settings for the imaging device even if the user's personal imaging skills are not very high.
[0126] Furthermore, according to the imaging device of Example 2, even if a user has high imaging skills, when taking a selfie, the user may not be able to capture the image as desired because the orientation of the imaging device relative to the subject, the way the imaging device is held, the location where the photographer stands, etc. are different from those in normal imaging. Even in such a case, according to the imaging device of Example 2, the user can take a more suitable image by using the AI assistant function.
[0127] Furthermore, with the imaging device according to the second embodiment, the user can receive advice on imaging in colloquial language, which increases the affinity between the imaging device and the user and allows the user to intuitively understand the content of the advice.
[0128] Furthermore, according to the imaging device of Example 2, by using generation AI, particularly multimodal LLM, it is possible to handle image information and text information simultaneously, and there is no need to switch the communication destination generation AI server depending on the type of information, thereby simplifying the processing of the control unit.
[0129] Furthermore, according to the imaging device of Example 2, by devising the creation of instruction information, the user can receive advice that meets more detailed requests. For example, by creating instruction information such as "Please tell me the optimum imaging conditions for the imaging device without using technical terms," even a beginner who does not understand technical terms can receive advice that the user can understand.
[0130] Furthermore, according to the imaging device of Example 2, it is possible to preserve the sense of realism and memories (memories) of the original image when capturing the image, and also to eliminate (process) any extraneous objects other than the object that is desired to be memorable when capturing the image, thereby making it possible to create a high-quality image without editing work later.
[0131] In Example 2, it is also possible to have the generation AI recognize an object that enters from outside the angle of view as a target subject or as an unwanted subject. In the former case, if you want to take a picture of a subject near the center of the image in the horizontal direction, by setting this in advance, you can request the generation AI to assist you by predicting the timing when the subject will approach the center and instructing you on the timing to release the shutter.
[0132] In addition, in the second embodiment, when the LLM dialogue window and the image display window are displayed on the display unit, if the display unit has only one screen, the windows may be displayed side by side on the single screen, or the windows may be displayed by switching between them. If the display unit has two screens, for example, one on the top and one on the back of the imaging device, the LLM dialogue window may be displayed on one screen, and the image display window may be displayed on the other screen.
[0133] In addition, in Example 2, when an original image and a processed image, or a through image and a model image, are displayed on the display unit, if the display unit has only one screen, these two images may be displayed by switching between them one by one, or two images may be displayed in parallel on one screen. If the display unit has two screens, one image may be displayed on each screen.
[0134] In addition, in Example 2, the preferred composition may be specified as one, or the user may select from multiple. Examples of preferred compositions include a composition in which a landmark fits within the angle of view, a composition in which a subject person fits within the angle of view, a composition in which both a landmark and a subject person fit within the angle of view, and a composition that creates a three-dimensional effect. Furthermore, even if there is one idea of a preferred composition, the generation AI may propose multiple preferred arrangements, etc.
[0135] In addition, in the second embodiment, a function may be realized that allows a user to later reconvert an image processed by the generation AI (by converting a part of the background) into an image desired by the user. A moving object that enters the angle of view from outside may be identified as an unwanted object, and the unwanted object may be removed from the captured image by processing.
[0136] Although the second embodiment is an example in which the assist function using the generation AI is applied to capturing still images, it can also be applied to continuous shooting or video capture. Continuous shooting refers to capturing still images continuously (e.g., at approximately 3 to 30 frames per second) while the shutter button is pressed. Video capture refers to capturing still images at a predetermined frame rate (e.g., 30 to 60 fps).
[0137] In addition, in Example 2, the generated AI may be constructed in a server connected to the imaging device via a network, or may include both constructed in the server and built in the imaging device. That is, depending on various conditions such as the situation or the state of the imaging device or the user, the control unit in the imaging device may transmit and receive data to and from the generated AI constructed in the server, or may transmit and receive data to and from a local generated AI built in the imaging device.
[0138] In addition, in Example 2, the imaging device is, for example, a digital camera, a digital movie camera, a personal computer with an imaging function, a smartphone, a tablet terminal, etc. The imaging device may be, for example, a combination of a digital camera and a smartphone, etc., which are connected to each other so that they can communicate with each other. In this case, for example, an imaging unit included in the digital camera is used as the imaging unit in Example 2, and a control unit included in the digital camera and a control unit included in the smartphone, etc. are used as the control unit in Example 2. Communication between the imaging device and the generation AI server may be performed via a communication unit included in the smartphone, etc.
[0139] Note that the imaging support method based on the processing flow executed in the imaging device and generation AI according to Example 2 is one embodiment of the present invention. Also, a program for causing one or more processors (CPU, MPU, etc.) constituting the control unit of the imaging device according to Example 2 to execute each process so as to implement the imaging support method is also one embodiment of the present invention. Furthermore, a tangible, non-transitory recording medium on which the program is recorded is also one embodiment of the present invention. The program may be downloaded from a server and used.
[0140] The technology according to this embodiment makes it possible to provide a more suitable AI imaging assistant technology. As with the first embodiment, this AI imaging assistant technology is expected to be introduced into higher quality, more reliable infrastructure. Furthermore, the introduction of this technology into infrastructure can contribute to supporting economic development and human welfare, with a focus on affordable and fair access for all. This contributes to the achievement of the United Nations' Sustainable Development Goal (SDG), "Build resilient infrastructure, promote inclusive and sustainable industrialization, innovate and build infrastructure."
[0141] Furthermore, the technology according to this embodiment makes it possible to provide a more suitable AI imaging assistant technology. As with the first embodiment, this AI imaging assistant technology, when introduced into public transportation, can contribute to improving traffic safety through the expansion of public transportation and realizing access to a safe, affordable, and easily accessible sustainable transportation system for all people. This contributes to "Sustainable cities and communities," one of the Sustainable Development Goals (SDGs) advocated by the United Nations.
[0142] Various embodiments have been described above in detail. However, the present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments are detailed descriptions of the entire system in order to clearly explain the present invention, and the present invention is not necessarily limited to a system including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.
[0143] The functions performed by the components described herein may be implemented in circuitry or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), a CPU (a Central Processing Unit), conventional circuits, and / or combinations thereof, programmed to perform the described functions. A processor includes transistors and other circuits and is considered to be circuitry or processing circuitry. A processor may also be a programmed processor that executes programs stored in memory.
[0144] In this specification, a circuitry, unit, or means is hardware that is programmed to realize or performs the described functions, which may be any hardware disclosed herein or any hardware known to be programmed to realize or perform the described functions.
[0145] If the hardware is a processor considered to be a type of circuitry, the circuitry, means, or unit is a combination of the hardware and software used to configure the hardware and / or processor.
[0146] 10010...artificial intelligence response output device, 10011...display unit, 10028...local LLM processing unit, 1107...operation input unit, 1110...control unit, 1132...communication unit, 1140...audio output unit, 1139...microphone, 1160...video control unit, 1170...storage unit, 1180...imaging unit, 1108...non-volatile memory, 1109...memory, 1113...orientation sensor, 1189...recording medium, 1195...current position acquisition unit, 1200...imaging device, 1201...imaging function unit, 19000...internet, 19001...large-scale language model server, 20001...multimodal large-scale language model server
Claims
1. An imaging device comprising an imaging unit, a display unit, and a control unit, wherein the control unit performs the following processes: sending instruction information to a generating artificial intelligence requesting imaging support information to support a user in taking an image using the imaging unit; receiving response information including the imaging support information from the generating artificial intelligence; and displaying the imaging support information included in the received response information or information based on the imaging support information on the display unit.
2. An imaging device according to claim 1, comprising a location information acquisition unit that acquires location information of the imaging device, and the instruction information is information that requests, as the imaging support information, imaging location information that indicates an imaging location where an image of a landmark corresponding to the position indicated by the location information acquired by the location information acquisition unit can be captured.
3. An imaging device according to claim 1, comprising a location information acquisition unit that acquires location information of the imaging device, and the instruction information is information that requests, as the imaging support information, guidance information for guiding the user to an imaging position where an image of a landmark corresponding to the position indicated by the location information acquired by the location information acquisition unit can be captured.
4. An imaging device according to claim 3, wherein the imaging position is a standing or crouching position of the user such that the entire landmark fits within the imaging area.
5. An imaging device according to claim 1, wherein the instruction information is information requesting advice information for advising the user on how to take an image as the imaging support information.
6. An imaging device according to claim 1, wherein the instruction information is information requesting extraction result information representing the result of extracting at least one of a moving object and a person not to be imaged in the captured image obtained by the imaging unit as the imaging support information.
7. An imaging device according to claim 1, wherein the instruction information is information requesting, as the imaging support information, a processed image in which at least one of a moving object and a person not to be imaged has been removed from the captured image obtained by the imaging unit.
8. An imaging device according to claim 7, wherein the control unit executes a process of displaying the original image and the processed image of the obtained captured image on the display unit, and a process of storing the original image and the processed image in a memory unit in association with each other in response to an operation by a user.
9. An imaging device according to claim 1, wherein the control unit detects that the user is spending a certain amount of time or more deciding on a composition, and executes the following processes: a process of inquiring of the user as to whether imaging assistance is required; a process of accepting the user's response to the inquiry as to whether imaging assistance is required; and a process of transmitting the instruction information to the generating artificial intelligence when the response indicating that imaging assistance is required is accepted.
10. An imaging device according to claim 1, wherein the generative artificial intelligence includes a multimodal large-scale language model, and the control unit executes a process of displaying on the display unit information including text information in natural language as the received imaging support information or information based on the imaging support information.
11. An imaging device according to claim 1, wherein the generative artificial intelligence includes a multimodal large-scale language model, and the control unit executes a process of displaying on the display unit information including text information representing an interactive expression as the received imaging support information or information based on the imaging support information.
12. An imaging device according to claim 1, wherein the generating artificial intelligence is constructed in a server connected to the imaging device via a communication network.
13. An imaging device according to claim 1, wherein the generative artificial intelligence includes one constructed on a server connected to the imaging device via a communication network and one built into the imaging device.
14. An imaging support method in which a control unit in an imaging device performs the following processes: sending instruction information to a generating artificial intelligence requesting imaging support information to support a user in taking an image using the imaging unit; receiving response information including the imaging support information from the generating artificial intelligence; and displaying the imaging support information included in the received response information or information based on the imaging support information on a display unit.
15. A program for causing a processor constituting a control unit in an imaging device to execute the following processes: sending instruction information to a generating artificial intelligence requesting imaging support information to assist a user in capturing images using the imaging unit; receiving response information including the imaging support information from the generating artificial intelligence; and displaying the imaging support information included in the received response information or information based on the imaging support information on a display unit.
Citation Information
Patent Citations
Imaging apparatus, information terminal, control method of imaging apparatus, and control method of information terminal
JP2019140561A
Assist device and assist method
JP2022055656A
Scene-aware video dialogue
JP2023510430A
System and method for controlling entity
JP2024035150A
Distributed ultrasound including a wearable ultrasound device
WO2024149612A1