Information processing device, and operation method and operation program for information processing device
By converting user requests into a standardized format and re-training the language model based on feedback, the device enhances the accuracy of interpreting image acquisition requests, addressing the inefficiencies of conventional methods.
Patent Information
- Application Number
- PCT/JP2025/025148
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-29
- Filing Date
- 2025-07-14
- Publication Date
- 2026-02-05
AI Technical Summary
Existing information processing devices struggle to accurately interpret user requests for image acquisition due to variations in linguistic expressions among users, necessitating comprehensive learning of individual user expressions, which is burdensome and inefficient.
The device employs a processor that converts user requests into a standardized format using a first language model, derives image acquisition conditions with a second language model, and re-trains the first model based on user feedback, reducing the need for extensive user-specific learning.
This approach allows for more accurate interpretation of user requests while minimizing the burden of learning diverse linguistic expressions, enabling better alignment with user intentions.
Smart Images

Figure JP2025025148_05022026_PF_FP_ABST
Abstract
Description
Information processing device, operation method and operation program for information processing device
[0001] The technology of the present disclosure relates to an information processing device, an operating method for an information processing device, and an operating program.
[0002] Japanese Patent Application Laid-Open Publication No. 2019-168848 describes an information provision system that includes a feature information generation unit that generates requested feature information indicating user statement information or image features transmitted from a user terminal regarding a desired photograph by the user, and a search unit that acquires photographing information associated with search feature information similar to the requested feature information from a search database that stores photographing information, which is information for photographing an image, in association with search feature information indicating character strings related to an image photographed according to the photographing information or the features of the image itself, and provides the photographing information to the user terminal.
[0003] One embodiment of the technology disclosed herein provides an information processing device, an operating method for an information processing device, and an operating program for an information processing device that are capable of more accurately interpreting user requests than conventional methods, while reducing the burden of comprehensively learning the linguistic expressions for each user, even when the linguistic expressions of requests related to image acquisition differ from user to user.
[0004] The information processing device relating to the technology of the present disclosure includes a processor, which converts a first prompt, which is a user request regarding image acquisition, into a second prompt different from the first prompt using a first language model, derives image acquisition conditions corresponding to the second prompt using the second language model, and re-trains the first language model based on user instructions regarding the image acquisition conditions.
[0005] The processor may present the image acquired based on the image acquisition conditions to the user and accept instructions from the user.
[0006] The processor may be mounted on a camera, and the image acquisition conditions may include shooting conditions of the camera.
[0007] The photographing conditions may include at least one of an F-number, a shutter speed, an ISO sensitivity, and a white balance setting.
[0008] When converting from the first prompt to the second prompt, the processor may accept auxiliary information in addition to the first prompt, and may output the second prompt based on the first prompt and the auxiliary information.
[0009] The processor may be mounted on the camera, the image acquisition conditions may include the shooting conditions of the camera, and the auxiliary information may include at least one of information regarding the image, the shooting environment, the shooting time, and the sound around the camera.
[0010] When deriving image acquisition conditions in response to the second prompt, the processor may accept auxiliary information in addition to the second prompt, and may derive the image acquisition conditions based on the second prompt and the auxiliary information.
[0011] The processor may be mounted on the camera, the image acquisition conditions may include the shooting conditions of the camera, and the auxiliary information may include at least one of information regarding the image, the shooting environment, the shooting time, and the sound around the camera.
[0012] When relearning the first language model, the processor may use, in addition to the user's instructions, at least one of the image recognition results for the image that is the subject of the instructions, the shooting environment, and information regarding the shooting time for the relearning.
[0013] The processor may train the first language model based on user answers input to prepared questions during initial setup of the camera.
[0014] At least one of the first language model and the second language model may be stored in a portable storage device that is detachable from the camera, or in an external storage device that is separate from the camera and that can communicate with the processor.
[0015] In the case where the first language model is stored in a portable storage device that is detachable from the camera or an internal storage device built into the camera, the processor may have the capability to transfer the first language model to another camera.
[0016] If the first language model is stored in a portable storage device that is detachable from the camera or in an internal storage device built into the camera, it may be possible to exclude the first language model from the initialization targets when initializing the internal information of the camera.
[0017] The image acquisition conditions may be image acquisition conditions in a virtual space.
[0018] The image acquisition conditions in the virtual space are the shooting conditions of a virtual camera in the virtual space, and may be conditions that mimic at least one of the F-number, shutter speed, ISO sensitivity, and color temperature included in the shooting conditions of a real camera.
[0019] The image acquisition conditions may be image generation conditions in an image generation model configured using a machine learning model.
[0020] A method for operating an information processing device relating to the technology of the present disclosure is a method for operating an information processing device equipped with a processor, in which the processor converts a first prompt, which is a user request regarding image acquisition, into a second prompt different from the first prompt using a first language model, derives image acquisition conditions corresponding to the second prompt using the second language model, and re-trains the first language model based on user instructions regarding the image acquisition conditions.
[0021] The operating program of an information processing device relating to the technology of the present disclosure is an operating program of an information processing device equipped with a processor, and causes the processor to perform processes including converting a first prompt, which is a user request regarding image acquisition, into a second prompt different from the first prompt using a first language model, deriving image acquisition conditions corresponding to the second prompt using the second language model, and re-learning the first language model based on user instructions regarding the image acquisition conditions.
[0022] According to the technology disclosed herein, even if the linguistic expressions of requests regarding image acquisition differ from user to user, it is possible to more accurately interpret user requests compared to conventional methods while reducing the burden of comprehensively learning the linguistic expressions for each user.
[0023] 24 is a diagram showing an example of a camera as an information processing device. FIG. 25 is a diagram showing an example of the hardware configuration of a camera. FIG. 26 is a diagram showing an example of a functional outline of language processing of a processor. FIG. 27 is a diagram showing a specific example of language processing when relearning is not required. FIG. 28 is a diagram showing a specific example of language processing when relearning is required. FIG. 29 is a diagram showing a specific example of the language processing shown in FIG. 5 after relearning. FIG. 30 is a diagram showing another specific example of language processing when relearning is required. FIG. 31 is a diagram showing a specific example of the language processing shown in FIG. 32 after relearning. FIG. 33 is a diagram showing a main flowchart of a shooting procedure using a camera. FIG. 34 is a flowchart showing the language processing procedure (1 / 2). FIG. 35 is a flowchart showing the language processing procedure (2 / 2). FIG. 36 is a diagram showing the diversity of a user's linguistic expressions. FIG. 37 is another diagram showing the diversity of a user's linguistic expressions. FIG. 38 is a diagram showing a comparative example. FIG. 39 is a diagram showing an overview of the technology of the present disclosure in comparison with FIG. 14. FIG. 39 is a diagram showing the effect of the second embodiment. FIG. 39 is a diagram showing a modification of the second embodiment. FIG. 39 is a diagram showing a third embodiment. FIG. 39 is a flowchart showing the procedure for initial customization of a language model. FIG. 39 is a diagram showing a storage location of a language model. FIG. 39 is a flowchart showing the procedure for initializing internal information. FIG. 39 is an example in which an information processing device is applied to a device that generates a virtual space. FIG. 39 is an example in which an information processing device is applied to an image generation device. FIG. 39 is a diagram showing an example of a functional outline of language processing of the processor of FIG. 3
[0024] First Embodiment The camera 10 shown in FIG. 1 is an example of an information processing device according to the technology of the present disclosure. The camera 10 is capable of verbally accepting a request from a user U regarding image acquisition. The request from the user U regarding image acquisition is, for example, a request regarding the finished image captured by the camera 10. More specifically, as shown in FIG. 1, when the user U takes a photograph of a human subject H with the camera 10, the user verbalizes and speaks the request regarding image acquisition, such as "I want to take an impressive photograph of the person." The camera 10 is capable of accepting a verbal expression of the user U's request via voice. The verbalized request from the user U regarding image acquisition, which is input to the camera 10, is an instruction given by the user U to the camera 10 and is an example of a "first prompt" in the technology of the present disclosure.
[0025] The camera 10 includes a processor 40. When the processor 40 receives the first prompt, it interprets the meaning of the first prompt through language processing and derives image acquisition conditions. The image acquisition conditions include shooting conditions for photographing the subject H with the camera 10. The shooting conditions include at least one of an F-number, a shutter speed, an ISO sensitivity, and a white balance setting. The shooting conditions may include an exposure setting mode such as a shutter speed priority mode or an aperture priority mode. The image acquisition conditions may also include image processing conditions after photographing in addition to the shooting conditions set on the camera 10 during photographing. The image processing conditions include various settings such as a film simulation mode that simulates the color characteristics of photographic film by image processing on the digital data of the image, a high dynamic range (HDR) mode that relatively widens the dynamic range of the photographed image, and gradation processing that determines the density contrast of the photographed image. The processor 40 sets the derived image acquisition conditions and controls photographing based on the set image acquisition conditions.
[0026] As an example, the camera 10 is provided with a display 15 and a live view display function that displays captured images on the display 15 in real time. Using the live view display function, the camera 10 presents images captured based on the derived image acquisition conditions to the user U as live view images LV. The user U checks the live view images LV, and if there are any requests regarding the finished image, the user U inputs a first prompt, which is a linguistic expression of the request, to the camera 10. Alternatively, if the finished live view images LV do not match the user U's intentions, the user U inputs an instruction to change the image acquisition conditions to the camera 10. Naturally, the camera 10 can also accept such an instruction to change the image acquisition conditions. Then, the camera 10 executes actual capture control based on the finally set image acquisition conditions. Actual capture is a concept distinct from live view capture, and unlike live view images LV, which are sequentially erased after display, actual capture is a type of capture in which images are recorded in a state that allows them to be played back after capture.
[0027] Fig. 2 shows an example of the internal configuration of the camera 10. As shown in Fig. 2, the camera 10 is, as an example, an interchangeable lens digital camera. The camera 10 is composed of a main body 11 and an imaging lens 12 that is interchangeably attached to the main body 11. The imaging lens 12 is attached to the front side of the main body 11 via a camera-side mount 11A and a lens-side mount 12A.
[0028] The main body 11 is provided with an operation unit that includes an operation device 13 such as a dial and a release button, as well as a display 15 with a touch panel function. These operation units accept operations by the user. These operation units enable the setting of various operation modes of the camera 10 as well as the setting of image acquisition conditions, including shooting conditions. The operation modes of the camera 10 include, for example, a still image shooting mode, a video shooting mode, and an image display mode.
[0029] In addition, in the still image shooting mode, modes for setting image acquisition conditions including shooting conditions include a prompt acceptance mode in which the first prompt can be input, and a normal mode in which the first prompt cannot be input. In the normal mode, acceptance of the first prompt is disabled, but image acquisition conditions can be set by manual operation or automatic setting.
[0030] The display 15 is used to play back and display a live view image LV and a captured actual image MP as captured images, and also to display various setting screens.
[0031] The main body 11 is also provided with a communication I / F (Interface) 16, a microphone 17, and a speaker 18. The communication I / F 16 controls transmission of wireless or wired communication with external devices. The microphone 17 is used to record audio during video capture and to input a first prompt issued by the user U. The speaker 18 is used to output audio during video playback and to output warning sounds, etc.
[0032] The main body 11 and the imaging lens 12 are electrically connected by electrical contacts 11B provided on the camera-side mount 11A coming into contact with electrical contacts 12B provided on the lens-side mount 12A.
[0033] The imaging lens 12 includes an objective lens 30, a focus lens 31, a rear end lens 32, and an aperture 33. The components are arranged along the optical axis A of the imaging lens 12 in the following order from the objective side: objective lens 30, aperture 33, focus lens 31, and rear end lens 32. The objective lens 30, focus lens 31, and rear end lens 32 constitute an imaging optical system. The type, number, and arrangement order of the lenses that make up the imaging optical system are not limited to the example shown in FIG. 2 .
[0034] The imaging lens 12 also has a lens driver 34. The lens driver 34 is configured with, for example, a CPU (Central Processing Unit), RAM (Random Access Memory), and ROM (Read Only Memory). The lens driver 34 is electrically connected to a processor 40 in the main body 11 via electrical contacts 12B and 11B.
[0035] The lens driving unit 34 drives the focus lens 31 and the diaphragm 33 based on a control signal transmitted from the processor 40. The lens driving unit 34 controls the driving of the focus lens 31 based on a control signal for focus control transmitted from the processor 40 in order to adjust the focus position of the imaging lens 12. The processor 40 performs focus position detection using a phase difference method, for example.
[0036] The diaphragm 33 has an aperture whose diameter is variable around the optical axis A. The lens driver 34 controls the drive of the diaphragm 33 based on an aperture adjustment control signal sent from the processor 40 to adjust the amount of light incident on the light receiving surface 20A of the image sensor 20. Controlling the diaphragm 33 adjusts the effective aperture of the imaging lens 12, and adjusts the F-number, which is the focal length of the imaging lens 12 divided by the effective aperture. The amount of incident light increases as the F-number becomes smaller (i.e., the image becomes brighter), and decreases as the F-number becomes larger (i.e., the image becomes darker). Furthermore, the depth of field, which represents the range in focus along the optical axis (i.e., the depth direction), decreases as the F-number becomes smaller (i.e., the range in focus is narrower), and decreases as the F-number becomes larger (i.e., the range in focus is wider).
[0037] The main body 11 also contains an image sensor 20, a processor 40, and a memory 42. The processor 40 controls the lens driver 34, the image sensor 20, the memory 42, the operation device 13, the display 15, the communication I / F 16, the microphone 17, and the speaker 18.
[0038] The processor 40 is configured, for example, by a CPU. In this case, the processor 40 executes various processes based on a program 43 stored in the memory 42. The processor 40 may be configured by a collection of multiple pieces of hardware, for example, multiple integrated circuit (IC) chips. The program 43 is an example of an "operating program" according to the technology of the present disclosure. The memory 42 is configured, for example, by at least one of various storage devices such as a RAM, a flash memory, or a hard disk drive. The memory 42 may also include a ROM.
[0039] The image sensor 20 is, for example, a CMOS (Complementary Metal Oxide Semiconductor) image sensor. The image sensor 20 is positioned so that its optical axis A is perpendicular to the light-receiving surface 20A and is located at the center of the light-receiving surface 20A. Light that has passed through the imaging lens 12 is incident on the light-receiving surface 20A. A plurality of pixels are formed on the light-receiving surface 20A, and each pixel generates a signal by performing photoelectric conversion. The image sensor 20 outputs an image signal by photoelectrically converting the light incident on each pixel.
[0040] The memory 42 also stores a language model 44. The language model 44 is a machine learning model that recognizes text or speech expressed in a natural language and understands its meaning. For example, the language model 44 is based on large language models such as Generative Pre-trained Transformers (GPT) and Bidirectional Encoder Representations from Transformers (BERT). When a first prompt expressed in a natural language is input, the language model 44 interprets the meaning of the first prompt and derives image acquisition conditions.
[0041] FIG. 3 shows an overview of the language processing functions of the processor 40. The processor 40 performs language processing in accordance with a program 43 stored in the memory 42. The processor 40 performs language processing using a language model 44. The language model 44 includes a first language model 44A and a second language model 44B. The first language model 44A is used to convert a first prompt into a second prompt that is different from the first prompt. The processor 40 converts the first prompt into the second prompt using the first language model 44A. The second language model 44B is used to derive image acquisition conditions corresponding to the second prompt. The processor 40 derives image acquisition conditions corresponding to the second prompt using the second language model 44B. As described above, the image acquisition conditions include shooting conditions.
[0042] Furthermore, if the user U instructs the processor 40 to change the image acquisition conditions derived from the second prompt and the actual shooting is performed under the changed image acquisition conditions, the processor 40 performs a re-learning process of the first language model 44A based on the image acquisition conditions of the actual shooting in addition to the first prompt.
[0043] 4 to 8, the language processing of the processor 40 will be described using an example in which a linguistic expression such as "I want to take an impressive photograph of a person" is input to the first language model 44A as a first prompt. As shown in FIG. 4, the first language model 44A converts the first prompt "I want to take an impressive photograph of a person" into a second prompt such as "I want to make the person stand out." When the second prompt is input to the second language model 44B, the second language model 44B derives a condition SC1 as an image acquisition condition, which is "open the aperture by one stop while focusing on the person." Opening the aperture by one stop corresponds to reducing the F-number. Reducing the F-number results in a shallower depth of field. For example, when a person is in focus, the background behind the person will be out of focus, resulting in a blurred background.
[0044] Here, for convenience, the image acquisition conditions are described in a simplified manner, but as mentioned above, there are various image acquisition conditions. In reality, the processor 40 sets various image acquisition conditions other than the F-number based on the image to be captured. For example, to achieve a standard exposure, the processor 40 sets the ISO sensitivity and shutter speed other than the F-number. The processor 40 also sets the white balance and other settings depending on the content of the first prompt.
[0045] The processor 40 displays a live view image LV captured by applying the derived condition SC1. If the user U is satisfied with the live view image LV captured by applying this condition SC1 as the intended result, the user inputs a command to perform actual photography by pressing the release button without changing the condition SC1. This causes actual photography to be performed, and an actual image MP is acquired in which the background is blurred and the person in the subject H stands out. Specific example 1 shown in Figure 4 is a case in which the user U does not issue a command to change the image capture conditions derived by the second language model 44B, and in this case, re-learning of the first language model 44B is not necessary.
[0046] Specific Example 2-1 shown in FIG. 5 is similar to Specific Example 1 shown in FIG. 4 in the process from converting the first prompt to the second prompt, deriving condition SC1 based on the second prompt, and displaying the live view image LV based on the derived condition SC1. In Specific Example 2-1 shown in FIG. 5, user U checks the live view image LV, feels that the finished image differs from his or her intention, and issues a change instruction to change the image acquisition conditions. The change instruction, for example, is to return the aperture to its original setting and adjust the white balance (WB) setting to increase redness. In Specific Example 2-1, user U's intention in the first prompt is to increase the redness of the subject H's skin to make the subject more striking. White balance is used to adjust the color temperature, which varies depending on the type of light source and weather, but it can also be used to change the color tone of the image. If returning the aperture to its original setting does not result in a standard exposure, the exposure is adjusted by adjusting the ISO sensitivity or shutter speed.
[0047] Although not shown in FIG. 5, when the image acquisition conditions are changed, the image acquisition conditions that are applicable to the live view image LV are applied, and the live view image LV is updated.
[0048] If the user U is satisfied with the finished image after the changes, the user U inputs a shooting instruction for the actual shooting under the changed image acquisition conditions. As a result, the actual shooting is performed and an actual image MP in which the skin has an increased reddish tinge is acquired.
[0049] When the image acquisition conditions for the actual photographing are changed based on an instruction from the user U to change the image acquisition conditions derived by the second language model 44B in this way, the processor 40 re-trains the first language model 44A. This is an example of the technique disclosed herein in which the processor 40 re-trains the first language model 44A based on an instruction from the user regarding the image acquisition conditions. Specific example 2-1 shown in FIG. 5 is a case in which re-training of the first language model 44A is necessary.
[0050] The reason for relearning the first language model 44A is as follows. Specific example 2-1 shown in FIG. 5 is considered to be a case in which, before relearning, the image finish intended by the user U in the first prompt, "I want to take an impressive photograph of a person," does not match the image finish obtained by the processor 40 interpreting the first prompt using the first language model 44A and the second language model 44B. Therefore, the first language model 44A is relearned so that the processor 40 can accurately interpret the intention conveyed by the user U in the first prompt, "I want to take an impressive photograph of a person." The relearning uses the first prompt and image acquisition conditions corresponding to a change instruction from the user U. The processor 40 performs relearning so that when the first prompt, "I want to take an impressive photograph of a person," is input to the first language model 44A, it is converted into the second prompt, "Increase the redness of the person's skin."
[0051] Specific example 2-2 shown in FIG. 6 illustrates language processing by the processor 40 after relearning shown in specific example 2-1. After relearning, if the user U inputs a first prompt such as "I want to take an impressive photo of a person," the first language model 44A converts the input first prompt into a second prompt such as "Increase the redness of the person's skin." The second language model 44B then derives a condition SC2, "Adjust the white balance so as to increase the redness of the skin," based on the second prompt. A live view image LV to which the condition SC2 has been applied is then displayed. By relearning in this manner, the processor 40 can accurately interpret the meaning of the first prompt initially input by the user U.
[0052] Specific Example 3-1 shown in FIG. 7 is similar to Specific Example 1 shown in FIG. 4 in the process from converting the first prompt to the second prompt, deriving condition SC1 based on the second prompt, and displaying the live view image LV based on the derived condition SC1. In Specific Example 3-1 shown in FIG. 7 , similar to Specific Example 2-1 shown in FIG. 5 , user U checks the live view image LV, feels that the finished image differs from his or her intention, and issues a change instruction to change the image acquisition conditions. The changes in Specific Example 3-1 differ from those in Specific Example 2-1, and include, for example, returning the aperture to its original position and turning on the flash. User U's intention with this change instruction is to use flash photography to illuminate the face of the person in subject H and emphasize the contrast between the face and the background, thereby making the person stand out. Because the condition of turning on the flash cannot be applied to the live view image LV, for example, an indicator 51 indicating that the flash should be turned on is displayed on the live view image LV (see FIG. 8 ).
[0053] If the user U is satisfied with the finished image after the changes, he / she inputs a shooting instruction for the actual shooting under the changed image acquisition conditions. This causes the actual shooting to be performed, and flash photography is performed during the actual shooting. The actual image MP with the face brightened by the flash photography is acquired.
[0054] 7, similarly to the specific example 2-1 shown in Fig. 5, the image acquisition conditions for the actual shooting are changed based on a change instruction from the user U for the image acquisition conditions derived by the second language model 44B, and therefore the processor 40 re-learns the first language model 44A. The re-learning shown in Fig. 7 is also an example of the technique disclosed herein in which the processor 40 re-learns the first language model 44A based on a user instruction for the image acquisition conditions.
[0055] The reason for relearning the first language model 44A in Specific Example 3-1 is the same as in Specific Example 2-1: to enable the processor 40 to accurately interpret the intention conveyed by the user U in the first prompt, "I want to take an impressive photograph of the person." As described above, the relearning uses the first prompt and image acquisition conditions according to a change instruction from the user U. When the first prompt, "I want to take an impressive photograph of the person," is input to the first language model 44A, the processor 40 performs relearning so that the first prompt, "I want to take an impressive photograph of the person," is converted into the second prompt, "I will brighten the face by using a flash."
[0056] Specific example 3-2 shown in FIG. 8 illustrates language processing by processor 40 after relearning shown in specific example 3-1. After relearning, if user U inputs a first prompt such as "I want to take an impressive photo of a person," first language model 44A converts the input first prompt into a second prompt such as "I want to brighten the face by using flash photography." Then, second language model 44B derives condition SC3, "Turn on the flash," based on the second prompt. Then, indicator 51 indicating that flash photography of condition SC3 will be applied during actual photography is displayed on live view image LV. By relearning in this manner, processor 40 can accurately interpret the meaning of the first prompt initially input by user U.
[0057] The operation of the above configuration will be described below with reference to the flowcharts shown in Figures 9 to 11. Figure 9 is a main flowchart outlining the photographing process of camera 10, and Figures 10 and 11 are sub-flowcharts showing the process of the prompt reception mode in which the first prompt can be used.
[0058] As shown in FIG. 9 , when the power of the camera 10 is turned on in step ST1100, the process proceeds to step ST1200, where the processor 40 starts live view display. If the user U does not select the prompt reception mode in step ST1300 (N in step ST1300), the processor 40 proceeds to normal mode in step ST1400. In normal mode, the user U sets image acquisition conditions, including shooting conditions, by manually operating the operation device 13 or the like. Alternatively, by selecting automatic setting, the setting of the image acquisition conditions is left to the camera 10. The user U checks the composition and angle of view using the live view image LV displayed on the display 15 and determines the composition and angle of view. In this state, the user issues an instruction to perform actual shooting using the release button. If an instruction to perform actual shooting is input in step ST1500 (Y in step ST1500), the processor 40 proceeds to step ST1600, where the processor 40 executes the shooting process for the actual shooting. If the normal mode is selected, processor 40 repeats the above process until the power is turned off in step ST1700.
[0059] On the other hand, if the prompt reception mode is selected in step ST1300 (Y in step ST1300), the process proceeds to step ST2010 shown in FIG.
[0060] 10, the processor 40 waits for input of a first prompt from the user U. As shown in FIGS. 4 to 8, when the first prompt is input by the user U speaking, the process proceeds to step ST2020, where the processor 40 interprets the meaning of the first prompt using the first language model 44A and converts it into the linguistic expression of a second prompt that is different from the first prompt.
[0061] Next, the process proceeds to step ST2030, where processor 40 derives image acquisition conditions based on the second prompt using second language model 44B. In step ST2040, processor 40 updates the shooting conditions for the live view image LV based on the image acquisition conditions derived from the second prompt. As a result, in step ST2050, the live view image LV is also updated.
[0062] In step ST2060, the processor 40 waits for the user U to input an instruction to change the image acquisition conditions derived from the second prompt by the second language model 44B during the live view display. As shown in Figures 5 and 7, if the image acquisition conditions derived from the second prompt do not match the image acquisition conditions intended by the user U, the user U inputs an instruction to change the image acquisition conditions.
[0063] If the user U issues an instruction to change the image acquisition conditions in step ST2060 (Y in step ST2060), the processor 40 returns to step ST2040, updates the image acquisition conditions and the live view image LV, and proceeds to step ST2060 again. If the user U does not issue an instruction to change the image acquisition conditions in step ST2060 (N in step ST2060), the processor 40 proceeds to step ST2070. If an instruction to perform actual photography is input by pressing the release button in step ST2070, the processor 40 proceeds to step ST2080 and executes actual photography. Subsequently, the processor 40 proceeds to step ST2090 shown in FIG. 11 .
[0064] In step ST2090 shown in FIG. 11 , the processor 40 determines whether the user U has issued a change instruction to the image acquisition conditions derived by the second language model 44B from the second prompt during the actual shooting. If a change instruction has been issued (Y in step ST2090), the processor 40 proceeds to step ST2100 and retrains the first language model based on the change instruction. As a result of this retraining, as shown in FIGS. 6 and 8 , the processor 40 can use the retrained first language model 44A to accurately interpret the meaning intended by the user U in the linguistic expression of the first prompt entered by the user U. This enables the processor 40 to derive image acquisition conditions consistent with the user U's intention, allowing the user U to acquire the actual image MP consistent with his or her intention.
[0065] Thereafter, in step ST2110, processor 40 determines whether the selection of the prompt reception mode is finished, and if an end instruction is given, returns to step ST1400 in Fig. 9. If an end instruction is not given, processor 40 proceeds to step ST2120 and repeats the processing from step ST2010 onwards until the power is turned off.
[0066] As described above, camera 10, shown as an example of an information processing device of the technology of the present disclosure, includes processor 40. Processor 40 converts a first prompt, which is a request from user U regarding image acquisition, into a second prompt that is different from the first prompt using first language model 44A, derives image acquisition conditions corresponding to the second prompt using second language model 44B, and re-learns first language model 44A based on the user's instructions regarding the image acquisition conditions. This reduces the burden of comprehensively learning linguistic expressions for each user, even when the linguistic expressions of requests regarding image acquisition differ from user to user, and enables more accurate interpretation of user requests compared to conventional methods.
[0067] That is, when a user U verbalizes the finished image, if the user U is a professional photographer or other user familiar with camera setting conditions and photography techniques, he or she is familiar with expressions that could be called standard language in that field. Therefore, when the same finish is intended, the same linguistic expression is often used. However, when a user U is not familiar with camera setting conditions and photography techniques, the linguistic expression may differ from user U even when the same finish is intended, or conversely, the same linguistic expression may mean different finishes for different users U.
[0068] For example, as shown in the above embodiment and in Figure 12, even if the first prompt is expressed in the same language as "I want to take an impressive photo of a person," the intention may differ for each of users U1, U2, and U3. User U1's intention is closer to the intention expressed in the language expression "to make the person stand out," user U2's intention is closer to the intention expressed in the language expression "to increase the redness of the person's skin," and user U3's intention is closer to the intention expressed in the language expression "to brighten the face by using a flash."
[0069] 13, even if the linguistic expressions of the first prompt are different, the intention may be the same among users U1, U2, and U3. In the example shown in FIG. 13, the linguistic expression of user U1's first prompt is "I want to take an impressive photo of the person," and the linguistic expression of user U2's first prompt is "I want to take a cool photo of the person." Furthermore, the linguistic expression of user U3's first prompt is "I want to make the person shine." However, the intention implied in each of these linguistic expressions by users U1, U2, and U3 is the common linguistic expression of "I want to make the person stand out."
[0070] As such, the language expressions of users U are diverse and have individual personalities and tendencies. If the language model 44 is configured using only the second language model 44B, as in the comparative example shown in Fig. 14, it is necessary to have the second language model 44B comprehensively learn the language expressions of each user U as shown in Fig. 12 and Fig. 13. However, it is extremely difficult to have a single second language model 44B learn all of the diverse language expressions of multiple users U, and the learning load is also extremely large.
[0071] Therefore, in the technology disclosed herein, as shown in FIG. 15 , the language model 44 is separated into two, a first language model 44A and a second language model 44B, and the first language model 44A can be customized for each user U based on their own language expression. The second language model 44B is standardized, while the first language model 44A is customized for each user U. As a result, for example, the first language model 44A customized for each user U can convert a first prompt with a variety of language expressions for each user U into a second prompt with a standard language expression. Meanwhile, the second language model 44B is trained so that it can derive corresponding image acquisition conditions based on the second prompt with a standard language expression. In this way, by using the first language model 44A that can be customized to fit the language expression for each user U, the load of the training process for the second language model 44B can be reduced.
[0072] In the above embodiment, the processor 40 presents the user U with a live view image LV (an example of an image) acquired based on the image acquisition conditions, and accepts a change instruction (an example of an instruction) from the user U. This allows the user U to check the live view image LV to which the image acquisition conditions have been applied, making it easier for the user U to determine whether the image acquisition conditions are in line with their intentions.
[0073] In the above embodiment, the processor 40 is mounted on the camera 10, and the image acquisition conditions include the shooting conditions of the camera 10. This makes it easy to reflect the user U's requests regarding image acquisition when taking a picture with the camera 10. The shooting conditions also include at least one of the F-number, shutter speed, ISO sensitivity, and white balance setting. These are typical items of shooting conditions, making it easy to specify the finished image.
[0074] Although the above embodiment shows an example in which the first prompt is received by voice, the first prompt may be received as text data instead of voice, or as an image of handwritten characters.
[0075] Furthermore, although the example has been shown in which the instruction by the user U to change the image acquisition conditions derived from the second prompt is received by manual operation of the operation device 13, the instruction to change may also be received by voice.
[0076] Furthermore, the first language model 44A has a function of converting a first prompt into a second prompt that is different from the first prompt, but may end up converting the first prompt into a second prompt that has the same linguistic expression as the first prompt. Furthermore, in the initial state, the first language model 44A may output the first prompt as the second prompt as is.
[0077] In the above embodiment, the information processing device according to the technology of the present disclosure has been described using an interchangeable lens digital camera as an example, but the technology of the present disclosure can also be applied to a digital camera with an integrated lens. It can also be applied to a digital camera built into a smart device, etc. In this way, the technology of the present disclosure can be applied to various imaging devices.
[0078] 16 and 17 , when the processor 40 of the camera 10 converts a first prompt into a second prompt using a first language model 44A, the processor 40 can accept auxiliary information in addition to the first prompt and outputs the second prompt based on the first prompt and the auxiliary information. In the first embodiment, when the first language model 44A is used to convert a first prompt into a second prompt, an example was shown in which only the first prompt was input to the first language model 44A. In contrast, the second embodiment is an example in which auxiliary information is input to the first language model 44A in addition to the first prompt.
[0079] The auxiliary information includes at least one of information related to an image, a shooting environment, a shooting time, and surrounding audio. The image is, for example, a live view image LV. The shooting environment includes the shooting location and weather. The shooting environment is detected, for example, by image analysis of the live view image LV. Alternatively, the shooting location may be detected based on GPS (Global Positioning System) information. The shooting time may be a time of day or a time period such as daytime or nighttime. The surrounding audio is audio collected by the camera 10 through the microphone 17.
[0080] In addition to the first prompt, such auxiliary information is input to the first language model 44A. Therefore, compared to when only the first prompt is used, it is easier to convert the first prompt into an appropriate second prompt depending on the shooting environment and shooting time. As a result, it is easier to derive appropriate image acquisition conditions depending on the shooting environment and shooting time.
[0081] For example, as shown in FIG. 17 , consider a case where the shooting time is input as auxiliary information in addition to the first prompt, "I want to take an impressive photograph of the person." As conceptually shown in FIG. 17 , if the shooting time is not at night, the processor 40 inputs the first prompt, "I want to take an impressive photograph of the person," and the auxiliary information, "The shooting time is not at night," into the language model 44 including the first language model 44A, and converts it into a second prompt, "I want to take an impressive photograph of the person." This derives a condition SC1 for blurring the background. On the other hand, if the shooting time is at night, the processor 40 inputs the first prompt, "I want to take an impressive photograph of the person," and the auxiliary information, "The shooting time is at night," into the language model 44 including the first language model 44A, and converts it into a second prompt, "I will brighten the face with flash photography." This derives a condition SC3 for taking flash photography. By using the auxiliary information in this way, the processor 40 can convert the same first prompt into a different second prompt depending on the auxiliary information.
[0082] Although a live view image LV captured by the camera 10 has been given as an example of an image used as auxiliary information, it need not be a live view image LV. For example, a reference image that matches the user U's intended finish may be used. The reference image may be a captured image that has already been saved, or an image downloaded from the Internet. An example of a method for using an image used as auxiliary information is to analyze the content of the subject (e.g., a person or a landscape), composition, and angle of view of the image using image analysis processing such as image recognition processing, and then input the analysis results into the first prompt. While the image analysis processing performed by the processor 40 is omitted in FIG. 16 , such image analysis processing (see FIG. 19 ) is required when using an image as auxiliary information.
[0083] The user U may also input auxiliary information by voice. For example, the user U may input information such as "forest," "ocean," "mountain," and "city" by voice as information representing the shooting environment. Furthermore, the user U may input information such as "people," "dogs," and "scenery" by voice as information related to the subject. Furthermore, this information may be input as text instead of voice.
[0084] The processor 40 shown in FIG. 18 is a modified example of the second embodiment. When deriving image acquisition conditions from a second prompt using a second language model 44B, the processor 40 can accept auxiliary information in addition to the second prompt and derives image acquisition conditions based on the second prompt and the auxiliary information. That is, the example shown in FIG. 18 is an example in which auxiliary information is input to the second language model 44B. An example of the auxiliary information is the same as that shown in FIG. 16 . By inputting auxiliary information to the second language model 44B in addition to the second prompt in this manner, the same effect as that shown in FIG. 17 can be obtained. That is, the camera 10 shown in FIG. 18 can derive appropriate image acquisition conditions according to the shooting environment, shooting time, and the like, compared to when only the second prompt is input.
[0085] "Third Embodiment" In the third embodiment shown in FIG. 19 , when relearning the first language model 44A, the processor 40 uses, in addition to an instruction from the user U (for example, a change instruction), at least one of the image recognition results of the image (for example, the main image MP) that is the target of the instruction, the shooting environment, and information about the shooting time. That is, in the third embodiment, when relearning the first language model 44A, in addition to the image acquisition conditions at the time of the main image capture that reflect the user U's change instruction to the image acquisition conditions derived from the second prompt, the above information other than the image acquisition conditions is used. Such relearning is effective when auxiliary information is used in the first language model 44A. The effect of the third embodiment is similar to that of the second embodiment shown in FIG. 17 . Compared to relearning based only on the user U's change instruction, the first language model 44A can convert the first prompt into a more appropriate second prompt depending on the shooting time, etc.
[0086] "Modifications" Furthermore, the technology of the present disclosure can be modified in various ways as described below.
[0087] (Initial Customization of First Language Model) As shown in FIG. 20 , during initial setup of the camera 10, the processor 40 may train the first language model 44A based on answers input by the user U to questions prepared in advance. For example, the camera 10 is capable of initial customization of the first language model 44A during initial setup. When the initial customization of the first language model 44A begins, the processor 40 presents a question to the user U. This question is for adapting the first language model 44A to the linguistic expression tendencies of the user U. For example, a prepared image may be presented, and the question may ask how the user U would verbalize the finished image. The processor 40 has the user answer this question, and if there is an answer, updates the first language model 44A based on the answer. Multiple such questions may be asked. When the questions are finished, the initial customization ends.
[0088] By performing initial customization on the first language model 44A in this way, the learning period for the user U's language expressions can be shortened compared to when there is no initial customization.
[0089] (Storage location of language model) As shown in Figure 21, at least one of the first language model 44A and the second language model 44B may be stored in a portable storage device that can be attached and detached to the camera 10, or in an external storage device separate from the camera 10 that can communicate with the processor 40.
[0090] As described above, the language model 44 is customized for the user U. Therefore, for example, when the user U replaces the camera 10 or uses multiple cameras 10, it is preferable that the language model 44 customized for the user U can be easily transferred to another camera 10. For this reason, it is preferable that the language model 44 be stored in a portable storage device such as a memory card 71. Alternatively, it is preferable that the language model 44 be stored in an external storage device 72 that can communicate with the processor 40 of the camera 10, such as an external hard disk drive or a smart device other than the camera 10. In this way, compared to storing the language model 44 in the built-in memory 42, it is possible to easily transfer the customized language model 44 to a new camera 10 that replaces an old camera 10 that has been used previously, or to multiple cameras 10.
[0091] Of the language models 44, it is considered that the first language model 44A is often the one that is customized by the user U, so it is preferable to store at least the first language model 44A of the language models 44 in a portable storage device or an external storage device 72.
[0092] Furthermore, the advantage of storing the second language model 44B of the language models 44 in a portable storage device or external storage device 72 is that, for example, the second language model 44B, which is used in combination with the first language model 44A, can be migrated together. If the version of the second language model 44B that can be used in combination with the first language model 44A is limited, it is convenient to be able to easily migrate the second language model 44B in addition to the first language model 44A. Furthermore, it may be convenient to be able to easily migrate the second language model 44B, for example, when only the second language model 44B is updated by the manufacturer.
[0093] Furthermore, when the first language model 44A is stored in a portable storage device that is detachable from the camera 10 or an internal storage device (for example, the memory 42) built into the camera 10, the processor 40 may have a function of transferring the first language model 44A to another camera 10. As shown in Fig. 21 , the processor 40 of the camera 10 transfers the first language model 44A to another camera 10 via the communication I / F 16. This also allows the customized first language model 44A to be easily transferred to another camera 10.
[0094] 22 , in a case where the first language model 44A is stored in a portable storage device that is detachable from the camera or an internal storage device (for example, the memory 42) built into the camera 10, it may be possible to exclude the first language model 44A from being initialized when initializing the internal information of the camera 10. This has the advantage that update information of the first language model 44A customized by the user U is not lost when the internal information of the camera 10 is initialized.
[0095] 22 , when initializing the internal information, the processor 40 can accept a designation to exclude language models 44 including the first language model 44A from the initialization targets. When the designation to exclude the language model 44 is made, the processor 40 initializes the internal information excluding the language model 44. On the other hand, when the designation to exclude the language model 44 is not made, the processor 40 initializes the internal information including the language model 44.
[0096] (Virtual Camera) In the above embodiment, an example has been described in which the information processing device according to the technology of the present disclosure is applied to a real camera 10. However, as shown in FIG. 23 , the information processing device may also be applied to a case in which image capture is performed in a virtual space. That is, the "image acquisition conditions" according to the technology of the present disclosure may be image acquisition conditions in a virtual space. For example, as shown in FIG. 23 , a VR (Virtual Space) headset 81, which is a device for generating a virtual space, may be the information processing device according to the technology of the present disclosure.
[0097] As shown in FIG. 23 as an example, in a virtual space generated by a headset 81, a subject in the virtual space can be photographed using a virtual camera 83 in the virtual space. The processor of the headset 81 or the controller 82 can receive a first prompt, such as "I want to take an impressive photograph of a person." This processor is similar to the processor 40 in each of the above-described embodiments. The processor then converts the first prompt into a second prompt using a language model 44 and derives image acquisition conditions from the second prompt. The image acquisition conditions thus derived are applied to the virtual camera 83. The image acquisition conditions in the virtual space are the shooting conditions of the virtual camera 83 in the virtual space, and are conditions that mimic at least one of the F-number, shutter speed, ISO sensitivity, and color temperature included in the shooting conditions of the real camera 10.
[0098] This allows the virtual camera 83 in the virtual space to capture an image that meets the intention of the user U. The embodiment shown in Fig. 23 also has the same effect as the above-described embodiments, that is, even when the linguistic expressions of requests for image acquisition differ from user to user, it is possible to more accurately interpret the user's request than in the past while reducing the burden of comprehensively learning the linguistic expressions for each user.
[0099] (Image Generation Model) Furthermore, as shown in FIGS. 24 and 25 , the information processing device according to the technique of the present disclosure can also be applied to an image generation device 91 that uses an image generation model configured using a machine learning model. In this case, the image acquisition conditions are image generation conditions in the image generation model. The image generation device 91 is, for example, a personal computer on which the image generation model is installed. A user U generates an image using the image generation device 91. The image generation device 91 is capable of receiving a first prompt from the user U. Then, similar to the processor 40 of the camera 10 in each of the above embodiments, the processor 40 of the image generation device 91 interprets the received first prompt using language processing and derives image generation conditions as image acquisition conditions. The processor 40 sets the derived image generation conditions and performs image generation.
[0100] As shown in Figure 25, the processor 40 of the image generation device 91 performs language processing and re-learning processing similar to that of the processor 40 of the camera 10 shown in Figure 3. The processing in Figure 25 differs from that shown in Figure 3 in that there is no concept of photography, and an image is simply generated. In Figure 25 as well, if a user issues a change instruction and the image acquisition conditions derived from the second prompt do not match the final image acquisition conditions, the first language model 44A is re-learned. This reduces the burden of comprehensively learning the language expressions for each user, even when the language expressions used to express requests for image acquisition vary from user to user, and enables more accurate interpretation of user requests compared to conventional methods.
[0101] 24 has been described using an example of a personal computer in which the image generating device 91 is directly operated by the user U, the image generating device 91 may be a server present on a network. The image generating device 91 may also be realized in the form of a mobile device such as a smartphone.
[0102] The above description can be used to understand the technologies described in the following supplementary items. [Supplementary Item 1] An information processing device including a processor, wherein the processor converts a first prompt, which is a user request regarding image acquisition, into a second prompt different from the first prompt using a first language model, derives image acquisition conditions corresponding to the second prompt using the second language model, and re-learns the first language model based on an instruction from the user regarding the image acquisition conditions. [Supplementary Item 2] The information processing device according to Supplementary Item 1, wherein the processor presents to the user an image acquired based on the image acquisition conditions, and accepts an instruction from the user. [Supplementary Item 3] The information processing device according to Supplementary Item 1 or Supplementary Item 2, wherein the processor is mounted on a camera, and the image acquisition conditions include shooting conditions of the camera. [Supplementary Item 4] The information processing device according to Supplementary Item 3, wherein the shooting conditions include at least one of F-number, shutter speed, ISO sensitivity, and white balance setting. [Supplementary Item 5] The information processing device of any one of Supplementary Items 1 to 4, wherein, when converting from a first prompt to a second prompt, the processor is capable of accepting auxiliary information in addition to the first prompt, and outputs the second prompt based on the first prompt and the auxiliary information. [Supplementary Item 6] The information processing device of Supplementary Item 5, wherein the processor is mounted on a camera, the image acquisition conditions include shooting conditions of the camera, and the auxiliary information includes at least one of information on the image, the shooting environment, the shooting time, and audio around the camera. [Supplementary Item 7] The information processing device of any one of Supplementary Items 1 to 6, wherein, when deriving image acquisition conditions according to the second prompt, the processor is capable of accepting auxiliary information in addition to the second prompt, and derives the image acquisition conditions based on the second prompt and the auxiliary information. [Supplementary Item 8] The information processing device of Supplementary Item 7, wherein the processor is mounted on a camera, the image acquisition conditions include shooting conditions of the camera, and the auxiliary information includes at least one of information on the image, the shooting environment, the shooting time, and audio around the camera.[Supplementary Item 9] The information processing device according to any one of Supplementary Items 3, 4, 6, and 8, wherein, when relearning the first language model, the processor uses, in addition to a user instruction, at least one of information related to the image recognition result of the image that is the target of the instruction, information related to the shooting environment, and information related to the shooting time, for the relearning process. [Supplementary Item 10] The information processing device according to any one of Supplementary Items 3, 4, 6, 8, and 9, wherein the processor trains the first language model based on user answers inputted to prepared questions during initial setup of the camera. [Supplementary Item 11] The information processing device according to any one of Supplementary Items 3, 4, 6, and 8 to 10, wherein at least one of the first language model and the second language model is stored in a portable storage device that is detachable from the camera, or in an external storage device that is separate from the camera and can communicate with the processor. [Supplementary Item 12] The information processing device according to any one of Supplementary Items 3, 4, 6, and 8 to 11, wherein, when the first language model is stored in a portable storage device detachable from the camera or an internal storage device built into the camera, the processor has a function of transferring the first language model to another camera. [Supplementary Item 13] The information processing device according to any one of Supplementary Items 3, 4, 6, and 8 to 12, wherein, when the first language model is stored in a portable storage device detachable from the camera or an internal storage device built into the camera, the first language model can be excluded from initialization targets when initializing internal information of the camera. [Supplementary Item 14] The information processing device according to Supplementary Item 1, wherein the image acquisition conditions are image acquisition conditions in a virtual space. [Supplementary Item 15] The information processing device according to Supplementary Item 14, wherein the image acquisition conditions in the virtual space are shooting conditions of a virtual camera in the virtual space, and are conditions that mimic at least one of the F-number, shutter speed, ISO sensitivity, and color temperature included in the shooting conditions of a real camera.[Supplementary Item 16] The information processing device according to Supplementary Item 1, wherein the image acquisition conditions are image generation conditions in an image generation model configured with a machine learning model. [Supplementary Item 17] A method for operating an information processing device having a processor, wherein the processor converts a first prompt, which is a user request for image acquisition, into a second prompt different from the first prompt, using a first language model, deriving image acquisition conditions corresponding to the second prompt, using the second language model, and re-learning the first language model based on a user instruction regarding the image acquisition conditions. [Supplementary Item 18] An operating program for an information processing device having a processor, wherein the operating program causes the processor to execute processes including: converting a first prompt, which is a user request for image acquisition, into a second prompt different from the first prompt, using the first language model, deriving image acquisition conditions corresponding to the second prompt, using the second language model, and re-learning the first language model based on a user instruction regarding the image acquisition conditions.
[0103] In the above-described embodiments, various processes in the information processing device are executed by an arbitrary computer. Furthermore, the arbitrary computer may execute these processes by a processor as hardware, a program as software, or a combination thereof. In such a case, the processor is configured to execute the various processes in the present embodiment in cooperation with the program, and may function as each unit or each means in the present embodiment. Furthermore, the order in which the processes are executed by the processor is not limited to the order described above and may be changed as appropriate.
[0104] The given computer may be a general-purpose computer, a computer for specific applications, a workstation, or any other system capable of executing each process. The processor may be configured with one or more pieces of hardware, and the type of hardware is not limited. For example, the processor may be configured with hardware such as a central processing unit (CPU), a micro processing unit (MPU), a programmable logic device such as a field programmable gate array (FPGA), a dedicated circuit for executing specific processes such as an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or a neural processing unit (NPU). The type of hardware may also be a combination of different types of hardware. When multiple pieces of hardware are configured to execute one or more processes of a certain processor, the multiple pieces of hardware may be located in devices physically separated from each other, or may be located in the same device. Furthermore, in any embodiment, the order of the processes performed by the processor is not limited to the order described above and may be changed as appropriate. The hardware is configured by an electric circuit (circuitry) or the like that combines circuit elements such as semiconductor elements.
[0105] Furthermore, the program may be software, such as firmware or microcode. The program may also be, for example, a group of program modules, each function of which may be implemented by a processor configured to perform the respective function. The program may be program code and / or multiple code segments stored in one or more non-transitory computer-readable media (e.g., storage media and / or other storages). The program may be stored across multiple non-transitory computer-readable media that reside in physically separate devices. The program code or code segment may represent a procedure, a function, a subprogram, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. The program code or code segment may be connected to another code segment or a hardware circuit by sending or receiving information, data, arguments, parameters, or memory contents.
[0106] The technology of the present disclosure can also be appropriately combined with the various embodiments and / or various modified examples described above. Furthermore, the technology is not limited to the above embodiments, and various configurations can be adopted without departing from the spirit of the present disclosure. Furthermore, the technology of the present disclosure also covers, in addition to programs, storage media that non-temporarily store programs. The storage medium is, for example, a computer-readable non-temporary storage medium such as a USB (Universal Serial Bus) memory, a flexible disk, or a CD-ROM (Compact Disc Read Only Memory). The program may also be provided online via a network such as the Internet. The technology of the present disclosure also covers, in addition to programs, program products. A program product includes any type of product for providing a program. Like a program, a program product may be provided stored on a computer-readable non-temporary storage medium or provided online.
[0107] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[0108] In this specification, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."
[0109] The disclosure of Japanese Patent Application No. 2024-122677, filed on July 29, 2024, is incorporated herein by reference in its entirety. In addition, all documents, patent applications, and technical standards described herein are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard was specifically and individually indicated to be incorporated by reference.
Claims
1. An information processing device comprising a processor, which converts a first prompt, which is a user request regarding image acquisition, into a second prompt different from the first prompt using a first language model, derives image acquisition conditions corresponding to the second prompt using the second language model, and re-trains the first language model based on instructions from the user regarding the image acquisition conditions.
2. The information processing device according to claim 1, wherein the processor presents the image acquired based on the image acquisition conditions to the user and accepts the instruction from the user.
3. The information processing device according to claim 1, wherein the processor is mounted on a camera, and the image acquisition conditions include shooting conditions of the camera.
4. The information processing device according to claim 3, wherein the photographing conditions include at least one of an F-number, a shutter speed, an ISO sensitivity, and a white balance setting.
5. An information processing device as described in claim 1, wherein when converting from the first prompt to the second prompt, the processor is capable of accepting auxiliary information in addition to the first prompt, and outputs the second prompt based on the first prompt and the auxiliary information.
6. The information processing device according to claim 5, wherein the processor is mounted on a camera, the image acquisition conditions include the shooting conditions of the camera, and the auxiliary information includes at least one of information regarding the image, the shooting environment, the shooting time, and the sound around the camera.
7. An information processing device as described in claim 1, wherein when deriving the image acquisition conditions in response to the second prompt, the processor is capable of accepting auxiliary information in addition to the second prompt, and derives the image acquisition conditions based on the second prompt and the auxiliary information.
8. An information processing device as described in claim 7, wherein the processor is mounted on a camera, the image acquisition conditions include the shooting conditions of the camera, and the auxiliary information includes at least one of information regarding the image, the shooting environment, the shooting time, and the sound around the camera.
9. An information processing device as described in claim 3, wherein when the first language model is retrained, the processor uses at least one of the following information for the retraining: the image recognition result of the image that is the subject of the instruction, the shooting environment, and the shooting time, in addition to the instruction from the user.
10. An information processing device according to claim 3, wherein the processor trains the first language model based on user responses to questions prepared in advance during initial setup of the camera.
11. An information processing device as described in claim 3, wherein at least one of the first language model and the second language model is stored in a portable storage device that can be attached and detached to the camera, or in an external storage device separate from the camera and capable of communicating with the processor.
12. An information processing device as described in claim 3, wherein when the first language model is stored in a portable storage device that can be attached and detached to the camera or an internal storage device built into the camera, the processor has the function of transferring the first language model to another camera.
13. An information processing device as described in claim 3, wherein when the first language model is stored in a portable storage device that can be attached and detached to the camera or an internal storage device built into the camera, the first language model can be excluded from initialization when initializing internal information of the camera.
14. The information processing device according to claim 1, wherein the image acquisition conditions are image acquisition conditions in a virtual space.
15. An information processing device as described in claim 14, wherein the image acquisition conditions in the virtual space are the shooting conditions of a virtual camera in the virtual space, and are conditions that mimic at least one of the F-number, shutter speed, ISO sensitivity, and color temperature included in the shooting conditions of a real camera.
16. The information processing device according to claim 1, wherein the image acquisition conditions are image generation conditions in an image generation model configured using a machine learning model.
17. A method for operating an information processing device having a processor, wherein the processor converts a first prompt, which is a user request regarding image acquisition, into a second prompt different from the first prompt using a first language model, derives image acquisition conditions corresponding to the second prompt using a second language model, and re-learns the first language model based on instructions from the user regarding the image acquisition conditions.
18. An operating program for an information processing device having a processor, the operating program causing the processor to execute processes including: converting a first prompt, which is a user request regarding image acquisition, into a second prompt different from the first prompt using a first language model; deriving image acquisition conditions corresponding to the second prompt using a second language model; and re-learning the first language model based on instructions from the user regarding the image acquisition conditions.
Citation Information
Patent Citations
Method of displaying a photographing mode by using lens characteristics, computer-readable storage medium of recording the method and an electronic apparatus
US20150189167A1
Server device and photographing device
WO2014103441A1
Learning system and data collection device
WO2021200503A1