A shooting guidance method, an electronic device, and a computer-readable storage medium
By outputting guiding information in the camera application interface of electronic devices to instruct users on adjusting shooting methods, the problem of poor photo quality caused by insufficient user skills and inspiration is solved, thus improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2024-12-31
- Publication Date
- 2026-06-30
AI Technical Summary
Electronic devices often produce poor-quality photos due to users' lack of shooting skills and inspiration, resulting in a poor user experience.
A shooting guidance method is provided, which displays a preview image in the camera application interface and outputs guidance information to instruct users to adjust shooting elements, composition, lighting, color matching, filter style or camera mode, etc., to improve photo quality.
With guidance from the information provided, users can adjust their shooting methods, improve photo quality, and enhance the user experience.
Smart Images

Figure CN122317399A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more particularly to a shooting guidance method, an electronic device, and a computer-readable storage medium. Background Technology
[0002] In the field of artificial intelligence, an intelligent agent (or agent) is an entity capable of perceiving its environment and reacting based on the information it perceives. For example, a voice assistant on an electronic device (such as a mobile phone) is an intelligent agent software that can interact with the user through natural language. Specifically, the user can issue voice commands to the voice assistant, which can then respond to those commands and, combined with information detected within the electronic device, automatically perform a series of operations to complete the voice command.
[0003] In daily life, users can use voice assistants to help them take photos quickly and conveniently. Specifically, Figure 1 This diagram illustrates the interface changes when taking a photo using a voice assistant. Figure 1 As shown in Figure (a), the user points their phone 100 at the scene they want to photograph based on their shooting inspiration and skills, and then issues a voice command to the voice assistant, "Please take a picture for me." The voice assistant responds to this voice command; if it detects that the camera application interface is not currently displayed, it will open the camera application and... Figure 1 As shown in Figure (b), the camera application interface 101 is displayed, and then the camera application's photo-taking function is invoked to take a picture of the scene that the phone 100 can currently capture. In addition, users can also point the phone 100 at the scene they want to photograph based on their own shooting inspiration and skills, and click as shown... Figure 1 The shutter control 1011 shown in Figure (b) allows the mobile phone 100 to take a picture of the scene that the mobile phone 100 can capture in response to the user's operation.
[0004] However, regardless of the method used to take photos, electronic devices may produce photos of poor quality or that do not meet the user's expectations due to the user's lack of shooting skills (such as shooting angle) and inspiration, resulting in a poor user experience. Summary of the Invention
[0005] This application provides a shooting guidance method, an electronic device, and a computer-readable storage medium, which provide shooting guidance to users, thereby improving shooting quality and enhancing user experience.
[0006] To achieve the above objectives, this application adopts the following technical solution:
[0007] In a first aspect, this application provides a shooting guidance method applied to an electronic device. The method includes: displaying the interface of a camera application, the interface of which includes a preview image; and, when a shooting guidance mode is enabled, outputting guidance information for the preview image, wherein the guidance information includes suggestion information to guide the user to adjust the shooting method, the shooting method including one or more of the following: selection of shooting elements, composition, lighting, color matching, filter style, camera mode, or parameter settings corresponding to the camera mode.
[0008] The camera app's interface is a preview interface, used to display preview images.
[0009] In this application, the guidance information displayed on the electronic device can provide shooting guidance to the user. After seeing the guidance information, the user can adjust the shooting methods such as the selection of shooting elements, composition, lighting, color matching, filter style, camera mode or the parameter settings corresponding to the camera mode, so as to take high-quality photos or videos. High-quality photos or videos are generally photos or images that the user is more satisfied with, thereby improving the user's shooting experience.
[0010] In one possible implementation of the first aspect, the guidance information further includes: evaluation information and / or reason information; wherein the evaluation information is an evaluation of the current shooting method corresponding to the preview image, and the reason information is used to explain the suggestion information.
[0011] The guidance information includes not only suggestions to guide users in adjusting shooting methods, but also explanations of the reasons behind those suggestions. This provides users with a more comprehensive understanding of the rationale behind these adjustments. In this way, users not only follow the suggestions but also understand the reasons behind them, resulting in a superior user experience.
[0012] The guidance information includes suggestions to help users adjust their shooting methods, as well as evaluation information on the current shooting method corresponding to the preview image. This way, users can not only adjust their shooting methods according to the suggestions, but also identify shortcomings in their shooting process, resulting in a higher user experience.
[0013] In one possible implementation of the first aspect, the guidance information is displayed as text or played as audio. The output methods for the suggestion, evaluation, and reason information in the guidance information can be the same, all played as audio, or all displayed as text. The output methods for the suggestion, evaluation, and reason information in the guidance information can also be different, such as playing at least one of the suggestion, evaluation, and reason information as audio, and displaying the other information as text.
[0014] Displaying the guidance information as text allows users to see it intuitively, resulting in a better user experience.
[0015] In one possible implementation of the first aspect, the guidance information is obtained based on a first preset model; the first preset model has the ability to suggest, evaluate, and provide reasons for the shooting method of the preview image.
[0016] The first presupposed model is the aesthetics expert model mentioned below.
[0017] In one possible implementation of the first aspect, the first preset model is set in an electronic device; the method further includes: inputting a preview image into the first preset model; determining a question to be asked, wherein the question to be asked includes one or more of the following: a question for evaluating the preview image, a question for suggesting a shooting method for the preview image, a question for providing reasons for suggesting a shooting method for the preview image, and a question related to guiding the user to adjust the selection of shooting elements, adjust composition, adjust lighting, adjust color matching, adjust filter style, adjust camera mode, or adjust the parameter settings corresponding to the camera mode; inputting the question to be asked into the first preset model, and obtaining a response to the question to be asked from the first preset model, wherein the response to the question to be asked is the response of the first preset model to the question to be asked based on the preview image; and determining guidance information based on the response to the question to be asked.
[0018] This means the electronic device can communicate with a first preset model to obtain guidance information. The number of questions to be asked can be one or more. The more questions and the more content they contain, the richer the guidance information, the more detailed the shooting guidance provided to the user, and the higher the user experience.
[0019] In one possible implementation of the first aspect, the first preset model is set in a server; the method further includes: sending a preview image to the server; determining a question to be asked, wherein the question to be asked includes a question for evaluating the preview image, a question for suggesting a shooting method for the preview image, a question for providing reasons for suggesting a shooting method for the preview image, and one or more questions related to guiding the user to adjust the selection of shooting elements, adjust composition, adjust lighting, adjust color matching, adjust filter style, adjust camera mode, or adjust the parameter settings corresponding to the camera mode; sending the question to be asked to the server; receiving a response to the question to be asked from the server, wherein the response to the question to be asked is a response from the server to the question to be asked based on the preview image using the first preset model; and determining guidance information based on the response to the question to be asked.
[0020] The first preset model has a large amount of data, which can be placed on a server to reduce the load on electronic devices.
[0021] In one possible implementation of the first aspect, determining the question to be asked includes: determining the question to be asked based on one or more of the following: understanding the user's voice command, understanding the preview image, setting information of the parameters of the system tools applied by the camera when capturing the preview image, and the answer to the question to be asked by the first preset model.
[0022] In user interaction with the camera application of an electronic device, the user issues voice commands to the corresponding application, and the electronic device acquires these voice commands. There may be one or more questions to be asked. The understanding of the preview image can be based on a first preset model.
[0023] The electronic device inputs one or more of the following into a second preset model: the user's voice command understanding, the understanding of the preview image, the parameter settings of the system tools used by the camera when capturing the preview image, and the answer to the question from the first preset model, to determine the question to be asked. The second preset model is the photography expert agent mentioned below.
[0024] If the user does not issue a voice command to the corresponding application on the electronic device, the electronic device determines the question to be asked based on one or more of the following: its understanding of the preview image and the parameter settings of the system tools corresponding to the camera application when the preview image was captured.
[0025] If the user does not issue a voice command to the corresponding application on the electronic device, the electronic device determines the question to be asked based on one or more of the following: understanding the user's voice command, understanding the preview image, and the parameter settings of the system tools corresponding to the camera application when the preview image was captured.
[0026] In one possible implementation of the first aspect, determining the question to be asked includes: inputting one or more of the following into a second preset model: understanding the user's voice command, understanding the preview image, setting information of the parameters corresponding to the system tools of the camera application when capturing the preview image, and the answer to the question to be asked from a first preset model, and determining the question to be asked based on the second preset model.
[0027] The second preset model can be the photography expert agent mentioned below.
[0028] In one possible implementation of the first aspect, the question to be asked is a pre-defined question.
[0029] In one possible implementation of the first aspect, the method further includes: auxiliary information for outputting the suggestion information; the auxiliary information is used to assist the user in understanding the suggestion information; wherein the auxiliary information includes any one or more of the following: a reference image, a level, reference lines, and bounding box information of the target object in the preview image, used to guide the user to re-capture the preview image.
[0030] In one possible implementation of the first aspect, before outputting auxiliary information for the suggestion information, the method further includes: if preset information is included in the question to be asked and / or the answer to the question to be asked, then a preset tool corresponding to the preset information is invoked to generate auxiliary information.
[0031] Secondly, this application provides a model training method applied to a second electronic device, the method comprising:
[0032] Obtain first training data, which includes photographic photos, questions posed to the photographic photos, and corresponding response data; use the photographic photos and questions posed to the photographic photos as input samples of the first preset model, and use the corresponding response data as output samples of the first preset model to train the first preset model, thereby obtaining the trained first preset model.
[0033] Since the first training data includes photographs, questions posed to the photographs, and corresponding responses, the trained first pre-set model is capable of responding to the questions posed to the photographs.
[0034] In one possible implementation of the second aspect, the method further includes: acquiring first training data, wherein the first training data includes photographic photos, data describing the content in the photographic photos, data of questions raised about the photographic photos, and response data corresponding to the question data; using the photographic photos, data describing the content in the photographic photos, and data of questions raised about the photographic photos as input samples of a first preset model, using the response data corresponding to the question data as output samples of the first preset model, and retraining the first preset model to obtain a trained first preset model.
[0035] Thirdly, this application provides an electronic device that includes at least a memory and one or more processors. The memory stores computer instructions, which, when executed by the one or more processors, cause the electronic device to perform any of the methods described in the first aspect above.
[0036] Fourthly, this application provides a computer storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform any of the methods described in the first aspect above.
[0037] Fifthly, this application provides a computer program product that, when run on an electronic device, causes the electronic device to perform any of the methods described in the first aspect above. Attached Figure Description
[0038] Figure 1 This diagram illustrates the interface changes when taking photos using a voice assistant.
[0039] Figure 2 A schematic diagram illustrating the interface changes during the activation process of a shooting guidance mode is shown.
[0040] Figure 3 This diagram illustrates the interface changes of a mobile phone 100 in providing guidance and assistance information to the user.
[0041] Figure 4 This diagram illustrates the interface changes of a mobile phone 100 in providing guidance information to a user.
[0042] Figure 5 This illustration shows a scenario where a user takes a photo.
[0043] Figure 6 This diagram illustrates another variation in the interface used by a mobile phone 100 to provide guidance and assistance information to the user.
[0044] Figure 7 This diagram illustrates another variation in the interface used by a mobile phone 100 to provide guidance and assistance information to the user.
[0045] Figure 8 A schematic diagram of the structure of the mobile phone 100 provided in an embodiment of this application is shown;
[0046] Figure 9 A flowchart illustrating a shooting guidance method is shown;
[0047] Figure 10 This diagram illustrates another process for deriving guiding and auxiliary information.
[0048] Figure 11 This illustration shows a scenario where a photography expert agent communicates with an aesthetics expert model.
[0049] Figure 12 This illustration shows another scenario where a photography expert agent communicates with an aesthetics expert model.
[0050] Figure 13 This illustration shows another scenario where a photography expert agent communicates with an aesthetics expert model.
[0051] Figure 14 This illustration shows another scenario where a photography expert agent communicates with an aesthetics expert model.
[0052] Figure 15 A schematic diagram of images 151 to 154 in Table 1 is shown;
[0053] Figure 16 This diagram illustrates the difference in the ability of a general multimodal large model and an aesthetic expert model to evaluate photographic images.
[0054] Figure 17 This diagram illustrates a workflow for a photography expert agent to raise and resolve questions.
[0055] Figure 18 This diagram illustrates the process of an aesthetics expert model handling questions posed by a photography expert agent.
[0056] Figure 19 This diagram illustrates a process for obtaining guidance information through interaction between a mobile phone and a server. Detailed Implementation
[0057] The technical solutions of the embodiments of this application are described below with reference to the accompanying drawings. In the description of the embodiments of this application, the terminology used in the following embodiments is for the purpose of describing specific embodiments only and is not intended to limit the application. As used in the specification and appended claims of this application, the singular expressions "a," "the," "the," "the," and "this" are intended to also include expressions such as "one or more," unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, "at least one" and "one or more" refer to one or more (including two). The term "and / or" is used to describe the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0058] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized. The term "connection" includes direct connections and indirect connections, unless otherwise stated. "First" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated.
[0059] In the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0060] As mentioned in the background above, electronic devices may produce photos of poor quality or that do not meet the user's expectations due to the user's lack of shooting skills (such as shooting angle) and shooting inspiration, resulting in a poor user experience.
[0061] To address this technical problem, this application provides a shooting guidance method, comprising: an electronic device displaying a camera application interface, the camera application interface including a preview window for displaying a preview image captured by the electronic device. When the shooting guidance mode is enabled, guidance information is output to the preview image to instruct the user on adjusting the shooting method. Instructing the user on adjusting the shooting method may include guiding the user to adjust one or more of the following: selecting shooting elements, adjusting composition, etc. That is, the electronic device can provide suggestions to the user on adjusting the shooting method based on the preview image. In this way, the user can readjust the shooting method based on the suggestions and take a photo using the electronic device. This increases the probability that the electronic device will take a high-quality photo or a photo that meets the user's expectations, thus improving the user experience.
[0062] It is understood that the shooting guidance method of this application embodiment is adapted to scenarios of taking photos, and can also be adapted to scenarios of shooting videos. In the scenario of shooting videos, the electronic device displays the interface of a camera application. In recording mode, the interface of the camera application includes a preview window, which is used to display a preview image captured by the electronic device. When the shooting guidance mode is enabled, guidance information is output to the preview image to guide the user in adjusting the shooting method. Guiding the user in adjusting the shooting method may include guiding the user to adjust one or more of the following: selecting shooting elements, adjusting composition, etc.
[0063] The following uses a scene of taking photos as an example to specifically introduce the shooting guidance method provided in the embodiments of this application.
[0064] The following is an example of how to enable a shooting guidance mode.
[0065] Figure 2 This diagram illustrates the interface changes during the activation of a shooting guidance mode. Figure 2 As shown in Figure (a), the camera application interface 101 of the mobile phone 100 generally includes four areas, which are, from bottom to top of the mobile phone 100, the shutter area, the bottom menu area, the framing area (or preview window), and the top menu area.
[0066] The controls located in the shutter area, from left to right, include: a shortcut control for the Gallery (or Album) application, a shutter control, and a camera switching control. The bottom menu area provides multiple camera mode controls, as well as parameter settings controls for the current camera mode. The camera mode controls, from left to right, include: large aperture mode, night mode, portrait mode, still mode, video mode, movie mode, professional mode, and more controls. The parameter settings controls for the current camera mode are displayed above these camera mode controls. Figure 2 As shown in (a), the current camera mode is the photo mode, and the zoom function 107 is the parameter setting control for the photo mode, displayed above the aforementioned camera mode control. The interactive interface displayed in the top menu area includes, from left to right, controls such as Smart Object Recognition 102, Artificial Intelligence (Ai) Photography 103, Flash 104, Filters 105, and Settings 106. It should be noted that the controls and control layout included in the camera application interface 101 described above are only an example. The controls and control layout included in the camera application interface may differ for different products, and this application embodiment does not impose specific limitations on this.
[0067] Continue reading Figure 2In Figure (a), the mobile phone 100 can activate the camera application's shooting guidance mode in response to the user's activation operation of the Ai camera 103. In some embodiments, the user's activation operation of the Ai camera 103 can be a tap operation, a long press operation, or other touch operation. Alternatively, the user's activation operation of the Ai camera 103 can be a voice wake-up operation. For example, if the user says to the voice assistant, "Help me activate the shooting guidance mode in the camera application," the mobile phone 100 will respond to this operation and activate the shooting guidance mode. However, it is not limited to these methods.
[0068] With the shooting guidance mode enabled, the phone can display something like this: Figure 2 The camera application interface is shown in Figure (b). Figure 2 The camera application interface shown in Figure (b) is similar to... Figure 2 One of the differences in the camera app interface shown in Figure (a) is that the icon for AiPhotography103 is different. When the shooting guidance mode is not enabled, as shown... Figure 2 As shown in Figure (a), the AiPhotography103 icon has a diagonal line. When the shooting guidance mode is enabled, as shown... Figure 2 As shown in Figure (b), the diagonal line on the Ai Photography 103 icon has disappeared. That is, the presence of a diagonal line on the Ai Photography 103 icon indicates that the shooting guidance mode is not enabled, while the absence of a diagonal line on the Ai Photography 103 icon indicates that the shooting guidance mode is enabled.
[0069] In some embodiments, in response to the user's activation of the Ai photography 103, the mobile phone 100 may also display guidance information to the user, reminding the user that the shooting guidance mode has been activated. For example, after the mobile phone 100 activates the shooting guidance mode in response to the user's activation operation, the mobile phone 100 displays the following: Figure 2 The camera application interface shown in Figure (b) displays the guidance message 108: "Shooting guidance mode is enabled" within a preset time after the shooting guidance mode is activated. The preset time can be 30 milliseconds, etc., but is not limited to this. The content of the guidance message reminding the user that the shooting guidance mode is enabled is not limited to this; it can also include phrases such as "Photography Master is enabled" or "Photography Expert is enabled."
[0070] In some embodiments, after the mobile phone 100 responds to the user's operation to start the Ai photography 103 and activates the shooting guidance mode of the camera application, a shooting guidance mode persistent indicator may also be displayed. For example, as Figure 2The persistent shooting guidance mode indicator 109 shown in Figure (b) indicates that the shooting guidance mode is enabled. It can be understood that the mobile phone 100 can display guidance information within the preset display area where the persistent shooting guidance mode indicator 109 is located; that is, the mobile phone 100 displays guidance information in the display area near the persistent shooting guidance mode indicator. Therefore, the persistent shooting guidance mode indicator can prompt the user to view the guidance information in the display area near the persistent shooting guidance mode indicator.
[0071] The location of the "Permanent Shooting Guide Mode" icon 109 in the camera application interface is set according to the actual situation. For example, as... Figure 2 As shown in Figure (b), the camera application interface 101 displays a shooting guide mode persistent indicator 109, which may be displayed in the bottom menu area, but is not limited to this.
[0072] The shape of the permanent shooting guide mode icon 109 is set according to the actual situation. For example, as... Figure 2 As shown in Figure (b), the permanent marker 109 in the shooting guidance mode can be a pentagram or other shapes, but is not limited to these.
[0073] In some other embodiments, after the mobile phone 100 responds to the user's operation to enable Ai Photography 103 and activates the camera application's shooting guidance mode, the persistent shooting guidance mode indicator may not be displayed. Instead, only the Ai Photography 103 icon without diagonal lines may be displayed, meaning that the absence of diagonal lines on the Ai Photography 103 icon is used to indicate that the shooting guidance mode has been enabled. Ai Photography 103 is merely one exemplary entry control for enabling the shooting guidance mode. The entry point for enabling the shooting guidance mode can also be other controls, and different products may have different entry points for enabling the shooting guidance mode. This application embodiment does not impose specific limitations on this.
[0074] It is understood that the guidance information is used to instruct users to adjust their shooting methods. The guidance information may include suggestions to guide users to recapture the preview image. These suggestions may be displayed as text on the camera application's interface. That is, the guidance information is output in text format containing these suggestions. This embodiment illustrates the example of a mobile phone 100 displaying guidance information for the preview image as text on the camera application's interface. It is understood that in other embodiments, the guidance information may be output in other ways besides this method, such as through voice or animation.
[0075] For example, User A is playing on the beach with their family in the evening. User A doesn't know what other scenes they can capture besides the sea and their children playing. So, after enabling the shooting guidance mode in the camera app of Phone 100, User A can give the camera app a voice command saying "Please give me some shooting inspiration" to get suggestions for re-capturing preview images.
[0076] Specifically, Figure 3 This diagram illustrates the interface changes of a mobile phone 100 when providing guidance and assistance information to the user. (For example...) Figure 3 As shown in Figure (a), mobile phone 100 displays preview image 1 in the preview window of the camera application interface 101. Preview image 1 includes a beach, seaside, and sky. Mobile phone 100 responds to the voice command "Please give me some photo inspiration" issued by user A, such as... Figure 3 As shown in Figure (b), the camera application interface 101 displays suggested information 110 to guide the user to recapture the preview image: the current time is suitable for shooting silhouette style, the beach scene allows for various poses, and the composition can be referenced in Figures A and B. Figures A and B each include one pose. In this way, the user can adjust the shooting method based on this suggested information, instruct the camera application to take a picture again, and obtain a photo that meets the user's expectations.
[0077] In some embodiments, to enable users to intuitively see and understand the guidance information, the electronic device may also output auxiliary information. The auxiliary information is used to assist the user in understanding the suggestions in the guidance information. The auxiliary information includes a reference image used to guide the user to re-capture the preview image. For example, please continue reading... Figure 3 ,like Figure 3 As shown in Figure (b), the preview window of the camera application interface 101 on the mobile phone 100 also displays reference images, Figure A and Figure B, to guide the user to recapture the preview image. Figure A shows a standing posture on a rocky shore, while Figure B shows a pose with arms outstretched facing the sea. In this way, user A can take photos of similar high quality or that meet user A's requirements based on the compositional inspiration from Figures A and B, thus improving user A's photography experience.
[0078] It can be understood that the guidance information includes not only suggestions to guide users to re-capture the preview image, but also explanations of the suggestions and evaluations of the preview image. In other words, the guidance information is output as a text structure combining suggestions, explanations, and evaluations. The evaluations are assessments of the method used to capture the preview image. The explanations clarify the suggestions.
[0079] For example, User B stands in front of a large snow sculpture, wanting to capture the moment and hoping to take a high-quality photo of themselves with the snow sculpture. However, User B feels their photography skills are lacking. Therefore, after enabling the shooting guidance mode in the camera app of Phone 100, they can issue a voice command to the camera app, "I want to take a photo with the snow sculpture, please provide me with shooting guidance," to receive evaluation information on the preview image, the reasons for the evaluation, and suggestions to guide the user to retake the preview image.
[0080] Specifically, Figure 4 This diagram illustrates a change in the interface of a mobile phone 100 when providing guidance information to the user. For example... Figure 4 As shown in Figure (a), mobile phone 100 displays preview image 2 in the preview window of the camera application interface 101. Preview image 2 includes a photo of the snow sculpture and user B. Responding to user B's voice command, "I want to take a photo with the snow sculpture, please provide me with photo-taking guidance," mobile phone 100 generates an evaluation message for preview image 2: the human figure in the preview image is too small. The camera application of mobile phone 100 also generates suggested information to guide the user to re-capture the preview image, along with corresponding reasoning: it is suggested to hold the phone back, move the phone up and down first, then left and right, to adjust the image, use 2x telephoto, and highlight the main human figure. Here, "suggest holding the phone back, moving the phone up and down first, then left and right" and "2x telephoto" are suggested information, while "to adjust the image" and "highlight the main human figure" are reasoning information.
[0081] like Figure 4 As shown in Figure (b), the preview window of the camera application interface 101 on the mobile phone 100 displays guidance information 112, including the aforementioned evaluation information, reason information, and suggestion information: "The current image of the person is relatively small. It is recommended to hold the phone back and move it up and down first, then left and right, to adjust the image. Use 2x telephoto to highlight the main subject." In this way, the user can readjust the shooting method according to this guidance information, instruct the camera application to take a picture again, and obtain an image that meets the user's expectations, thus improving the user experience.
[0082] For example, Figure 5 The illustration depicts a scenario where a user takes a photo. User B holds the phone, moves back 100 degrees, adjusts the frame, and then... Figure 5 As shown, the parameters in zoom function 107 are set to 2x telephoto, resulting in preview image 3 in the preview window. One difference between preview image 3 and preview image 2 is that the proportion of the human figure in preview image 3 is greater than that in preview image 2. Then, the mobile phone 100 responds to user B's click of the shutter control to take a picture, resulting in a higher quality photo and improving user B's photography experience.
[0083] It is understood that, in addition to the text structure mentioned above, the guidance information may be a combination of any one or more of the following: suggested information to guide the user to re-encode the preview image, reason information corresponding to the suggested information, and evaluation information of the preview image.
[0084] In the two examples above, the illustration focuses on how a single interaction between the phone 100 and the user can yield a photo that meets the user's expectations. It's understandable that the phone 100 can also interact with the user multiple times to obtain a photo that meets their expectations. A specific example could be the user interacting with the camera app on the phone 100 to obtain a photo that meets their expectations. Examples will be provided below.
[0085] For example, Figure 6 and Figure 7 This diagram illustrates the interface changes of a mobile phone 100 when providing guidance and assistance information to the user. (For example...) Figure 6 As shown in Figure (a), user C uses mobile phone 100 to take a picture of the current scenery. The preview window in the camera application interface 101 of mobile phone 100 displays preview image 4. The content of preview image 4 includes: the lake surface, the arched bridge on the lake surface and the reflection of the arched bridge in the water, the trees around the arched bridge and the reflection of the trees around the arched bridge in the water. At this time, user C issues a voice command to the camera application of mobile phone 100, "Guide me to take a picture." Mobile phone 100 responds to this voice command, as follows: Figure 6 As shown in Figure (b), the preview window of the camera application interface 101 displays guidance information 113: The current image is cluttered; it is recommended to remove the large area of leaf reflections and use a telephoto lens to highlight the bridge arch. To allow users to intuitively see and understand the guidance information, the phone 100 can also display auxiliary information, such as selection information for the target object. Specifically, in this scenario, the phone 100 will also display the suggested highlighting bridge arch and the cluttered leaves to be removed in a selection manner in the preview window of the camera application interface 101, such as... Figure 6 As shown in Figure (b), the mobile phone 100 displays boxes 114 and 115 of the bridge hole and box 116 of the leaves in the preview window.
[0086] After seeing the guidance information 113, the user can select telephoto mode to capture the current scene on the camera application interface 101, and the phone 100 will respond to this operation, such as... Figure 6 As shown in Figure (c), preview image 5 is displayed in the preview window. Preview image 5 includes: the lake surface, a partial arched bridge on the lake surface, the reflection of the partial arched bridge on the lake surface, trees around the arched bridge, and the reflections of the trees around the arched bridge in the water. The partial arched bridge includes a bridge opening.
[0087] like Figure 6 As shown in Figure (c), for the preview image 5, the mobile phone 100 displays guidance information 117 in the preview window of the camera application interface 101: Adjust the lens to place the bridge arch in the center of the frame to create a sense of balance and minimize the presence of vehicles passing on the arch. Furthermore, to allow users to intuitively see and understand the guidance information, the mobile phone 100 can also display auxiliary information, such as reference lines and selection information of target objects in the preview image. Specifically, in this scenario, the mobile phone 100 displays reference lines and selects passing vehicles on the arch in the preview window of the camera application interface 101, such as... Figure 6 As shown in Figure (c), mobile phone 100 displays reference lines and boxes 118 of passing vehicles on the arch bridge in the preview window.
[0088] The user adjusts the camera lens according to this guidance, and the phone responds accordingly. Figure 7 As shown in Figure (d), the mobile phone 100 displays preview image 6 in the preview window of the camera application interface 101. The preview image 6 includes: the lake surface, a partial arched bridge on the lake surface, the reflection of the partial arched bridge on the lake surface, the trees around the arched bridge, and the reflection of the trees around the arched bridge in the water. The bridge arch is located in the center of the entire image, and there are no passing vehicles on the arched bridge.
[0089] If user C is satisfied with the preview image 6, they can click the shutter control in the camera application interface 101 to have the phone 100 capture the image they are satisfied with. If user C wants to further optimize the preview image 6, such as changing the filter style, etc. Figure 7 As shown in Figure (d), you can send a message to the camera app on your phone: "Do you have any suitable shooting styles to recommend? Please help me convert and preview the effect."
[0090] like Figure 7 As shown in Figure (d), the preview window of the camera application interface 101 on the mobile phone 100 displays guidance information 119: Recommended film style. The camera application of the mobile phone 100 converts the preview image 6 into a film style and displays the converted image: preview image 7. The preview image 7 has the same image content as the preview image 6, the difference being that the preview image 7 is in film style, while the preview image 6 does not have the film style added.
[0091] like Figure 7 As shown in Figure (e), for the preview image 7, the mobile phone 100 displays guidance information 120 in the preview window of the camera application interface 101: "The composition is full; you can wait for the boat to pass through the bridge arch to take the picture, enhancing the dynamism and atmosphere of the image." Upon seeing this guidance information, if a boat passes through the bridge arch, the user can then... Figure 7As shown in Figure (f), the user can click the shutter control in the camera application interface 101 to allow the phone 100 to take a satisfactory image. This demonstrates that the camera application of the phone 100 can interact with the user multiple times until a satisfactory image is captured, resulting in a high level of human-computer interaction experience.
[0092] The above scenario illustrates the case where auxiliary information may include the bounding box information of the target object in the preview image, or the auxiliary information may include the bounding box information of the target object in the preview image and reference lines. It is understood that auxiliary information may also include a level. Auxiliary information may include any one or a combination of multiple of the following: the bounding box information of the target object in the preview image, reference lines, and a level.
[0093] The suggestion, evaluation, and reason information can be output in the same way. In some embodiments, the mobile phone 100 can output the suggestion, evaluation, and reason information by voice. In some embodiments, the mobile phone 100 can display the suggestion, evaluation, and reason information as text.
[0094] The output methods for suggestion information, evaluation information, and reason information can also be different. In some embodiments, the mobile phone 100 can output at least one of the suggestion information, evaluation information, and reason information in a voice manner, and display the suggestion information, evaluation information, and reason information that is not output in a voice manner in a text manner.
[0095] The following description, in conjunction with the accompanying drawings, details the photo prompting method provided in the embodiments of this application.
[0096] In this embodiment, the aforementioned electronic device is an electronic device that supports photography. Specifically, the electronic device can be a portable electronic device or other suitable electronic device. For example, the electronic device can be a mobile phone, tablet personal computer, laptop computer, personal digital assistant (PDA), camera, personal computer, laptop computer, wearable device, augmented reality (AR) glasses, AR headset, virtual reality (VR) glasses, or VR headset, etc. The following description uses a mobile phone as an example to illustrate this embodiment.
[0097] Taking mobile phones as an example, Figure 8 A schematic diagram of the structure of the mobile phone 100 provided in an embodiment of this application is shown.
[0098] like Figure 8As shown, the mobile phone 100 may include a processor 810, an external memory interface 820, an internal memory 821, a universal serial bus (USB) interface 830, a charging management module 840, a power management module 841, a battery 842, antenna 1, antenna 2, a mobile communication module 850, a wireless communication module 860, an audio module 870, a speaker 870A, a receiver 870B, a microphone 870C, a headphone jack 870D, a sensor module 880, buttons 890, a motor 891, an indicator 892, a camera 893, a display screen 894, and a subscriber identification module (SIM) card interface 895, etc. The sensor module 880 may include a touch sensor 880K, etc.
[0099] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the mobile phone 100. In other embodiments of this application, the mobile phone 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0100] The processor 810 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. The controller may serve as the central nervous system and command center of the mobile phone 100. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.
[0101] The charging management module 840 receives charging input from the charger. The power management module 841 connects to the battery 842, the charging management module 840, and the processor 810. The power management module 841 receives input from the battery 842 and / or the charging management module 840 to power the processor 810, internal memory 821, external memory, display 894, camera 893, and wireless communication module 860, etc.
[0102] The wireless communication function of mobile phone 100 can be realized through antenna 1, antenna 2, mobile communication module 850, wireless communication module 860, modem processor and baseband processor.
[0103] Antennas 1 and 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 850 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use on mobile phone 100. Wireless communication module 860 can provide wireless communication solutions, including wireless local area networks (WLAN) (such as Wi-Fi), Bluetooth, Global Navigation Satellite System (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR), for use on mobile phone 100. In some embodiments, antenna 1 of mobile phone 100 is coupled to mobile communication module 850, and antenna 2 is coupled to wireless communication module 860, enabling mobile phone 100 to communicate with networks and other devices via wireless communication technology.
[0104] The buttons 890 include the power button, volume buttons, etc.
[0105] The mobile phone 100 can achieve shooting functions through the camera 893, ISP, digital signal processor, video codec, etc.
[0106] Camera 893 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, mobile phone 100 may include one or N cameras 893, where N is a positive integer greater than 1. In this embodiment, when the camera application is enabled, camera 893 can capture images or videos corresponding to the current scene in real time.
[0107] The ISP (Image Signal Processor) is used to process data fed back from the camera 893. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 893.
[0108] A digital signal processor (DSP) is used to process digital signals. Besides digital image signals, it can also process other digital signals. For example, when a mobile phone 100 is selecting a frequency, the DSP performs Fourier transforms on the frequency energy.
[0109] Video codecs are used to compress or decompress digital video. Mobile phone 100 can support one or more video codecs. Thus, mobile phone 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0110] The mobile phone 100 implements its display function through a GPU and a display screen 894. The GPU is a microprocessor for image processing, connected to the display screen 894. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 810 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0111] The display screen 894 is used to display images, videos, etc. The display screen 894 includes a display panel. In this embodiment, the display screen 894 is used to display the interface of a camera application. The camera application interface includes a preview image, which can be obtained by the camera 893 capturing an image corresponding to the current scene when the camera application is enabled. The display screen 894 is also used to display guidance information and auxiliary information.
[0112] The mobile phone 100 can achieve audio functions such as music playback and recording through the audio module 870, speaker 870A, receiver 870B, microphone 870C, headphone jack 870D, and application processor.
[0113] The audio module 870 is used to convert digital audio information into analog audio signal output, and also to convert analog audio input into digital audio signal. The audio module 870 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 870 may be located in the processor 810, or some functional modules of the audio module 870 may be located in the processor 810.
[0114] Microphone 870C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 870C, inputting the sound signal into microphone 870C. Mobile phone 100 can be equipped with at least one microphone 870C. In some embodiments, mobile phone 100 can be equipped with two microphones 870C, which, in addition to collecting sound signals, can also achieve noise reduction. In other embodiments, mobile phone 100 can also be equipped with three, four, or more microphones 870C, which can collect sound signals, reduce noise, identify the sound source, and achieve directional recording functions, etc. In the embodiments of this application, the user... Figure 2 As shown in Figure (a), the activation operation of Ai Photography 103 can be a user's voice wake-up operation for Ai Photography 103. For example, the user says to the voice assistant, "Help me activate the shooting guidance mode in the camera application," and the mobile phone 100 responds to this operation and activates the shooting guidance mode. Specifically, when the user says to the mobile phone 100, "Help me activate the shooting guidance mode in the camera application," the microphone 870C of the mobile phone 100 acquires the voice signal and sends it to the voice assistant of the mobile phone 100. The voice assistant of the mobile phone 100 recognizes the semantics of the voice signal and activates the shooting guidance mode.
[0115] Touch sensor 880K, also known as a "touch panel," can be located on display screen 894. The touch sensor 880K and display screen 894 together form a touchscreen, also known as a "touch screen." Touch sensor 880K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 894. In other embodiments, touch sensor 880K may also be located on the surface of mobile phone 100, in a different position than display screen 894. In this embodiment, the user... Figure 2 As shown in Figure (a), the Ai Camera 103 can be activated by the user through a tap or long press, or other touch operations. Specifically, the user taps... Figure 2 As shown in Figure (a), the Ai camera 103's touch sensor 880K detected the event and activated the shooting guidance mode.
[0116] Figure 9 A flowchart illustrating a shooting guidance method is shown, such as... Figure 9 As shown, the process includes the following steps:
[0117] 901. Display the camera application interface, which includes a preview window that displays a preview image.
[0118] It's understandable that the preview image could be a real-time view captured by the phone.
[0119] 902. In response to the user's operation of enabling the shooting guidance mode, display the interface of the corresponding camera application that has enabled the shooting guidance mode.
[0120] After the mobile phone 100 activates the shooting guidance mode, it can execute the shooting guidance method provided in this embodiment. The way the mobile phone 100 responds to the user's activation of the shooting guidance mode has been described above and will not be repeated here. For example, the interface of the camera application corresponding to the activation of the shooting guidance mode can be as follows: Figure 2 The corresponding application interface is shown in Figure (b).
[0121] The difference between the interface of the camera application with the shooting guidance mode enabled and the interface of the camera application without the shooting guidance mode enabled has been described above and will not be repeated here. In this embodiment, both the interface of the camera application with the shooting guidance mode enabled and the interface of the camera application in step 901 include a preview image. The preview images in the two camera application interfaces may be the same or different. The difference between the interface of the camera application with the shooting guidance mode enabled and the interface of the camera application in step 901 is that the icon for Artificial Intelligence (Ai) photography 103 is different, and the interface of the camera application with the shooting guidance mode enabled has an additional shooting guidance mode persistent flag 109 compared to the interface of the camera application in step 901.
[0122] In some embodiments, the mobile phone 100 may enable the shooting guidance mode before step 901. For example, if the mobile phone 100 opens the camera application before step 901, it enables the shooting guidance mode in response to the user's operation. In this way, when the user opens the camera application again, the shooting guidance mode is already enabled, and the mobile phone 100 can directly execute step 903 after executing step 901.
[0123] This embodiment describes an example where the shooting guidance mode is off by default. Therefore, the shooting guidance mode can be enabled based on the user's operation before or after step 901. In other embodiments, the shooting guidance mode can be enabled by default, thus eliminating the need for the phone 100 to enable the shooting guidance mode before step 901 or to enable it in step 902.
[0124] 903. Output guidance and auxiliary information for the preview image.
[0125] The guidance information is used to instruct users to adjust their shooting methods. Instructing users to adjust their shooting methods can refer to one or more of the following: adjusting the selection of shooting elements, adjusting composition, adjusting lighting and shadow, adjusting color matching, adjusting filter style, adjusting camera mode, or adjusting the parameter settings corresponding to the camera mode.
[0126] The selection of shooting elements refers to the selection of various components such as people, objects, and scenery that appear in the scene to be photographed during the photography process.
[0127] Adjusting composition refers to modifying the position, proportions, and angles of various elements in a scene during photography. Element position refers to moving the main subject or important elements in the frame, such as placing the subject at the golden ratio point or following the rule of thirds. Proportion refers to adjusting the size ratios of various elements in the image to make them more harmonious and highlight the main subject or important elements. Angle changes refer to changing the shooting angle, such as high angle, low angle, eye-level, or bird's-eye view, to showcase the subject from different perspectives.
[0128] Lighting and shadow effects refer to enhancing the sense of depth and three-dimensionality in an image by adjusting light and shadow. Adjusting color matching refers to adjusting the color relationships within an image, such as the use of contrasting colors, complementary colors, and analogous colors.
[0129] Adjusting filter styles refers to the effects applied to photos or videos after shooting.
[0130] The more comprehensive the guidance information, the more precise the adjustments users can make based on that information, and the better the images will meet their expectations.
[0131] This application embodiment uses the example of displaying guidance information for preview images on the camera application interface of mobile phone 100. It is understood that in some other embodiments, the guidance information may be output in the form of voice in addition to this method.
[0132] Adjusting the camera mode or its corresponding parameter settings has been described above. The following provides an example of adjusting the parameter settings for the camera mode. For instance, ... Figure 2 As shown in (a), the current camera mode is the photo mode, and the zoom function 107 is a parameter setting control for the photo mode. Users can use the zoom function 107 to adjust the zoom ratio, such as 1x, 2x, 5x, etc. This embodiment mainly uses the photo mode as an example, but it is not limited to this. The shooting guidance method provided in this embodiment is also applicable to other camera modes. When the camera mode is another mode (such as large aperture, night scene, portrait, video, movie, professional, more, etc.), the guidance information also includes suggested parameter settings corresponding to that camera mode, and users can also adjust the parameter settings corresponding to that camera mode.
[0133] This application embodiment is illustrated by taking the example of a mobile phone 100 outputting guidance information and auxiliary information for a preview image. It is understood that in some other embodiments, the mobile phone 100 may output guidance information for a preview image to guide the user to take a picture.
[0134] In some embodiments, even without the user issuing a voice command to interact with the camera application on the mobile phone 100, the mobile phone 100 can obtain and display guidance information based on the preview image upon detecting the preview image. In other embodiments, when the user issues a voice command to interact with the camera application on the mobile phone 100, the mobile phone 100 displays guidance information for the preview image on the camera application's interface.
[0135] Guiding and auxiliary information can be determined in various ways. The following section introduces a specific method for determining guiding and auxiliary information. Figure 10 A schematic diagram illustrating the process of deriving guiding and auxiliary information is shown. The determination of guiding and auxiliary information can be achieved through multiple functional modules. For example... Figure 10 As shown, the guidance and auxiliary information are defined by multiple functional modules, including a photography expert agent, a toolset module, and an aesthetic expert model. The photography expert agent has the ability to understand user voice commands, ask questions to the aesthetic expert, summarize the answers to questions, and invoke the toolset module. The aesthetic expert model has the ability to answer questions posed by the photography expert agent. The toolset module includes system tools for the camera application, a web application programming interface (API), and an expert model library, etc. Figure 2 As shown in Figure (a), the various controls included in the camera application interface 101 described above can be referred to as the system tools of the camera application. The API and expert model library will be described in detail below.
[0136] The following is an example of a specific technical solution for determining boot information based on a boot information determination module. For example... Figure 10 As shown, the process includes the following steps:
[0137] Step 1: The photography expert agent acquires multimodal information related to the preview image.
[0138] In computer science and human-computer interaction, multimodality generally refers to the use of multiple different input and output methods to achieve human-computer interaction. For example, a multimodal application can use multiple input methods such as voice and gestures, and multiple output methods such as images and text to interact with the user simultaneously.
[0139] In some embodiments, the multimodal information associated with the preview image may include the preview image itself. When the user issues a voice command to interact with the camera application in the mobile phone 100, the multimodal information may also include the user's voice command.
[0140] In some embodiments, the aforementioned multimodal information related to the preview image may include the preview image itself and setting information of parameters corresponding to the system tools of the camera application when capturing the preview image. When the user issues a voice command to interact with the camera application on the mobile phone 100, the multimodal information further includes the user's voice command. The system tools of the camera application include, for example... Figure 2 The various functional controls and their corresponding functional tools are shown in Figure (a).
[0141] In some embodiments, multimodal information related to the preview image is stored in the memory of the mobile phone 100, and the photography expert agent can retrieve the multimodal information related to the preview image from the memory. In other embodiments, the photography expert agent can obtain the parameter settings information corresponding to the system tools of the camera application based on its understanding of the camera application's interface.
[0142] Step 2: The photography expert agent communicates with the aesthetics expert model to determine the guiding information.
[0143] The guidance information includes any one or more of the following: evaluation information on the shooting method of the preview image, suggestions to guide users to adjust the shooting method, and reasons for the suggestions.
[0144] The photography expert agent and the aesthetics expert model can engage in one or more rounds of dialogue to obtain more accurate guidance information. Specifically, the photography expert agent asks the aesthetics expert model at least one question, the aesthetics expert model provides the photography expert agent with the answers to the questions posed by the photography expert agent, and the photography expert agent summarizes the answers to obtain more accurate guidance information.
[0145] The following examples illustrate two methods by which a photography expert agent obtains guidance information.
[0146] One approach utilizes prompt technology to facilitate multi-turn dialogues between a photography expert agent and an aesthetics expert model, thereby obtaining guidance information. Structured prompts refer to questions arranged in a specific order, guiding the aesthetics expert model to think along pre-defined lines and arrive at conclusions. The photography expert agent can then ask questions to the aesthetics expert model in the pre-defined order. Correspondingly, the aesthetics expert model provides responses in the same order and outputs multiple responses back to the photography expert agent. The photography expert agent then summarizes these responses to obtain the guidance information.
[0147] In some embodiments, the structured prompts may be pre-stored in the mobile phone 100. After the photography expert agent inputs a preview image into the aesthetics expert model, the photography expert agent can ask questions to the aesthetics expert model according to the pre-stored structured prompts.
[0148] The questions arranged in a specific order can include: questions asking for feedback on the preview image; questions offering suggestions on how to shoot the preview image; questions providing the reasons behind those suggestions; and questions related to one or more of the following: adjusting the selection of shooting elements, adjusting composition, adjusting lighting, adjusting color matching, adjusting filter styles, adjusting camera modes, or adjusting the corresponding parameter settings for camera modes. The more detailed the questions, the richer the guidance information provided, and the more likely users are to capture images that better meet their expectations.
[0149] For example, the structured prompts could be: Question 1: Evaluate the photo; Question 2: Which elements should be emphasized in the composition; and Question 3: How should light and shadow be adjusted in the composition? For example, after a user issues the voice command "Guide me to take a photo," the photography expert agent sends the preview image and the structured prompts sequentially to the aesthetic expert model. The aesthetic expert model, utilizing its ability to understand image content, its photographic knowledge, and its ability to aesthetically evaluate images, obtains aesthetic evaluation results, the elements to be emphasized in the composition, and the results of adjusting light and shadow, and sends these results back to the photography expert agent. The photography expert agent summarizes the aesthetic evaluation results, the elements to be emphasized in the composition, and the results of adjusting light and shadow to derive the guidance information.
[0150] Another approach is to use Large Language Model (LLM) technology to enable multi-turn dialogues between a photography expert agent and an aesthetics expert model to obtain guidance information. A Large Language Model (LLM) is a deep learning model trained on a large amount of text data. LLM can handle various natural language tasks, such as question answering and dialogue. This type of model supports diverse and generalized natural language processing, meaning it can flexibly respond to various user commands.
[0151] Specifically, in some embodiments, the photography expert agent determines the question to be asked and then sends the question to the aesthetics expert model. The aesthetics expert model outputs the answer to the received question back to the photography expert agent. Afterward, based on its understanding of the received answer, the photography expert agent generates another question and continues to ask it to the aesthetics expert model. The aesthetics expert model outputs the corresponding answer to the received question back to the photography expert agent, and this cycle continues until the photography expert agent stops asking questions to the aesthetics expert model. In this process, there is a one-to-one correspondence between questions and answers. The photography expert agent summarizes the answers and derives guidance information.
[0152] The questions to be asked can include: questions to evaluate the preview image; questions to suggest shooting methods for the preview image; questions explaining the reasons for suggesting shooting methods for the preview image; and questions related to guiding users to adjust one or more of the following: selection of shooting elements, composition, lighting, color matching, filter style, camera mode, or corresponding parameter settings. The more detailed the questions, the richer the guidance information provided, and the more likely users are to capture images that better meet their expectations.
[0153] For example, Figure 11 This illustration depicts a scenario where a photography expert agent communicates with an aesthetics expert model. Figure 11 As shown, for Image 1, after the user sends a message to the camera app to "guide me to take a picture", the photography expert agent and the aesthetics expert model have the following three rounds of communication.
[0154] The photography expert Agent asks the photography expert Agent question 1: If you were a photographer and wanted to take a picture in this place, what elements would you remove and highlight in the picture?
[0155] An aesthetics expert model answered question 1: If I were a photographer, I would remove the white boat in the foreground and highlight the cruise ship and blue sky in the picture.
[0156] The photography expert Agent asks the photography expert Agent 2: Please frame the position of the white boat in the foreground of the picture.
[0157] The aesthetics expert model answers question 2 as follows: <ref> White boat< / ref> <box> (267,833),(418,999)< / box> <box> (820,833),(997,988)< / box> .
[0158] Photographer Agent asks Photographer Agent Question 3: I want to highlight the cruise ship and blue sky in the picture. Should I use a telephoto lens or a wide-angle lens?
[0159] The aesthetics expert model answers question 3: If you want to highlight the cruise ship and the blue sky, a telephoto lens is more suitable. A telephoto lens can bring distant objects closer and blur background elements, making the cruise ship and the blue sky stand out more. A wide-angle lens is more suitable for shooting landscapes, as it can include more scenery in the frame, but it may make background elements appear more cluttered.
[0160] <ref> White boat< / ref> <box> (267,833),(418,999)< / box> <box> (820,833),(997,988)< / box> The four coordinates of the selected white boat are (267,833), (418,999), (820,833), and (997,988).
[0161] During the three rounds of communication between the photography expert agent and the aesthetics expert model, the photography expert agent, based on its LLM (Limited Learning Model) understanding of the user's voice commands, arrives at question 1. Then, based on the aesthetics expert model's response to question 1, it raises question 2, and based on the aesthetics expert model's response to question 2, it raises question 3. Finally, the three responses from the three rounds of communication are summarized to obtain guiding information. For example, the guiding information might be: "It is recommended to remove the white boat in the foreground of the image to highlight the cruise ship and the blue sky. If you want to emphasize the cruise ship and the blue sky, a telephoto lens would be more suitable, as it can bring distant objects closer and blur background elements, making the cruise ship and the blue sky stand out more. A wide-angle lens is more suitable for shooting landscapes, as it can capture more scenery in the frame, but it may make background elements appear more cluttered."
[0162] Question 1 relates to guiding users on adjusting the selection of shooting elements, and Question 3 relates to adjusting the parameter settings corresponding to the camera mode. The guidance information obtained from the answers to Questions 1 and 3 includes suggestions on adjusting the selection of shooting elements and the parameter settings corresponding to the camera mode. Thus, by seeing this guidance information, users can adjust the selection of shooting elements and the parameter settings corresponding to the camera mode to obtain images that better meet their expectations.
[0163] Figure 12 This illustration depicts a scenario where a photography expert agent communicates with an aesthetics expert model. Figure 12 As shown, for image 2, after the user sends a message to the camera app to "guide me to take a picture", the photography expert agent and the aesthetics expert model have the following two rounds of communication.
[0164] Photography expert Agent poses question 1 to aesthetics expert Model: If you were a photographer and wanted to take a picture in this place, what would you use as the main subject in your composition?
[0165] An aesthetics expert model answered question 1: If I were a photographer, I would use the blue and white cruise ship as the main subject in my composition.
[0166] Photography expert Agent poses question 2 to aesthetics expert Model: If you were a photographer and wanted to use a cruise ship as the main subject in your composition, where would you place the cruise ship in the frame?
[0167] The aesthetics expert model answers question 2: If a photographer wants to use the cruise ship as the main subject in the composition, he can place the cruise ship in the center of the picture or place it at the golden ratio point of the picture to create a sense of balance. In addition, the photographer can also use the outline of the cruise ship to guide the viewer's eye and enhance the sense of depth in the picture.
[0168] During the two rounds of communication between the photography expert agent and the aesthetics expert model, the photography expert agent, based on its LLM's understanding of the user's voice commands, arrives at question 1. Then, based on the aesthetics expert model's response to question 1, it poses question 2. Finally, by summarizing the two responses from the two rounds of communication, guiding information is obtained. For example, the guiding information might be: It is suggested that the blue and white cruise ship be used as the main subject in the composition. The cruise ship can be placed in the center of the frame, or at the golden ratio point, to create a sense of balance. Furthermore, the photographer can use the cruise ship's outline to guide the viewer's eye and enhance the image's depth.
[0169] Questions 1 and 2 above relate to adjusting composition. The guidance information derived from the answers to questions 1 and 2 includes suggestions for adjusting composition. By seeing this guidance, users can adjust the composition to obtain an image that better meets their expectations.
[0170] Figure 13 This illustration depicts a scenario where a photography expert agent communicates with an aesthetics expert model. Figure 13 As shown in Figure 3, after the user sends a message to the camera app to "guide me to take a picture," the photography expert agent and the aesthetics expert model engage in the following two rounds of communication.
[0171] Photography expert Agent asks aesthetics expert Model: Please evaluate the image and tell me how to improve the shooting effect?
[0172] The aesthetics expert model responded to question 1: The composition of the image is relatively full, but the colors are a bit monotonous. It is recommended to try different color combinations when shooting to increase the sense of depth in the image.
[0173] The photography expert Agent asks the aesthetics expert Model question 2: Please evaluate the image and then tell me how to adjust the brightness and saturation?
[0174] The aesthetics expert model responded to question 2: The image has rich colors and strong contrast between light and dark, but the overall image is too dark. It is recommended to appropriately increase the brightness and decrease the saturation to make the colors more vivid.
[0175] During the two rounds of communication between the photography expert agent and the aesthetics expert model, the photography expert agent, based on its LLM (Limited Learning Model) understanding of the user's voice commands, arrives at Question 1. Then, based on the aesthetics expert model's response to Question 1, it poses Question 2. Finally, by summarizing the two responses from the two rounds of communication, guiding information is obtained. For example, the guiding information might be: "The composition of the image is relatively full, but the colors are somewhat monotonous. It is recommended to try different color combinations when shooting to increase the depth of the image." or "The image has rich colors and strong contrast, but the overall image is too dark. It is recommended to appropriately increase the brightness while decreasing the saturation to make the colors more vivid."
[0176] Questions 1 and 2 above relate to adjusting color schemes and lighting. The guidance information derived from the answers to questions 1 and 2 includes suggestions for adjusting color schemes and lighting. By seeing this guidance, users can adjust color schemes and lighting to obtain images that better meet their expectations.
[0177] Figure 14 This illustration depicts a scenario where a photography expert agent communicates with an aesthetics expert model. Figure 14 As shown, for image 4, after the user asked the camera app, "Do you have any suitable personalized filter styles to recommend? Please convert them for me and then give me your feedback on the final image," the photography expert Agent and the aesthetics expert model engaged in the following round of communication.
[0178] Photography expert Agent poses question 1 to aesthetics expert Model: Are there any suitable personalized filter styles to recommend? And how would you evaluate images that have undergone filter style transformation?
[0179] The aesthetics expert model answers question 1: "Fresh" style is recommended, as the composition is full, the colors are bright, the image is vivid, and the theme is prominent.
[0180] During the aforementioned exchange between the photography expert agent and the aesthetics expert model, the photography expert agent, based on its LLM understanding of the user's voice commands, arrives at question 1. Then, based on the aesthetics expert model's response to question 1, it obtains guiding information. For example, the guiding information might be: "Recommend a 'fresh' style, with full composition, vibrant colors, lively imagery, and a prominent theme."
[0181] Question 1 above relates to adjusting filter styles. The guidance information obtained from Question 1 includes suggestions for adjusting filter styles. In this way, the phone can adjust the filter style to obtain an image that better meets the user's expectations.
[0182] In some embodiments, the aesthetics expert model can also answer questions posed by the photography expert agent based on the user's historical personality preferences. For example, in the scenario described above, the photography expert agent can recommend a "fresh" style to the aesthetics expert model based on the user's previous history of using the "fresh" style.
[0183] In some embodiments, the photography expert agent can determine the question to be asked based on one or more of the following: understanding the user's voice commands, understanding the preview image, parameter settings of the system tools used by the camera application when capturing the preview image, and the aesthetic expert model's response to the question.
[0184] For example, when a user interacts with the camera application on the phone 100 by issuing voice commands, the photography expert agent can obtain the user's voice commands and, based on the photography expert agent's understanding of the user's voice commands, derive the question to be asked.
[0185] For example, even without the user issuing voice commands to interact with the camera app on phone 100, the photography expert agent obtains a preview image and sends it to the aesthetic expert model. The aesthetic expert model then provides feedback to the photography expert agent regarding its understanding of the preview image, and the photography expert agent, based on the aesthetic expert model's understanding, derives the questions to be asked.
[0186] For example, a photography expert agent can obtain the parameter settings information of the system tools used by the camera application when capturing and previewing images, and derive the questions to be asked based on this information. Below is an example of a photography expert agent posing one or both of the following related questions to an aesthetics expert model based on the system tool parameter settings: adjusting the camera mode or adjusting the corresponding parameter settings.
[0187] For example, in the scenario of photographing a snow sculpture with a person, if the photography expert agent obtains that the zoom function 107 of the camera application is set to 1X when capturing preview image 2, the photography expert agent can ask the aesthetic expert model: "The current focal length is 1X, should the focal length be adjusted?" The aesthetic expert model's response could be: "It is recommended to change 1X to 2X." The photography expert agent summarizes this response in the guidance information, displaying the suggested parameter setting in the guidance information, such as... Figure 4 As shown in Figure (b), the camera application interface displays guidance information 112: The current image of the person is relatively small. It is recommended to hold the phone and move it back, first up and down and then left and right to adjust the image. Use the telephoto lens at 2x to highlight the main subject.
[0188] For example, a photography expert agent can derive the question to be asked based on the answers to the question from an aesthetics expert model.
[0189] The following is an example of the training process for an aesthetics expert model. The training process of an aesthetics expert model includes two steps: one is the construction of training data, and the other is the training of the aesthetics expert model using the training data.
[0190] Training data includes photographs, questions posed to those photographs, and corresponding responses. The electronic device training the model uses the photographs and questions from the training data as input samples to the aesthetic expert model, and the responses from the training data as output samples. The device adjusts the model parameters based on the error between the actual output and the output samples until the error is within a preset range, resulting in a well-trained aesthetic expert model. Thus, when questions are posed to photographs, the trained model can answer them. Furthermore, the more training data available, the higher the accuracy of the aesthetic expert model's responses, leading to a better user experience.
[0191] In some embodiments, to enhance the aesthetic expert model's ability to understand the content of images, the training data includes not only photographs, questions posed to the photographs, and corresponding responses, but also data describing the content within the photographs. After training the aesthetic expert model using photographs, questions posed to the photographs, and corresponding responses, the model can be retrained using the data describing the content within the photographs, the photographs themselves, the questions posed to the photographs, and corresponding responses, fine-tuning the model parameters to obtain a well-trained aesthetic expert model. In this way, the aesthetic expert model possesses the ability to answer questions about photographs based on its understanding of them.
[0192] The question data includes one or more of the following: questions related to commenting on photographs, suggested questions related to re-capturing images, or questions providing reasons for suggested questions. For example, a question related to commenting on photographs could be: Please rate the photo. A suggested question related to re-capturing images could be: Suggestions for improving the photo. A question providing reasons for suggested questions could be: Why would you suggest such improvements to the photo?
[0193] Training data can be obtained from newspapers or public online platforms.
[0194] Correspondingly, the response data corresponding to the question data includes one or more of the following: comment information corresponding to the photograph, suggestion information corresponding to the photograph, and reason information corresponding to the suggestion information. For example, the comment information corresponding to the photograph could be: The human figure in the preview image is too small. The suggestion information corresponding to the photograph could be: It is recommended to hold the phone back and retake the shot with a 2x telephoto lens. The reason information corresponding to the suggestion could be: It is recommended to hold the phone back to adjust the image, and to retake the shot with a 2x telephoto lens to highlight the subject, the human figure.
[0195] Because the problem data includes questions related to photographic images, suggested questions related to re-capturing images, and questions providing reasons for those suggestions, a well-trained aesthetic expert model is able to comment on, suggest, and provide reasons for those suggestions to the input images.
[0196] In this embodiment, the suggested questions related to re-capturing images can involve more granular inquiries than simply offering suggestions for improvement. For example, suggested questions might be related to guiding the user to adjust one or more of the following: the selection of shooting elements, composition, lighting, color matching, filter style, camera mode, or the parameter settings corresponding to the camera mode. The more detailed the questions, the richer the guidance information provided, and the more likely the user is to capture images that better meet their expectations.
[0197] General-purpose multimodal large language models (MLLMs) and large vision models (LVMs) typically possess multimodal understanding and reasoning capabilities, semantic understanding, and so on. Aesthetic expert models can be trained based on MLLMs or LVMs. By constructing the aforementioned training samples to train MLLMs or LVMs, they can acquire the ability to understand and answer questions related to commenting on input images, making suggestions, and providing reasons for those suggestions.
[0198] The following is an example of a training sample data format, which is shown below:
[0199]
[0200]
[0201] As shown in the data format above, 2015-01-01(2), 2015-01-01(3), 2015-01-01(4), and 2015-01-01(5) represent the names of the images. The data format above includes the storage path corresponding to the images with the aforementioned names. The electronic device training the model can obtain the images named 2015-01-01(2), 2015-01-01(3), 2015-01-01(4), and 2015-01-01(5) from the corresponding paths.
[0202] The above data format includes the question corresponding to the images with the image names 2015-01-01(2), 2015-01-01(3), 2015-01-01(4) and 2015-01-01(5): Please evaluate the image.
[0203] In the above data format, the comments on the image titled 2015-01-01(2) are: The main figures in the work are prominent, and the compression of the telephoto lens is strong. The position of the trophy echoes the logo of the competition, which should be the author's decision after consideration. The comments on the image titled 2015-01-01(3) are: The children in the photo are like little warriors about to go into battle. They are still full of passion in the harsh environment, which makes people think deeply when they see them. The comments on the image titled 2015-01-01(4) are: The expressions of the characters talking are very good, real and natural, which makes the audience quickly think of the positive energy theme of "police and people in harmony to ensure safety". The shutter control is good and makes the fan dynamic. The background selection has a strong local flavor. The comments on the image titled 2015-01-01(5) are: The swan's movements are stretched and the water splashes, which makes the picture full of dynamism. The dark background and the use of light to emphasize the shape of the main subject show the author's ingenuity.
[0204] The electronic device training the model can obtain from the above data format: the questions and corresponding answers corresponding to the images named 2015-01-01(2), 2015-01-01(3), 2015-01-01(4) and 2015-01-01(5), respectively, and the answers are the above comment information.
[0205] The following example illustrates the effect of the trained aesthetic expert model obtained from the embodiments of this application in answering questions. As shown in Table 1 below, Table 1 shows a question posed to images 151 to 154 and the aesthetic expert model's response to the question. Figure 15 A schematic diagram of images 151 to 154 in Table 1 is shown. It can be understood that since images 151 to 154 in Table 1 respectively correspond to... Figure 15 Images 151 to 154 in the text. Combined Figure 15 As shown in Table 1 below, after training the aesthetics expert model using the training data, the aesthetics expert model is able to answer a series of questions related to shooting, such as evaluations and suggestions.
[0206] Table 1
[0207]
[0208]
[0209] The following section will input the same photo into the general multimodal model and the aesthetic expert model respectively, and use the difference in the output results of the general multimodal model and the aesthetic expert model to briefly illustrate the difference in their ability to evaluate photographs.
[0210] Figure 16 This diagram illustrates a comparison of the ability of a general multimodal large model and an aesthetic expert model to evaluate photographic images. Figure 16 As shown, because the general multimodal large model was not trained on the aforementioned training samples, its photo recognition ability, image content understanding ability, and ability to provide aesthetic evaluation and suggestions are weak. In contrast, the aesthetic expert model, trained on the aforementioned training samples, exhibits stronger photo recognition, image content understanding, and aesthetic evaluation and suggestions capabilities. The aesthetic expert model can recognize images 161 and 162 and evaluate their content, while the general multimodal large model cannot recognize or evaluate images 161 and 162. For image 163, the aesthetic expert model accurately identifies the clouds and determines they are redundant, suggesting appropriate cropping. The general multimodal large model, however, cannot identify the clouds in image 163 and only provides a general evaluation and suggestion. For image 164, the aesthetic expert model can provide an evaluation and suggestion, while the general multimodal large model can only evaluate image 164.
[0211] In summary, the aesthetic expert model in this application embodiment outperforms the general multimodal large model in terms of photo recognition capabilities and the ability to evaluate and make suggestions on photos.
[0212] Step 3: The photography expert agent determines auxiliary information.
[0213] The Photography Expert Agent can determine whether the questions and / or answers from the Aesthetics Expert involve preset information. Preset information includes information related to tools within the toolset module; for example, it could be fields related to those tools. If preset information is determined to be involved in the guidance information, it indicates the need to invoke tools from the toolset module to assist the user in adjusting their shooting method. In this case, the Photography Expert Agent can invoke the tools from the toolset module to display the auxiliary information.
[0214] For example, <grounding>For fields related to the selection box,<ref_lines> For fields related to reference lines. The photography expert agent determined that the aesthetics expert's response contained...<ground ing> At that time, the object detection capability in the toolset module is invoked to outline the corresponding object. For example, the photography expert agent determines that the aesthetics expert's response contains...<ref_lines> When working with fields, call the reference lines in the toolset module.
[0215] The toolset module includes system tools, network application programming interfaces (APIs), expert model libraries, and more.
[0216] The photography expert agent obtains auxiliary information through system tools, network APIs, and expert model libraries, and presents this information on the camera application interface. This allows users to understand the guidance information more intuitively, improving the user experience.
[0217] Among them, the system tools refer to the system tools of the camera application introduced above, which will not be repeated here.
[0218] For example, in the aforementioned scene filmed under the bridge, such as Figure 6 As shown in Figure (c), the mobile phone 100 displays auxiliary information—reference lines—in the preview window of the camera application interface 101. Specifically, in one embodiment, during multiple rounds of communication between the photography expert agent and the aesthetics expert model, the photography expert agent can determine whether the aesthetics expert's questions and / or answers involve…<ref_lines> In this field, the photography expert Agent can call the reference lines in the toolset module to instruct the phone 100 to display the reference lines in the preview window of the camera application interface 101.
[0219] Photography expert agents can use web APIs to communicate between camera applications and other applications, web browsers, and databases to obtain relevant data.
[0220] For example, in the aforementioned beach shooting scene, such as Figure 3 As shown in Figure (b), the mobile phone 100 displays auxiliary information (Figures A and B) in the preview window of the camera application interface 101. Specifically, in one embodiment, during multiple rounds of communication between the photography expert agent and the aesthetics expert model, the photography expert agent can determine that the aesthetics expert's questions and / or answers involve the field of photo pose search. At this time, the photography expert agent can call the network API in the toolset module to obtain Figures A and B, and instruct the mobile phone 100 to display Figures A and B in the preview window of the camera application interface 101.
[0221] An expert model library refers to a set of pre-trained deep learning models specifically designed for image processing and analysis. It includes image understanding models such as image recognition, object detection, and image segmentation. A photography expert agent can call upon this library to process preview images and obtain processing results. These results include object detection, image recognition, and image segmentation. Object detection, in particular, is the process of identifying and locating objects of interest in an image or video. It involves not only identifying the categories of objects present in the image but also determining their specific locations, typically represented by bounding boxes.
[0222] For example, in the aforementioned scene filmed under the bridge, such as Figure 6 As shown in Figure (b), the mobile phone 100 displays auxiliary information in the preview window of the camera application interface 101: frames 114 and 115 for the bridge arch and frame 116 for the leaves. Specifically, in one embodiment, during multiple rounds of communication between the photography expert agent and the aesthetics expert model, the photography expert agent can determine that the aesthetics expert's questions and / or answers involve selection fields. At this time, the photography expert agent can call the object detection capability in the toolset module to instruct the mobile phone 100 to display frames 114 and 115 for the bridge arch and frame 116 for the leaves in the preview window of the camera application interface 101.
[0223] It is understandable that in some embodiments, the image processing capabilities of the expert model library are also possessed by the aesthetic expert model. In the multiple rounds of communication between the photography expert agent and the aesthetic expert model, the photography expert agent determines that the questions from the aesthetic expert involve information related to image processing capabilities. At this time, the photography expert agent can directly ask questions to the aesthetic expert, utilize the aesthetic expert's image processing capabilities, and obtain the image processing results. After obtaining the image processing results, the photography expert agent instructs the mobile phone 100 to display the image processing results in the preview window of the camera application interface 101.
[0224] For example, in the aforementioned scene filmed under the bridge, such as Figure 6 As shown in Figure (b), the mobile phone 100 displays auxiliary information in the preview window of the camera application interface 101: frames 114 and 115 for the bridge arch and frame 116 for the leaves. In the multiple rounds of communication between the photography expert agent and the aesthetic expert model, the photography expert agent determines that the aesthetic expert's question involves the field of frame selection. At this time, the photography expert agent can directly ask the aesthetic expert a question and use the aesthetic expert's object detection capability to obtain the frame information of the target object. After obtaining the frame information of the target object, the photography expert agent instructs the mobile phone 100 to display frames 114 and 115 for the bridge arch and frame 116 for the leaves in the preview window of the camera application interface 101.
[0225] In the aforementioned Figure 10 The description indicates that the guidance information determination module may include a photography expert agent, a toolset module, and an aesthetics expert model. Some questions posed by the photography expert agent can be answered by the aesthetics expert model, while others require invoking the toolset module to perform the corresponding tasks. The following section details the logical relationship between the questions posed by the photography expert agent and the invocation of the toolset module and the aesthetics expert model.
[0226] Figure 17 This diagram illustrates a workflow for a photography expert agent to raise and process questions. The following section focuses on the logical relationship between the questions raised by the photography expert agent and the calls to the toolset module and the aesthetic expert model. The specific implementation schemes for each step have been described previously and will not be repeated here. The workflow includes the following steps:
[0227] 1701. The photography expert agent used LLM to identify the problem.
[0228] Photography expert Agent used LLM to identify problems with the preview images.
[0229] 1702. The photography expert agent determines whether to invoke the aesthetics expert model.
[0230] If the Photography Expert Agent determines that the Aesthetic Expert Model should be invoked, then the Aesthetic Expert Model should be invoked to answer the corresponding question, i.e., step 1704 should be executed. If the Photography Expert Agent determines that the Aesthetic Expert Model should not be invoked, then it should determine whether a toolset module call is involved, i.e., step 1703 should be executed.
[0231] 1703. The photography expert agent determines whether the use of toolset modules is involved.
[0232] If the Photography Expert Agent determines that a toolset module call is involved, it will call the toolset module, i.e., execute step 1706. If the Photography Expert Agent determines that a toolset module call is not involved, it will determine whether all questions should be executed, i.e., execute step 1705.
[0233] 1704. The photography expert Agent calls upon the aesthetics expert model to obtain the answer to the question.
[0234] 1705. The photography expert agent determines whether all issues have been addressed.
[0235] If the Photography Expert Agent determines that all questions have been answered, it summarizes the responses to all questions, generates guidance information, and executes step 1707. If the Photography Expert Agent determines that not all questions have been answered, it continues to use LLM to determine the next question and executes step 1702.
[0236] 1706. The photography expert agent calls the tools in the toolset module to perform the task corresponding to the problem.
[0237] 1707. The photography expert agent compiles the answers to all questions and provides guidance information.
[0238] In the aforementioned Figure 10 In the description, the guidance information determination module may include a photography expert agent, a toolset module, and an aesthetics expert model. Some of the questions raised by the photography expert agent can be answered by the aesthetics expert model. The following is an example of how the aesthetics expert model answers the questions raised by the photography expert agent.
[0239] Figure 18 This diagram illustrates a flowchart of how an aesthetics expert model processes questions posed by a photography expert agent. The following flowchart focuses on the process by which the aesthetics expert model handles these questions. The specific implementation schemes for each step have been described previously and will not be repeated here. The flowchart includes the following steps:
[0240] 1801. Aesthetics expert model obtains preview images and questions.
[0241] 1802. The aesthetics expert model determines whether the problem is a task related to the location of the target object being queried.
[0242] For example, in a scene depicting a cruise ship, the task of "please frame the location of the white boat in the foreground of the image" is related to querying the location of the target object.
[0243] If the aesthetics expert model determines that the problem is related to the location of the target object, then the image detection capability is invoked to determine the location of the target object, i.e., step 1803 is executed.
[0244] If the aesthetics expert model determines that the question is not a task related to the location of the target object, then answer the question, determine the answer to the question, and proceed to step 1804.
[0245] 1803. The aesthetics expert model uses image detection capabilities to determine the position of the target object.
[0246] The aesthetics expert model has image detection capabilities and can utilize these capabilities to determine the location of target objects.
[0247] 1804. The aesthetics expert model determines the answer to the problem.
[0248] It is understandable that, in some embodiments, since the cloud-side server contains a large material library, the network API in the toolset module can be set on the cloud-side server. This makes it convenient for the network API to directly call the corresponding materials from the cloud-side server and send them back to the mobile phone upon request.
[0249] The aesthetic expert model can be set on the mobile device or on a cloud server. The above description uses the example of setting the aesthetic expert model on the mobile device; the following description uses the example of setting the aesthetic expert model on a cloud server. For example, Figure 19 This diagram illustrates a process for obtaining guidance information through interaction between a mobile phone and a server. Figure 19 As shown in the diagram, the figure includes a mobile phone 100 and a cloud server 200. The aesthetic expert model is set on the cloud server side. The mobile phone 100 sends a question to the cloud server 200. When the cloud server 200 receives the question sent by the mobile phone 100, it calls the aesthetic expert model, makes a response, and then sends the response back to the mobile phone 100.
[0250] Specifically, the interaction steps between mobile phone 100 and cloud server 200 may include the following:
[0251] 1901. Mobile phone 100 sends a preview image to cloud server 200.
[0252] 1902. Mobile phone 100 calls the photography expert Agent to determine the question to be asked and sends the question to cloud server 200.
[0253] The questions to be asked include those for evaluating the preview image, those for suggesting shooting methods for the preview image, those for providing reasons for suggesting shooting methods for the preview image, and questions related to guiding users to adjust one or more of the following: selection of shooting elements, composition, lighting, color matching, filter style, camera mode, or corresponding parameter settings of the camera mode. The determination of the questions to be asked has already been introduced above and will not be repeated here.
[0254] 1903. The cloud server 200 receives the question to be asked, calls the aesthetics expert model, obtains the answer to the question to be asked, and sends the answer to the mobile phone 100.
[0255] Steps 1902 and 1903 can be performed multiple times, meaning that mobile phone 100 can ask multiple different questions to cloud server 200.
[0256] 1904. Mobile phone 100 receives the answer to the question from cloud server 200, and determines the guidance information based on the answer to the question.
[0257] The answers to the questions to be asked are provided by the cloud server 200 using an aesthetic expert model based on the preview images.
[0258] The aesthetics expert model has a large amount of data, which can be placed on a cloud server 200 to reduce the load on electronic devices.
[0259] In summary, the guidance information displayed on the phone can instruct users to adjust one or more of the following: the selection of shooting elements, composition, lighting, color matching, filter style, camera mode, or corresponding parameter settings. This, to some extent, increases the probability that the electronic device will take high-quality photos or photos that meet the user's expectations, thus improving the user experience.
[0260] This application also provides a computer storage medium that includes computer instructions. When the computer instructions are executed on the electronic device, the electronic device causes the electronic device to perform various functions or steps performed by the mobile phone in the above method embodiment.
[0261] This application also provides a computer program product that, when run on a computer, causes the computer to perform the various functions or steps performed by the mobile phone in the above method embodiments.
[0262] It is understood that the electronic device provided in this application embodiment includes hardware structures and / or software modules corresponding to perform each function in order to achieve the above-mentioned functions. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.
[0263] This application embodiment can divide the above-described electronic device into functional modules based on the method example described above. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0264] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0265] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0266] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0267] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0268] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0269] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.< / grounding>
Claims
1. A shooting guidance method, characterized in that, Applied to electronic devices, the method includes: The interface of the camera application is displayed, and the interface of the camera application includes a preview image; When the shooting guidance mode is enabled, guidance information is output for the preview image. The guidance information includes suggestion information, which is used to guide the user to adjust the shooting method. The shooting method includes one or more of the following: selection of shooting elements, composition, lighting, color matching, filter style, camera mode, or parameter settings corresponding to the camera mode.
2. The method according to claim 1, characterized in that, The guidance information also includes: evaluation information and / or reasoning information; The evaluation information refers to the evaluation information of the current shooting method corresponding to the preview image, and the reason information is used to explain the suggestion information.
3. The method according to claim 1 or 2, characterized in that, The guidance information is obtained based on a first preset model; the first preset model has the ability to suggest, evaluate, and provide reasons for the suggestions regarding the shooting method of the preview image.
4. The method according to claim 3, characterized in that, The first preset model is set in the electronic device; the method further includes: Input the preview image into the first preset model; The questions to be asked include questions for evaluating the preview image, questions for suggesting shooting methods for the preview image, questions for providing reasons for suggesting shooting methods for the preview image, and questions related to one or more of the following: guiding the user to adjust the selection of shooting elements, adjusting composition, adjusting lighting, adjusting color matching, adjusting filter style, adjusting camera mode, or adjusting the parameter settings corresponding to the camera mode. The question to be asked is input into the first preset model, and the answer to the question to be asked is obtained from the first preset model. The answer to the question to be asked is the answer given by the first preset model based on the preview image. The guidance information is determined based on the answers to the questions to be asked.
5. The method according to claim 3, characterized in that, The first preset model is set in the server; the method further includes: Send the preview image to the server; The questions to be asked include questions for evaluating the preview image, questions for suggesting shooting methods for the preview image, questions for providing reasons for suggesting shooting methods for the preview image, and questions related to one or more of the following: guiding the user to adjust the selection of shooting elements, adjusting composition, adjusting lighting, adjusting color matching, adjusting filter style, adjusting camera mode, or adjusting the parameter settings corresponding to the camera mode. Send the question to be asked to the server; Receive the response to the question from the server, wherein the response to the question is the server's response to the question based on the preview image using the first preset model; The guidance information is determined based on the answers to the questions to be asked.
6. The method according to claim 4 or 5, characterized in that, The determination of the question to be asked includes: The question to be asked is determined based on one or more of the following: understanding the user's voice command, understanding the preview image, setting information of the parameters of the system tools applied by the camera when capturing the preview image, and the first preset model's response to the question to be asked.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: The system outputs auxiliary information to assist the user in understanding the suggested information. The auxiliary information includes any one or more of the following: a reference image, a level, reference lines, and bounding box information of the target object in the preview image, used to guide the user to re-capture the preview image.
8. The method according to claim 7, characterized in that, Before outputting the auxiliary information of the suggested information, the method further includes: If the question to be asked and / or the answer to the question to be asked includes preset information, then the preset tool corresponding to the preset information is invoked to generate the auxiliary information.
9. An electronic device, characterized in that, The electronic device includes at least a memory and one or more processors; the memory is used to store computer instructions, which, when executed by the one or more processors, cause the electronic device to perform the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 8.
11. A computer program product, characterized in that, When the computer program product is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1 to 8.