Shooting method and device, electronic equipment and readable storage medium

By integrating AI models in electronic devices, using user input to obtain emotional information and interactive text information, and automatically performing shooting operations, the problem of cumbersome user operations is solved, and simple recording of image-related information and multimedia expression of emotional experience is realized.

CN120128792APending Publication Date: 2025-06-10NANJING VIVO SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510342984.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the process of recording related information about images, the user's operations are cumbersome and requires multiple operations to add notes to the image.

Method used

By integrating AI models in electronic devices, using user input to obtain emotional information and interactive text information, and automatically perform shooting operations based on this information, directly generating images or videos containing relevant information.

Benefits of technology

The operation process of users when recording image-related information is simplified, avoiding the tedious process of manually adding information, and providing users with images or videos related to emotional information and interactive text information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128792A_ABST
    Figure CN120128792A_ABST
Patent Text Reader

Abstract

The invention discloses a shooting method and device, electronic equipment and a readable storage medium, and belongs to the technical field of shooting. The shooting method comprises the steps that the electronic equipment can receive first input; in response to the first input, first information is obtained through an artificial intelligence AI model, and the first information comprises at least one of emotion information of the user and interactive text information of the user and the AI model; and executing shooting operation according to the first information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of shooting technology, and particularly relates to a shooting method, device, electronic device and readable storage medium. Background Art

[0002] Generally, after a user uses an electronic device to capture an image, if they want to record more relevant information about the image, they can first click on the thumbnail of the image in the album interface to make the electronic device display the image, and then perform multiple operations in the album interface to make the electronic device add a note including the relevant information to the image. Subsequently, when the user views the image, they can view the note simultaneously to recall the situation when the image was captured.

[0003] However, since the user needs to perform multiple operations after capturing the image to make the electronic device add a note including the relevant information to the image, the operation of the user is cumbersome during the process of recording the relevant information of the image. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a shooting method, device, electronic device and readable storage medium, which can solve the problem of cumbersome user operation during the process of recording the relevant information of the image.

[0005] In a first aspect, the embodiments of this application provide a shooting method, which includes: an electronic device can receive a first input; and in response to the first input, use an artificial intelligence AI model to obtain first information, where the first information includes at least one of the following: the user's emotional information, the interaction text information between the user and the AI model; and perform a shooting operation according to the first information.

[0006] In a second aspect, the embodiments of this application provide a shooting device, which includes: a receiving module, a processing module and an execution module. Among them, the receiving module is used to receive a first input. The processing module is used to, in response to the first input received by the receiving module, use the AI model to obtain first information, where the first information includes at least one of the following: the user's emotional information, the interaction text information between the user and the AI model; and perform a shooting operation according to the first information obtained by the processing module.

[0007] In a third aspect, the embodiments of this application provide an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, it implements the steps of the method described in the first aspect.

[0008] Fourthly, an embodiment of the present application provides a readable storage medium, on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps of the method described in the first aspect are implemented.

[0009] Fifthly, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instructions to implement the steps of the method described in the first aspect.

[0010] Sixthly, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the steps of the method described in the first aspect.

[0011] In an embodiment of the present application, an electronic device can obtain first information according to a first input of a user by using an AI model. The first information includes at least one of the user's emotion information and the interaction text information between the user and the AI model, and perform a shooting operation according to the first information. In this way, when the user inputs to the electronic device, an image or video related to the user's emotion information and / or the interaction text information between the user and the AI model can be captured, and when the user views the captured image, the user's emotion information and / or the interaction text information between the user and the AI model can also be obtained, without the user manually adding information to the image. Therefore, the cumbersome operation of the user in the process of recording the relevant information of the image is avoided. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is one of the flow diagrams of the shooting method provided by an embodiment of the present application;

[0013] Figure 2 is another flow diagram of the shooting method provided by an embodiment of the present application;

[0014] Figure 3 is yet another flow diagram of the shooting method provided by an embodiment of the present application;

[0015] Figure 4 is still another flow diagram of the shooting method provided by an embodiment of the present application;

[0016] Figure 5A is one of the interface diagrams of the mobile phone provided by an embodiment of the present application;

[0017] Figure 5B is another interface diagram of the mobile phone provided by an embodiment of the present application;

[0018] Figure 6 is yet another interface diagram of the mobile phone provided by an embodiment of the present application;

[0019] Figure 7 It is the fourth schematic diagram of the interface of the mobile phone provided by the embodiment of the present application;

[0020] Figure 8 It is the fifth schematic diagram of the interface of the mobile phone provided by the embodiment of the present application;

[0021] Figure 9 It is the sixth schematic diagram of the interface of the mobile phone provided by the embodiment of the present application;

[0022] Figure 10 It is the seventh schematic diagram of the interface of the mobile phone provided by the embodiment of the present application;

[0023] Figure 11 It is the fifth schematic diagram of the flowchart of the shooting method provided by the embodiment of the present application;

[0024] Figure 12 It is the eighth schematic diagram of the interface of the mobile phone provided by the embodiment of the present application;

[0025] Figure 13 It is the sixth schematic diagram of the flowchart of the shooting method provided by the embodiment of the present application;

[0026] Figure 14 It is the ninth schematic diagram of the interface of the mobile phone provided by the embodiment of the present application;

[0027] Figure 15 It is the schematic structural diagram of the shooting device provided by the embodiment of the present application;

[0028] Figure 16 It is one of the schematic diagrams of the hardware structure of the electronic device provided by the embodiment of the present application;

[0029] Figure 17 It is the second schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application fall within the protection scope of the present application.

[0031] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.

[0032] The following will, with reference to the accompanying drawings, through specific embodiments and their application scenarios, elaborate in detail on the shooting method, device, electronic device, and readable storage medium provided by the embodiments of this application.

[0033] Figure 1 The flowchart of the shooting method provided by the embodiments of this application is shown. As Figure 1 shown, the shooting method provided by the embodiments of this application may include the following steps 101 to 103.

[0034] Step 101: The electronic device receives a first input.

[0035] In some embodiments of this application, when the electronic device is displaying the settings application interface, it can enable the AI Island function according to the user's input to the AI function icon in the settings application interface. Thus, when the electronic device is in the powered-on state, the user can make a first input.

[0036] In some examples, the above-mentioned AI Island function can be understood as: a function of using an AI model to perform operations. Among them, the AI model can be deployed on the electronic device, or the server, or other electronic devices, and the operations can include but are not limited to file processing operations, image processing operations, user interaction operations, text recognition operations, music recognition operations, etc. Those skilled in the art can set them according to their needs, and the embodiments of this application do not limit them here.

[0037] In some examples, the above-mentioned powered-on state can be a state of being powered on and the screen being on, or a state of being powered on and the screen being off.

[0038] In some embodiments of this application, the above-mentioned first input can be the user's input to the electronic device, and this first input is used to trigger the electronic device to use the AI model to perform operations.

[0039] Exemplarily, the above first input includes but is not limited to: touch input of the electronic device by the user through a touch device such as a finger or a stylus (e.g., touch input on the display screen of the electronic device or the body of the electronic device), or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be specifically determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application can be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, a double click gesture, and a tap input; the click input in the embodiments of the present application can be a single click input, a double click input, or a click input of any number of times, etc., and can also be a long press input or a short press input.

[0040] For example, the above first input can be: the user performs a tapping input on the back of the electronic device in a preset tapping manner.

[0041] Step 102, in response to the first input, the electronic device uses the AI model to obtain first information, where the first information includes at least one of the following: the user's emotion information and the interaction text information between the user and the AI model.

[0042] In some embodiments of the present application, the above AI model can be the AI model associated with the above AI island function. The AI model can be a pre-trained model, so that the first information can be obtained using the AI model, and the number of the AI models can be at least one.

[0043] In some embodiments of the present application, the above emotion information can include at least one of the following: emotion type and emotion intensity information. Among them, the emotion intensity information can be used to indicate the intensity level of the user's emotion.

[0044] In some examples, the above emotion type can include but is not limited to a happy type, an angry type, a sad type, an excited type, a calm type, etc., which can be set by those skilled in the art according to needs and are not limited in the embodiments of the present application.

[0045] In some examples, the value of the above emotion intensity information can be an integer or a decimal, etc. The larger the value of the emotion intensity information, the higher the intensity level of the emotion indicated by the emotion intensity information.

[0046] For example, the value of the emotion intensity information can be a numerical value from 0.1 to 1. If the value of the emotion intensity information is 0.1, the intensity level of the emotion indicated by the emotion intensity information is the lowest; if the value of the emotion intensity information is 1, the intensity level of the emotion indicated by the emotion intensity information is the highest.

[0047] In some embodiments of the present application, the electronic device may first obtain the fourth information, and then input the fourth information into the AI model. The AI model determines the first information based on the fourth information and outputs the first information. The fourth information may include at least one of the following: historical user information, current user information. The historical user information may include, but is not limited to, historical chat information, historical images, historical input information, historical playback information, etc. The current user information may include, but is not limited to, current user input information (such as the second information in the following embodiments), current playback information, etc. Those skilled in the art can set it according to their needs, and the embodiments of the present application do not limit it here.

[0048] In some examples, a large amount of multi-modal data such as audio, video, and text of a large number of participants can be collected in advance, and the participants are allowed to rate and verify the emotional expressions presented according to the intensity of values within a preset range, so as to label the multi-modal data with detailed emotion type labels and emotion intensity information labels. Thus, the preset model can be trained using the labeled multi-modal data to obtain an AI model. Furthermore, the electronic device can input the fourth information into the AI model. The AI model determines the first information based on the fourth information and outputs the first information.

[0049] It should be noted that for the description of training the preset model using the labeled multi-modal data, reference can be made to the specific descriptions in the related art, and the embodiments of the present application will not elaborate here.

[0050] In some examples, the above historical chat information may be the content information of the user chatting using the application in the electronic device during the period from the current system time of the electronic device to a preset duration ago. The above historical images may be the images taken by the user using the electronic device during the period from the current system time of the electronic device to a preset duration ago. The above historical input information may be the information input by the user in the electronic device (such as text information, audio information, etc.) during the period from the current system time of the electronic device to a preset duration ago. The above historical playback information may be the content information of the file (such as audio or video) played by the user using the application in the electronic device during the period from the current system time of the electronic device to a preset duration ago.

[0051] In some examples, when the fourth information includes historical user information, the electronic device may obtain the historical user information from the usage information of at least one application in the electronic device.

[0052] In some examples, the above current user input information may be the information input by the user in the electronic device currently (such as text information, audio information, etc.). The above current playback information may be the content information of the file (such as audio or video) played by the user using the application in the electronic device currently.

[0053] In some examples, when the fourth information includes the current user information and the current user information includes the current user input information, the electronic device can directly collect the current input information; when the fourth information includes the current user information and the current user information includes the current playback information, the electronic device can obtain the current playback information from the usage information of at least one application in the electronic device.

[0054] Taking the fourth information including the current user information and the current user information including the current user input information as an example, the specific solution for obtaining the first information by using the AI model will be described below.

[0055] In some embodiments of the present application, the above-mentioned first information includes the emotion information. In some examples, in combination Figure 1 , such as Figure 2 shown, step 102 can be specifically implemented by the following step 102a and step 102b.

[0056] Step 102a: In response to a first input, the electronic device obtains second information, which includes at least one of the following of the user: voice information, N-frame facial image data, where N is a positive integer greater than 1.

[0057] In some embodiments of the present application, the above-mentioned voice information may be voice audio.

[0058] In some embodiments of the present application, when the second information includes the user's voice information, the electronic device can turn on the voice collection component (such as a microphone) of the electronic device to collect the voice information.

[0059] In some embodiments of the present application, the above-mentioned N-frame facial image data may be continuous image data or discontinuous image data.

[0060] It should be noted that the above-mentioned "continuous image data" can be understood as: among any two adjacent frames of image data collected, the N-frame image data for which the electronic device has not collected other image data.

[0061] In some embodiments of the present application, when the second information includes the user's N-frame facial image data, the electronic device can turn on the image collection component (such as a camera (such as a front camera)) of the electronic device to collect the N-frame facial image data.

[0062] In some examples, the electronic device can collect N-frame facial image data at a speed of X frames per second, where X is a positive integer greater than 1.

[0063] Step 102b: The electronic device determines the emotion information according to the second information through the AI model.

[0064] In some embodiments of the present application, the electronic device may input the second information into the AI model. After the electronic device inputs the second information into the AI model, the emotion analysis unit of the AI model may extract and analyze the features of the second information to obtain a determination message, and determine the emotion information according to the determination message.

[0065] Wherein, when the second information includes voice information, the determination message may include the rhythm intonation and the emotional tendency information of the voice text. In this case, the emotion analysis unit of the AI model may first extract and analyze the voice features of the voice information to obtain the rhythm intonation, then convert the voice information into the first text information, and extract and analyze the text features of the first text information to obtain the emotional tendency information of the voice text.

[0066] When the second information includes N frames of facial image data, the determination message may include the movement amplitude of the key points. In this case, the emotion analysis unit of the AI model may extract the key points from the N frames of facial image data and analyze the change of the key points in the N frames of facial image data to obtain the movement amplitude of the key points.

[0067] When the second information includes voice information and N frames of facial image data, the determination message may include the rhythm intonation, the emotional tendency information of the voice text, and the movement amplitude of the key points. In this case, the emotion analysis unit of the AI model may first extract and analyze the voice features of the voice information to obtain the rhythm intonation, convert the voice information into the first text information, extract and analyze the text features of the first text information to obtain the emotional tendency information of the voice text, then extract the key points from the N frames of facial image data and analyze the change of the key points in the N frames of facial image data to obtain the movement amplitude of the key points.

[0068] In some examples, the emotion analysis unit of the AI model may directly determine the information corresponding to the determination message as the emotion information.

[0069] As can be seen, since the electronic device can obtain at least one of the user's voice information and N frames of facial image data, the electronic device can accurately identify the user's emotion information through the AI model according to at least one of the voice information and N frames of facial image data, and output the emotion information. Therefore, the accuracy of the electronic device in identifying emotion information can be improved.

[0070] In some embodiments of the present application, the above-mentioned second information includes N frames of facial image data. In some examples, in combination with Figure 2 , as Figure 3 shown, the above-mentioned step 102b may be specifically implemented by the following step 102b1 and step 102b2.

[0071] Step 102b1: The electronic device detects the position change information of the key points on the user's face in the N-frame facial image data through the AI model.

[0072] In some embodiments of the present application, the above position change information is used to indicate the movement amplitude of the key points.

[0073] In some embodiments of the present application, the emotion analysis unit of the AI model can first detect each frame of facial image data to obtain the first position information of the key points in each frame of facial image, and determine the position change information of each key point in the N-frame facial image data according to the first position information of each key point in the N-frame facial image data.

[0074] Among them, the above first position information may be the coordinate information of the key point in the first coordinate system, and the first coordinate system may be the coordinate system corresponding to the facial image data where the key point is located. For example, the first coordinate system may be a coordinate system with two adjacent edge lines of the facial image data where the key point is located as the coordinate axes, or may be a coordinate system with the center point of the facial image data where the key point is located as the coordinate origin and straight lines parallel to two adjacent edge lines of the facial image data as the coordinate axes.

[0075] In some examples, the above position change information may include at least one change information of each key point. Thus, the emotion analysis unit of the AI model can determine one change information as the distance between the first position information of a key point in two frames of facial image data, and so on, to determine at least one change information of the one key point, and then determine at least one change information of each key point, so as to obtain the above position change information.

[0076] Step 102b2: The electronic device determines the emotion information through the AI model according to the position change information.

[0077] In some embodiments of the present application, the emotion analysis unit of the AI model determines the movement amplitude of the facial muscles corresponding to each key point according to at least one change information of each key point, so as to determine the movement amplitude of each facial muscle. Furthermore, the emotion analysis unit of the AI model can determine the information corresponding to the determined movement amplitude of each facial muscle as the emotion information.

[0078] It can be seen that since the electronic device can detect the position change information of the key points on the user's face in the N-frame face image data through the AI model, and can accurately obtain the movement amplitude of each facial muscle of the user through this position change information, therefore, the AI model can accurately determine the user's emotion information based on this position change information, that is, based on the movement amplitude of each facial muscle. In this way, the accuracy of the electronic device in recognizing emotion information can be improved.

[0079] In some embodiments of the present application, the above first information includes interaction text information. In some examples, in combination with Figure 1 , such as Figure 4 shown, the above step 102 can be specifically implemented by the following steps 102c to 102e.

[0080] Step 102c: The electronic device displays an interaction interface corresponding to the AI model in response to the first input.

[0081] In some embodiments of the present application, the above interaction interface is used for the user to interact with the AI model to obtain interaction text information.

[0082] In some embodiments of the present application, the electronic device can floatingly display the interaction interface on the current interface; or, the electronic device can switch the current interface to the interaction interface to display the interaction interface.

[0083] In some examples, the electronic device can floatingly display the interaction interface in a predetermined area of the display screen of the electronic device. Among them, the predetermined area can include any one of the following: the area within the preset range of the upper edge line of the display screen, the area within the preset range of the lower edge line of the display screen, the area within the preset range of the left edge line of the display screen, the area within the preset range of the right edge line of the display screen, and the area within the preset range of the center of the display screen.

[0084] For example, taking the electronic device as a mobile phone for illustration. As Figure 5A shown, the mobile phone displays the first interface 10, so that the user can perform a first input on the mobile phone, such as an input of tapping on the side of the mobile phone. As Figure 5B shown, after the user inputs by tapping on the side of the mobile phone, the mobile phone can floatingly display the interaction interface 12 in the predetermined area 11 on the first interface 10, so that the user can perform a second input. Among them, the predetermined area 11 can be the area within the preset range of the upper edge line of the display screen.

[0085] It should be noted that Figure 5B the predetermined area 11 is schematically shown by the area enclosed by the dashed line box, and the dashed line box may not be displayed during actual use.

[0086] In some embodiments of the present application, the above first information further includes emotion information, and the emotion information includes the emotion type of the user. The shooting method provided by the embodiments of the present application may further include the following step 201.

[0087] Step 201: The electronic device displays emotion text in the interaction interface, and the emotion text is the text that matches the emotion type among at least one preset text in the electronic device.

[0088] It should be noted that the present application embodiments do not limit the execution order of step 201 and step 102c; in one example, the electronic device may first execute step 102c and then execute step 201, that is, the electronic device may first display the interaction interface and then display the emotion text in the interaction interface; in another example, the electronic device may execute step 201 while executing step 102c, that is, the electronic device may display the emotion text in the interaction interface while displaying the interaction interface.

[0089] In some embodiments of the present application, at least one first correspondence is stored in the electronic device, and the first correspondence is the correspondence between a preset text and an emotion type. Thus, the electronic device can determine at least one text corresponding to the emotion type (i.e., the text that matches the emotion type) from at least one preset text in the electronic device according to the at least one first correspondence, and determine the above emotion text from the at least one text.

[0090] In some examples, when the emotion information is the emotion type of the user, the number of the at least one text corresponding to the emotion type may also be one, so that the electronic device can directly determine the one text as the above emotion text.

[0091] In other examples, when the emotion information includes the emotion type and emotion intensity information of the user, at least one second correspondence is further stored in the electronic device, and the second correspondence is the correspondence between a text and an emotion intensity information. Thus, the electronic device can determine a text corresponding to the emotion intensity information from at least one text as the above emotion text according to the at least one second correspondence.

[0092] In some embodiments of the present application, when the electronic device displays the emotion text, the electronic device may also play the voice corresponding to the emotion text in the first intonation.

[0093] In some examples, at least one third corresponding relationship is stored in the electronic device, and the third corresponding relationship is a corresponding relationship between a tone and an emotion type. Thus, the electronic device can determine a first tone corresponding to the emotion type according to the at least one third corresponding relationship, and play the voice corresponding to the emotion text according to the first tone.

[0094] For example, it is assumed that the emotion information includes the user's emotion type, and the emotion type is a happy type. Combining Figure 5B , such as Figure 6 shown, the mobile phone can display emotion text in the interaction interface 12, such as the text 13 "You seem to be in a great mood today. Is there anything good happening?", and play the voice corresponding to the text 13 "You seem to be in a great mood today. Is there anything good happening?" according to the first tone (such as a cheerful tone), that is, play the voice "You seem to be in a great mood today. Is there anything good happening?". Thus, the user can interact with the AI model according to the displayed text and the played voice.

[0095] It can be seen that since the electronic device can also display preset text matching the emotion type in the interaction interface, the user can quickly determine whether the emotion type recognized by the electronic device is correct by viewing the preset text, and can quickly trigger the electronic device to recognize the correct emotion type when the emotion type recognized by the electronic device is incorrect. Therefore, while improving the user experience, the accuracy of the electronic device in recognizing the user's emotion type can be improved.

[0096] In some embodiments of the present application, the above first information further includes emotion information, and the emotion information includes the user's emotion type. The shooting method provided by the embodiments of the present application may further include the following step 202.

[0097] Step 202: The electronic device displays a first emotion identifier in the interaction interface, and the first emotion identifier is used to indicate the emotion type.

[0098] It should be noted that the present application embodiments do not limit the execution order of step 202 and step 102c; in one example, the electronic device may first execute step 102c and then execute step 202, that is, the electronic device may first display the interaction interface and then display the first emotion identifier in the interaction interface; in another example, the electronic device may execute step 202 while executing step 102c, that is, the electronic device may display the first emotion identifier in the interaction interface while displaying the interaction interface.

[0099] In some embodiments of the present application, the above first emotion identifier may include at least one of the following: an emotion avatar, an emotion image, etc.

[0100] In some embodiments of the present application, at least one fourth correspondence is stored in the electronic device, and the fourth correspondence is a correspondence between a preset emotion identifier and an emotion type. Thus, the electronic device can determine, according to the at least one fourth correspondence, a first emotion identifier corresponding to the emotion type from at least one preset emotion identifier in the electronic device.

[0101] In some embodiments of the present application, the electronic device can display the first emotion identifier at a predetermined position in the interaction interface.

[0102] For example, assume that the emotion information includes the user's emotion type, and the emotion type is the happy type. Figure 6 , such as Figure 7 shown, the mobile phone can display the first emotion identifier in the interaction interface 12, such as the emotion avatar 14. The emotion avatar 14 can be a smiling avatar, so that the user can interact with the AI model according to the emotion avatar 14.

[0103] It can be seen that since the electronic device can also display the first emotion identifier matching the emotion type in the interaction interface, the user can quickly determine whether the emotion type recognized by the electronic device is correct by viewing the first emotion identifier, and can quickly trigger the electronic device to recognize the correct emotion type when the emotion type recognized by the electronic device is incorrect. Therefore, while improving the user experience, the accuracy of the electronic device in recognizing the user's emotion type can be improved.

[0104] Step 102d: The electronic device receives a second input from the user in the interaction interface.

[0105] In the embodiments of the present application, the second input is used to input interaction text information.

[0106] Exemplarily, the second input includes but is not limited to: the user's touch input to the electronic device through a touch device such as a finger or a stylus (such as a touch input to the display screen of the electronic device or the body of the electronic device), or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be specifically determined according to actual usage requirements, and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application can be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, a double click gesture, and a tap input; the click input in the embodiments of the present application can be a single click input, a double click input, or a click input of any number of times, and can also be a long press input or a short press input.

[0107] For example, the above second input may be: an input in which the user enters text information in the interaction interface, or an input in which the user inputs voice corresponding to the text information to the electronic device.

[0108] Step 102e: The electronic device responds to the second input and determines interaction text information according to the input content of the second input.

[0109] In some embodiments of the present application, when the second input is an input in which the user enters text information in the interaction interface, the electronic device may directly determine the text information entered by the user as the interaction text information. When the second input is an input in which the user inputs voice corresponding to the text information to the electronic device, the electronic device may first perform a voice-to-text conversion operation on the voice input by the user, and then determine the text information obtained by the voice-to-text conversion operation as the interaction text information.

[0110] In some embodiments of the present application, after determining the interaction text information, the electronic device may also display the interaction text information in the interaction interface, so as to facilitate the user to view whether the interaction text information is the text information required by the user.

[0111] Illustrate by way of example, in combination with Figure 7 , such as Figure 8 shown, after the user makes the second input, for example, the user inputs text information such as "Taking a walk in the park, the sun is so nice, thinking of a book by a writer, having more thoughts about life" to the mobile phone. Thus, the mobile phone can determine the text information "Taking a walk in the park, the sun is so nice, thinking of a book by a writer, having more thoughts about life" as the interaction text information, and display the interaction text information 15 of "Taking a walk in the park, the sun is so nice, thinking of a book by a writer, having more thoughts about life" in the interaction interface 12.

[0112] In this way, on the one hand, since the electronic device can directly display the interaction interface corresponding to the AI model according to the user's first input without the user performing multiple operations, the user's operation can be simplified; on the other hand, since the electronic device can display the interaction interface corresponding to the AI model, the user can quickly make the second input in this interaction interface, so that the electronic device can quickly determine the interaction text information required by the user; therefore, the efficiency of the electronic device in obtaining the interaction text information can be improved.

[0113] Step 103: The electronic device performs a shooting operation according to the first information.

[0114] In some embodiments of the present application, the electronic device may embed the first information into the image data collected by the camera of the electronic device to generate a captured image, thereby performing the shooting operation.

[0115] In some embodiments of the present application, the number of the captured images may be at least one.

[0116] In some examples, when the number of the captured images is at least two, after generating at least two captured images, the electronic device may generate a captured video according to the at least two captured images.

[0117] In some embodiments of the present application, the electronic device may add the first information to the description information of the image data and generate a captured image.

[0118] It can be understood that the description information of the generated captured image includes the first information.

[0119] An embodiment of the present application provides a shooting method. The electronic device may obtain first information according to a first input of a user by using an AI model. The first information includes at least one of the user's emotion information and the interaction text information between the user and the AI model, and perform a shooting operation according to the first information. In this way, when the user inputs to the electronic device, an image or video related to the user's emotion information and / or the interaction text information between the user and the AI model can be captured, and when the user views the captured image, the user's emotion information and / or the interaction text information between the user and the AI model can also be known, without the user manually adding information to the image. Therefore, the cumbersome operation of the user in the process of recording the relevant information of the image is avoided.

[0120] Moreover, since the captured image may include emotion information, a large amount of anonymized emotion information of users can be aggregated and analyzed, which can provide valuable information for enterprises and market research institutions. For example, understanding the emotion change trends of users in different regions and different age groups during a specific period (such as holidays and major events), so as to provide a basis for product research and development and marketing strategy formulation. For example, it is found that the emotions of young people in a certain region are generally relatively low in a certain season, which may be related to the climate or social environment factors in the region. Enterprises can specifically launch emotion regulation products or marketing activities suitable for young people in this region.

[0121] Moreover, since the captured image may include emotion information, parents or educational institutions can use this function to help children better understand and manage their emotions. For example, by reviewing photos and videos with emotion information, discussing with children their emotion feelings and coping methods in different situations, and cultivating children's emotional intelligence; some emotion challenge tasks can also be set, such as asking children to record the causes and coping strategies of their different emotions within a week, and using the mobile phone function to assist in recording and analysis to improve children's self-awareness and emotion regulation ability.

[0122] In some embodiments of the present application, before the above step 103, the shooting method provided by the embodiments of the present application may further include the following step 301 and step 302, and the above step 103 may be specifically implemented by the following step 103a.

[0123] Step 301: The electronic device displays a picture shooting control and a video shooting control. The picture shooting control is used to start the picture shooting mode, and the video shooting control is used to start the video shooting mode.

[0124] In some embodiments of the present application, the electronic device may display the picture shooting control and the video shooting control in the above interaction interface after or at the same time as obtaining the first information.

[0125] Step 302: The electronic device receives a fourth input from the user to one of the picture shooting control and the video shooting control.

[0126] In some embodiments of the present application, the above fourth input is used to select the shooting mode required by the user.

[0127] Exemplarily, the above fourth input includes but is not limited to: the user's touch input to the electronic device through a touch device such as a finger or a stylus (such as a touch input to the display screen of the electronic device or the body of the electronic device), or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be specifically determined according to actual usage requirements, and the embodiments of the present invention are not limited. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, a double click gesture, and a tap input; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, and may also be a long press input or a short press input.

[0128] For example, the above fourth input may be: the user's click input to one of the picture shooting control and the video shooting control.

[0129] For example, in combination with Figure 8 , as Figure 9 shown, the mobile phone displays the picture shooting control 16 and the video shooting control 17, so that the user can input to one of the displayed picture shooting control 16 and the video shooting control 17, such as the picture shooting control 16.

[0130] Step 103a: In response to the fourth input, the electronic device adopts the shooting mode corresponding to one control and performs a shooting operation according to the first information.

[0131] In some embodiments of the present application, the electronic device may activate a shooting application, enter a shooting mode corresponding to a control, collect image data through the shooting application, and then embed the first information into the image data to perform a shooting operation.

[0132] In some examples, the above shooting application may be understood as an application with a shooting function.

[0133] In some embodiments of the present application, when a control is a picture shooting control, the electronic device may perform a shooting operation in the picture shooting mode. It can be understood that the number of shooting images generated by the electronic device is one, and what the electronic device shoots is an image.

[0134] In some embodiments of the present application, when a control is a video shooting control, the electronic device may perform a shooting operation in the video shooting mode. It can be understood that the number of shooting images generated by the electronic device is at least two, and what the electronic device shoots is a video.

[0135] For example, in combination with Figure 9 , such as Figure 10 shown, after the user clicks on the picture shooting control 16, the mobile phone may activate the shooting application, display the interface 18 of the shooting application, and enter the picture shooting mode (such as the photo-taking mode) corresponding to the picture shooting control 16, and perform a shooting operation through the shooting application.

[0136] It should be noted that Figure 10 in

[0137] the mobile phone entering the photo-taking mode is indicated by showing a black frame on the shooting mode control. In actual use, the mobile phone may not display the black frame.

[0138] In some embodiments of the present application, before the above step 103, the shooting method provided by the embodiments of the present application may further include the following step 304, and the above step 103 may be specifically implemented by the following step 103b.

[0139] Step 304: The electronic device responds to the first input, obtains the first information using the AI model, and determines, through the AI model, a shooting filter corresponding to the first information from at least one preset filter.

[0140] In some embodiments of the present application, at least one fifth correspondence is stored in the AI model, and the fifth correspondence is a correspondence between a preset filter and an information, so that the AI model can determine a shooting filter corresponding to the first information from at least one preset filter according to the at least one fifth correspondence.

[0141] For example, assume that the first information includes the user's emotion information and the interaction text information between the user and the AI model. The emotion information includes an emotion type (such as a happy type), and the interaction text information includes the text information of "daytime amusement park". Then the AI model can, according to the at least one fifth correspondence, determine, from the at least one preset filter, the preset filter (such as the "Amusement Park Memories" filter) corresponding to the happy type and the text information of "daytime amusement park" as the shooting filter.

[0142] Step 103b: The electronic device performs a shooting operation according to the first information and the shooting filter.

[0143] Thus, it can be seen that since the electronic device can directly determine, through the AI model, the shooting filter corresponding to the first information from at least one preset filter, that is, directly determine the shooting filter that the user may need without the user manually selecting, the operation of the user during shooting can be simplified and the time consumption can be reduced. In this way, the shooting efficiency of the electronic device can be improved.

[0144] Of course, after determining the shooting filter, the electronic device can also adaptively adjust the filter parameters of the shooting filter, which will be illustrated by examples below.

[0145] In some embodiments of the present application, before the above step 103b, the shooting method provided by the embodiments of the present application may further include the following step 401 and step 402, and the above step 103b may be specifically implemented by the following step 103b1.

[0146] Step 401: The electronic device determines at least one adjustment parameter corresponding to the scene information of the environment where the electronic device is located, and the adjustment parameter is used to adjust the filter parameters of the shooting filter.

[0147] In some embodiments of the present application, the above scene information may include at least one of the following: ambient light, ambient color temperature, ambient type, etc. Among them, the ambient type may include at least one of the following: landscape type, portrait type, night scene type, etc.

[0148] In some embodiments of the present application, the above filter parameters include at least one of the following: brightness, hue, saturation, light and shadow parameters, vignetting processing parameters, etc.

[0149] In some embodiments of the present application, at least one sixth correspondence relationship is stored in the electronic device, and the sixth correspondence relationship is the correspondence relationship between the scene information and at least one parameter. Thus, the electronic device can determine at least one adjustment parameter corresponding to the scene information according to the at least one sixth correspondence relationship.

[0150] Step 402: The electronic device adjusts a corresponding filter parameter of the shooting filter according to each adjustment parameter.

[0151] In some embodiments of the present application, the electronic device can adjust the value of each filter parameter to the value of a corresponding adjustment parameter; or, the electronic device can increase or decrease the value of a corresponding adjustment parameter on the value of each filter parameter.

[0152] Step 103b1: The electronic device performs a shooting operation according to the first information and the adjusted shooting filter.

[0153] It can be seen that since the electronic device can also automatically adjust each filter parameter corresponding to the shooting filter according to each adjustment parameter corresponding to the scene information of the environment where the electronic device is located, without the need for user means to adjust, the operation of the user during shooting can be simplified and the time consumption can be reduced. Thus, the shooting efficiency of the electronic device can be improved.

[0154] Of course, after the shooting image is obtained, there may be a situation where the user wants to view the shooting image. At this time, the electronic device can display the first information while displaying the shooting image, so that the user can view the shooting image and the first information at the same time. The following will give an example.

[0155] In some embodiments of the present application, the above first information includes emotion information and interactive text information, and the emotion information includes the emotion type and emotion intensity information of the user. In some examples, in combination Figure 1 , such as Figure 11 shown, after the above step 103, the shooting method provided by the embodiments of the present application may further include the following step 501.

[0156] Step 501: When the electronic device displays the shooting image, it displays a second emotion identifier and interactive text information. The shooting image is an image obtained by performing a shooting operation, and the second emotion identifier is used to indicate the emotion type and emotion intensity information.

[0157] In some embodiments of the present application, when the user wants to view the shooting image, the user can trigger the electronic device to display a second application, and click and input on the identifier (such as a thumbnail) of the above shooting image in the interface of the second application, so that the electronic device can display the shooting image.

[0158] Wherein, the second application may be any one of the following: a file management application, an album application.

[0159] In some embodiments of the present application, the electronic device may floatingly display a second emotion identifier and interactive text information on the captured image.

[0160] For example, as Figure 12 shown, the mobile phone displays the captured image 19, and the mobile phone may display a second emotion identifier on the captured image 19, such as an emotion avatar 20 and an emotion intensity value 21. The emotion avatar 20 may be a smiling avatar, used to indicate that the user's emotion type is a happy type, and the emotion intensity value 21 is used to indicate the user's emotion intensity information; moreover, the mobile phone may display interactive text information on the captured image 19, such as the text information 22 of "Taking a walk in the park, the sun is so nice, thinking of a book by a writer, having more thoughts about life".

[0161] Thus, it can be seen that since when the electronic device displays the captured image, the electronic device can also simultaneously display the second emotion identifier and the interactive text information, so that the user can quickly determine the user's emotion type, emotion intensity information, and interactive text information when taking the captured image through the second emotion identifier and the interactive text information. Therefore, the user can quickly recall the scene when taking the captured image without the user viewing multiple captured images to determine, thereby simplifying the user's operation.

[0162] In some embodiments of the present application, in combination with Figure 1 , as Figure 13 shown, after the above step 103, the shooting method provided by the embodiments of the present application may further include the following steps 502 to 505.

[0163] Step 502: The electronic device displays an image sharing control.

[0164] In some embodiments of the present application, the above image sharing control is used to share the above captured image.

[0165] In some embodiments of the present application, when the electronic device displays the captured image, the electronic device may display an image sharing control on the captured image.

[0166] For example, as Figure 14 shown, the mobile phone displays the captured image 23, and the mobile phone may display an image sharing control 24 on the captured image 23, so that the user can perform a third input on the image sharing control 24.

[0167] Step 503: The electronic device receives a third input from the user to the image sharing control.

[0168] In some embodiments of the present application, the above third input is used to share the captured image with other electronic devices.

[0169] Exemplarily, the above third input includes, but is not limited to: the touch input of the electronic device by the user through a touch device such as a finger or a stylus (such as the touch input on the display screen of the electronic device or the body of the electronic device), or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be specifically determined according to actual usage requirements, and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application can be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, a double click gesture, and a tap input; the click input in the embodiments of the present application can be a single click input, a double click input, or a click input of any number of times, etc., and can also be a long press input or a short press input.

[0170] For example, the above third input can be: the click input of the user on the image sharing control.

[0171] Step 504: In response to the third input, the electronic device generates sharing text information according to the first information and the third information through the AI model. The third information includes at least one of the following: the shooting position information of the captured image, the person information in the captured image, and the shooting time information of the captured image.

[0172] In some embodiments of the present application, the above shooting position information may include at least one of the following: the province and city where the shooting location is located, the longitude and latitude of the shooting location, the coordinate information of the shooting location, etc. The above person information may include at least one of the following: person's name, person's gender, person's clothing information, person's hairstyle information, etc.

[0173] It should be noted that the description of generating sharing text information according to the first information and the third information through the AI model can refer to the specific description in the related art, and will not be elaborated in the embodiments of the present application.

[0174] Exemplarily, assuming that the first information includes emotion information and text interaction information, the emotion information includes an emotion type (such as a happy type), the text interaction information is the text information of "taking a walk in the park and feeling great!", and the third information includes shooting position information (such as Park C in City B, Province A), then the electronic device can input the happy type, the text information of "taking a walk in the park and feeling great!" and Park C in City B, Province A into the AI model, and generate sharing text information through the AI model, such as "The wonderful time in Park C in City B, Province A, is overflowing with happiness, and every moment is as wonderful as a dream!" text information.

[0175] Step 505: The electronic device publishes the captured image and the sharing text information through the first application.

[0176] In some embodiments of the present application, the above-mentioned first application may be a preset application or an application set by the user.

[0177] In some examples, when the first application is an application set by the user, after generating the sharing text information, the electronic device may display at least one application identifier, and each application identifier is used to indicate an application, so that the user can input the application identifier indicating the first application, so that the electronic device can publish the captured image and the sharing text information through the first application.

[0178] In some embodiments of the present application, the electronic device may open the first application, jump to the information publishing page of the first application, input the sharing text information in the text input box on the information publishing page, add the captured image to the image control to be sent on the information publishing page, and the electronic device may generate a click event on the publishing control on the information publishing page, so that the electronic device can publish the captured image and the sharing text information through the first application.

[0179] It can be seen that since the electronic device can display the image sharing control, generate the sharing text information required by the user through the AI model according to the third input of the user to the image sharing control, and publish the captured image and the sharing text information through the first application, without the need for the user to manually operate, therefore, the operation of the user in the process of publishing the captured image can be simplified and the time consumption can be reduced. In this way, the efficiency of the electronic device in publishing the captured image can be improved.

[0180] It can be understood that since the electronic device can publish the captured image and the sharing text information through the first application, this can facilitate the user to share their emotional experiences and wonderful moments with friends, family or fans, thereby further enhancing the convenience and richness of user emotional expression and social interaction.

[0181] In some embodiments of the present application, after the electronic device publishes the captured image and the sharing text information through the first application, when the captured image is displayed on other electronic devices, the other electronic devices may also display the above-mentioned second emotion identifier and the interactive text information, so that the users of the other electronic devices can also learn about the user's emotion type, emotion intensity information and the interaction content between the user and the AI model through the second emotion identifier and the interactive text information.

[0182] In some embodiments of the present application, after the electronic device publishes the captured image and the sharing text information through the first application, when the captured image is displayed on other electronic devices, the other electronic devices may also add the input content corresponding to the input to the description information of the captured image according to the input of the user of the other electronic devices.

[0183] It can be understood that since users of other electronic devices can not only view the captured images, but also interactively comment or reply based on them, or can jointly add emotional tags or recall stories, a group's common emotional memory bank can be formed, thereby promoting emotional communication among family members or friends.

[0184] Of course, there may also be a situation where the user wants to view the recall video. At this time, the electronic device can select multiple captured images and generate a recall video, which will be illustrated by examples below.

[0185] In some embodiments of the present application, in combination with Figure 3 , as Figure 13 shown, after step 103 above, the shooting method provided by the embodiments of the present application may further include the following step 601 and step 602.

[0186] Step 601: The electronic device obtains M captured images that meet the preset conditions from multiple captured images in the electronic device. M is a positive integer greater than 1. The captured image is an image obtained by performing a shooting operation. The preset conditions include any one of the following: matching the first information; the first information includes emotional information, and the change trend of the emotional information is a preset change trend.

[0187] In some embodiments of the present application, when the application interface of the second application is displayed, the application interface includes an option to view the recall video. Thus, the electronic device can obtain M captured images that meet the preset conditions from multiple captured images in the electronic device according to the input of the user for the option to view the recall video.

[0188] In some embodiments of the present application, the above M captured images all include the first information. It can be understood that the M captured images are all captured images generated by each step in the above embodiments.

[0189] In some embodiments of the present application, when the preset conditions include matching the first information and the first information includes emotional information, the electronic device can determine M captured images with matching emotional information.

[0190] For example, assuming that the preset conditions include matching the first information, the first information includes emotional information, and the emotional information includes an emotional type (such as a happy type), the electronic device can determine M captured images that all include the happy type. It can be understood that the electronic device can select M captured images in a single-dimensional manner.

[0191] In some embodiments of the present application, the above preset change trend may be a change trend from an expectant type to a happy type and then to a reluctant type, or a change trend from a reluctant type to a happy type and then to an expectant type.

[0192] In some embodiments of the present application, the above preset conditions may further include matching of third information, and the third information includes at least one of the following: shooting position information of the captured image, person information in the captured image, and shooting time information of the captured image.

[0193] For example, assume that the above preset conditions include matching of first information, the first information includes emotion information and interactive text information, the emotion information includes an emotion type (such as an excited type), and the interactive text information includes "skiing activity" text information. Then, the electronic device can determine M captured images that all include the excited type and the "skiing activity" text information. It can be understood that the electronic device can select M captured images in a multi-dimensional manner.

[0194] For another example, assume that the above preset conditions include matching of first information, the first information includes emotion information, the change trend of the emotion information is a preset change trend, and matching of third information (such as matching of shooting time information). The first information includes emotion information, and the emotion information includes emotion types (such as an expectant type, a happy type, and a reluctant type). Then, the electronic device can determine M captured images that all include the expectant type, the happy type, and the reluctant type and the shooting time information matches. It can be understood that M captured images can be selected with an emotional change trend of "from expectation to reluctance" during the vacation.

[0195] Step 602: The electronic device generates a memory video based on the M captured images.

[0196] In some embodiments of the present application, the electronic device may generate a video segment according to each captured image and generate a memory video based on the M video segments.

[0197] Thus, it can be known that since the electronic device can determine M captured images that meet the preset conditions from multiple captured images, that is, determine M captured images that meet the emotional needs of the user. Therefore, the electronic device can accurately generate a memory video that meets the emotional needs of the user based on the M captured images. In this way, the usability of the video generation function of the electronic device can be improved.

[0198] In some embodiments of the present application, the above first information includes emotion information. In some examples, the above step 602 may be specifically implemented by the following step 602a and step 602b.

[0199] Step 602a: The electronic device generates M corresponding video segments according to the M captured images and obtains M audio segments corresponding to the M emotion information in the M captured images.

[0200] It should be noted that for the description of the electronic device generating M video segments in one-to-one correspondence according to M captured images, reference can be made to the specific description in the related art, and the embodiments of the present application will not elaborate herein.

[0201] In some embodiments of the present application, the electronic device includes at least one seventh correspondence and at least one eighth correspondence. Each seventh correspondence is a correspondence between a third piece of information and multiple audio segments, and each eighth correspondence is a correspondence between an emotion information and an audio segment. Thus, the electronic device can first determine multiple audio segments according to at least one seventh correspondence, and then determine M audio segments corresponding to the M emotion information one by one from the multiple audio segments according to at least one eighth correspondence.

[0202] For example, assuming that the third piece of information includes the shooting time information of the captured image (such as Spring Festival or Christmas time information), the electronic device can first determine multiple audio segments related to Spring Festival or Christmas according to at least one seventh correspondence, and then determine M audio segments corresponding to the M emotion information one by one from the multiple audio segments according to at least one eighth correspondence.

[0203] Step 602b: The electronic device synthesizes the M video segments and the M audio segments to obtain a memory video.

[0204] In some embodiments of the present application, the electronic device can first splice the M video segments in sequence, splice the M audio segments in sequence, adjust the transition effect of the audio segments, and then synthesize the spliced video and the spliced audio to obtain a memory video.

[0205] It can be seen that since the electronic device can obtain M audio segments corresponding to the M emotion information one by one, that is, M audio segments matching the user's emotions, and synthesize the M video segments and the M audio segments, a memory video that meets the user's emotional needs can be accurately obtained. In this way, the usability of the video generation function of the electronic device can be improved.

[0206] For the shooting method provided by the embodiments of the present application, the execution subject can be a shooting device. In the embodiments of the present application, taking the shooting device executing the shooting method as an example, the shooting device provided by the embodiments of the present application is described.

[0207] Figure 15 The structural schematic diagram of the shooting device provided by the embodiments of the present application is shown. As Figure 15As shown in the figure, the photographing device 700 provided by the embodiment of the present application may include: a receiving module 701 and a processing module 702. Among them, the receiving module 701 is used to receive a first input. The processing module 702 is used to obtain first information by using an AI model in response to the first input received by the receiving module 701. The first information includes at least one of the following: the emotional information of the user, the interaction text information between the user and the AI model; and perform a photographing operation according to the first information.

[0208] The embodiment of the present application provides a photographing device. When a user inputs to an electronic device, an image or video related to the emotional information of the user and / or the interaction text information between the user and the AI model can be photographed. And when the user views the photographed image, the emotional information of the user and / or the interaction text information between the user and the AI model can also be known. There is no need for the user to manually add information to the image. Therefore, the cumbersome operation of the user in the process of recording the relevant information of the image is avoided.

[0209] In a possible implementation manner, the above first information includes emotional information. The processing module 702 is specifically used to obtain second information, and the second information includes at least one of the following of the user: voice information, N-frame facial image data, and N is a positive integer greater than 1; and determine the emotional information according to the second information through the AI model.

[0210] In a possible implementation manner, the above second information includes N-frame facial image data. The processing module 702 is specifically used to detect the position change information of the key points on the user's face in the N-frame facial image data through the AI model; and determine the emotional information according to the position change information through the AI model.

[0211] In a possible implementation manner, the above first information includes interaction text information. The photographing device 700 provided by the embodiment of the present application may further include: a display module. Among them, the display module is used to display an interaction interface corresponding to the AI model. The receiving module 701 is further used to receive a second input from the user in the interaction interface displayed by the display module. The processing module 702 is specifically used to determine the interaction text information based on the input content of the second input in response to the second input received by the receiving module 701.

[0212] In a possible implementation manner, the above first information further includes emotional information, and the emotional information includes the emotional type of the user. The display module is further used to display emotional text in the interaction interface, and the emotional text is a text matching the emotional type among at least one preset text in the photographing device 700.

[0213] In a possible implementation, the above-mentioned first information further includes emotion information, which includes the emotion type of the user. The above-mentioned display module is further configured to display a first emotion identifier in the interaction interface, and the first emotion identifier is used to indicate the emotion type.

[0214] In a possible implementation, the above-mentioned processing module 702 is further configured to, before performing a shooting operation according to the first information, determine a shooting filter corresponding to the first information from at least one preset filter through an AI model. Specifically, the processing module 702 is configured to perform a shooting operation according to the first information and the shooting filter.

[0215] In a possible implementation, the above-mentioned processing module 702 is further configured to, before performing a shooting operation according to the first information, determine at least one adjustment parameter corresponding to the scene information of the environment where the shooting device 700 is located, and the adjustment parameter is used to adjust the filter parameters of the shooting filter; and adjust a corresponding filter parameter of the shooting filter according to each adjustment parameter. Specifically, the processing module 702 is configured to perform a shooting operation according to the first information and the adjusted shooting filter.

[0216] In a possible implementation, the above-mentioned first information includes emotion information and interactive text information, and the emotion information includes the emotion type and emotion intensity information of the user. The shooting device 700 provided in the embodiment of the present application may further include: a display module. The display module is configured to, after the processing module 702 performs a shooting operation according to the first information, when displaying the shooting image, display a second emotion identifier and interactive text information, where the shooting image is an image obtained by performing the shooting operation, and the second emotion identifier is used to indicate the emotion type and emotion intensity information.

[0217] In a possible implementation, the above-mentioned processing module 702 is further configured to, after performing a shooting operation according to the first information, obtain M shooting images that meet preset conditions from multiple shooting images in the shooting device 700, where the shooting images are images obtained by performing the shooting operation, and M is a positive integer greater than 1; and generate a memory video based on the M shooting images. The above-mentioned preset conditions include any one of the following: matching the first information; the first information includes emotion information, and the change trend of the emotion information is a preset change trend.

[0218] In a possible implementation, the above-mentioned first information includes emotion information. Specifically, the processing module 702 is configured to generate M corresponding video segments according to the M shooting images, and obtain M audio segments corresponding to the M emotion information in the M shooting images; and synthesize the M video segments and the M audio segments to obtain a memory video.

[0219] In a possible implementation, the photographing device 700 provided in the embodiments of the present application may further include a display module and a publishing module. Among them, the display module is configured to display an image sharing control after the processing module 702 performs a photographing operation according to the first information. The receiving module 701 is further configured to receive a third input from the user on the image sharing control displayed by the display module. The processing module 702 is further configured to, in response to the third input received by the receiving module 701, generate a sharing text message according to the first information and the third information through an AI model. The publishing module is configured to publish the photographed image and the sharing text message generated by the processing module 702 through a first application. Among them, the third information includes at least one of the following: the photographing location information of the photographed image, the person information in the photographed image, and the photographing time information of the photographed image.

[0220] The photographing device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than the terminal. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0221] The photographing device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0222] The photographing device provided in the embodiments of the present application can implement Figures 1 to 14 each process implemented by the method embodiments. To avoid repetition, it will not be elaborated here.

[0223] In some embodiments of the present application, such as Figure 16As shown in the figure, an embodiment of the present application further provides an electronic device 800, including a processor 801 and a memory 802. A program or instruction that can run on the processor 801 is stored on the memory 802. When the program or instruction is executed by the processor 801, each step of the above-mentioned shooting method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.

[0224] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0225] Figure 17 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application.

[0226] The electronic device 100 includes but is not limited to: a radio frequency unit 101, a network module 102, an audio output unit 103, an input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, and a processor 110 and other components.

[0227] Those skilled in the art can understand that the electronic device 100 may further include a power supply (such as a battery) for supplying power to each component. The power supply can be logically connected to the processor 110 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 17 The structure of the electronic device shown in the figure does not limit the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements, which will not be elaborated here.

[0228] Among them, the user input unit 107 is used to receive a first input.

[0229] The processor 110 is used to, in response to the first input received by the user input unit 107, obtain first information by using an AI model. The first information includes at least one of the following: the user's emotion information, the interaction text information between the user and the AI model; and perform a shooting operation according to the first information.

[0230] An embodiment of the present application provides an electronic device. When a user inputs to the electronic device, an image or video related to the user's emotion information and / or the interaction text information between the user and the AI model can be shot, and when the user views the shot image, the user's emotion information and / or the interaction text information between the user and the AI model can also be known, without the user manually adding information to the image. Therefore, the cumbersome operation of the user in the process of recording the relevant information of the image is avoided.

[0231] In some embodiments of the present application, the above first information includes emotion information. The processor 110 is specifically configured to obtain second information, where the second information includes at least one of the following of the user: voice information, N-frame facial image data, and N is a positive integer greater than 1; and determine the emotion information according to the second information through an AI model.

[0232] In some embodiments of the present application, the above second information includes N-frame facial image data. The processor 110 is specifically configured to detect the position change information of the key points on the user's face in the N-frame facial image data through an AI model; and determine the emotion information according to the position change information through an AI model.

[0233] In some embodiments of the present application, the above first information includes interactive text information, and the display unit 106 is configured to display an interactive interface corresponding to the AI model.

[0234] The user input unit 107 is further configured to receive a second input from the user in the interactive interface displayed by the display unit 106.

[0235] The processor 110 is specifically configured to, in response to the second input received by the user input unit 107, determine the interactive text information based on the input content of the second input.

[0236] In some embodiments of the present application, the above first information further includes emotion information, and the emotion information includes the emotion type of the user. The display unit 106 is further configured to display emotion text in the interactive interface, and the emotion text is a text in at least one preset text in the electronic device that matches the emotion type.

[0237] In some embodiments of the present application, the above first information further includes emotion information, and the emotion information includes the emotion type of the user. The display unit 106 is further configured to display a first emotion identifier in the interactive interface, and the first emotion identifier is used to indicate the emotion type.

[0238] In some embodiments of the present application, the above processor 110 is further configured to, before performing a shooting operation according to the first information, determine a shooting filter corresponding to the first information from at least one preset filter through an AI model.

[0239] The above processor 110 is specifically configured to perform a shooting operation according to the first information and the shooting filter.

[0240] In some embodiments of the present application, the above processor 110 is further configured to, before performing a shooting operation according to the first information, determine at least one adjustment parameter corresponding to the scene information of the environment where the electronic device is located, and the adjustment parameter is used to adjust the filter parameters of the shooting filter; and adjust the corresponding one filter parameter of the shooting filter according to each adjustment parameter.

[0241] The above-mentioned processor 110 is specifically configured to perform a shooting operation according to the first information and the adjusted shooting filter.

[0242] In some embodiments of the present application, the above-mentioned first information includes emotion information and interactive text information, and the emotion information includes the user's emotion type and emotion intensity information. The display unit 106 is further configured to, after performing the shooting operation according to the first information, when displaying the shooting image, display a second emotion identifier and interactive text information, where the shooting image is an image obtained by performing the shooting operation, and the second emotion identifier is used to indicate the emotion type and emotion intensity information.

[0243] In some embodiments of the present application, the above-mentioned processor 110 is further configured to, after performing the shooting operation according to the first information, obtain M shooting images that meet preset conditions from multiple shooting images in the electronic device, where the shooting image is an image obtained by performing the shooting operation, and M is a positive integer greater than 1; and generate a memory video based on the M shooting images. Among them, the above-mentioned preset conditions include any one of the following: matching the first information; the first information includes emotion information, and the change trend of the emotion information is a preset change trend.

[0244] In some embodiments of the present application, the above-mentioned first information includes emotion information, and the processor 110 is specifically configured to generate M corresponding video segments according to the M shooting images, and obtain M audio segments corresponding to the M emotion information in the M shooting images; and synthesize the M video segments and the M audio segments to obtain a memory video.

[0245] In some embodiments of the present application, the above-mentioned display unit 106 is further configured to display an image sharing control after performing the shooting operation according to the first information.

[0246] The above-mentioned user input unit 107 is further configured to receive a third input from the user to the image sharing control displayed by the display unit 106.

[0247] The above-mentioned processor 110 is further configured to, in response to the third input received by the user input unit 107, input the first information and the third information into the AI model, and generate sharing text information according to the first information and the third information through the AI model.

[0248] The radio frequency unit 101 is configured to publish the shooting image and the sharing text information generated by the processor 110 through the first application.

[0249] Among them, the above-mentioned third information includes at least one of the following: the shooting position information of the shooting image, the person information in the shooting image, and the shooting time information of the shooting image.

[0250] It should be understood that in the embodiments of the present application, the input unit 104 may include a Graphics Processing Unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include two parts: a touch detection device and a touch controller. The other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0251] The memory 109 can be used to store software programs and various data. The memory 109 mainly includes a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 109 can include a volatile memory or a non-volatile memory, or the memory 109 can include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 109 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memories.

[0252] The processor 110 may include one or more processing units; optionally, the processor 110 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor may not be integrated into the processor 110 either.

[0253] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned shooting method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0254] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.

[0255] The embodiments of the present application further provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement each process of the above-mentioned shooting method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0256] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0257] The embodiments of the present application provide a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement each process of the shooting method embodiment as described above and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0258] It should be noted that in this text, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0259] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0260] The embodiments of the present application have been described above with reference to the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the spirit of the present application and the scope protected by the claims, can also make many forms, all of which fall within the protection scope of the present application.

Claims

1. A shooting method, characterized in that: include: receiving a first input; In response to the first input, first information is acquired using an artificial intelligence (AI) model, where the first information includes at least one of the following: user emotion information and interactive text information between the user and the AI ​​model; A photographing operation is performed according to the first information.

2. The method according to claim 1, characterized in that: The first information includes the emotion information, and the using the AI ​​model to obtain the first information includes: Acquire second information, where the second information includes at least one of the following items of the user: voice information, N frames of facial image data, where N is a positive integer greater than 1; The emotion information is determined by the AI ​​model according to the second information.

3. The method according to claim 1, characterized in that: The first information includes the interactive text information, and the using the AI ​​model to obtain the first information includes: Display the interactive interface corresponding to the AI ​​model; receiving a second input from the user in the interactive interface; In response to the second input, the interactive text information is determined based on input content of the second input.

4. The method according to claim 1, characterized in that: Before performing a shooting operation according to the first information, the method further includes: Determining, by the AI ​​model, a shooting filter corresponding to the first information from at least one preset filter; The performing a shooting operation according to the first information includes: The shooting operation is performed according to the first information and the shooting filter.

5. The method according to claim 4, characterized in that Before performing the shooting operation according to the first information and the shooting filter, the method further includes: Determining at least one adjustment parameter corresponding to scene information of an environment in which the electronic device is located, wherein the adjustment parameter is used to adjust a filter parameter of the shooting filter; According to each of the adjustment parameters, respectively adjust a corresponding filter parameter of the shooting filter; The performing the shooting operation according to the first information and the shooting filter includes: The shooting operation is performed according to the first information and the adjusted shooting filter.

6. The method according to claim 1, characterized in that The first information includes the emotion information and the interactive text information, and the emotion information includes the emotion type and emotion intensity information of the user; After performing the shooting operation according to the first information, the method further includes: In the case of displaying a captured image, a second emotion identifier and the interactive text information are displayed, the captured image is an image obtained by performing the capturing operation, and the second emotion identifier is used to indicate the emotion type and the emotion intensity information.

7. The method according to claim 1, characterized in that After performing the shooting operation according to the first information, the method further includes: Acquire M captured images satisfying a preset condition from a plurality of captured images in the electronic device, wherein the captured images are images obtained by performing the capturing operation, and M is a positive integer greater than 1; Based on the M captured images, generating a recall video; The preset condition includes any one of the following: matching the first information; The first information includes the emotional information, and a change trend of the emotional information is a preset change trend.

8. The method according to claim 7, characterized in that The first information includes the emotion information, and the generating of the recollection video based on the M captured images includes: Generate M video clips corresponding to each other according to the M captured images, and obtain M audio clips corresponding to the M pieces of emotional information in the M captured images; The M video clips and the M audio clips are synthesized to obtain the recall video.

9. The method according to claim 1, characterized in that: After performing the shooting operation according to the first information, the method further includes: Display image sharing controls; receiving a third input from a user to the image sharing control; In response to the third input, generating sharing text information according to the first information and the third information through the AI ​​model; Publishing the captured image and the shared text information through a first application; The third information includes at least one of the following: shooting location information of the shot image, person information in the shot image, and shooting time information of the shot image.

10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the shooting method according to any one of claims 1 to 9 are implemented.