A theme recommendation method, an electronic device, and a storage medium
By displaying candidate topics in the interface through a voice assistant, the problem of electronic devices being unable to quickly determine the topic when generating videos is solved, thus achieving an efficient video generation interactive process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2026-03-20
AI Technical Summary
When electronic devices generate videos, the large number of images in the gallery makes it difficult to quickly determine the theme of the video, resulting in low interactive efficiency.
The voice assistant displays candidate topics in the interface and quickly recommends topics that the user expects to generate videos. It includes interface prompts and result displays for the image recognition process and supports generating video topics after the image library is recognized.
It improves the efficiency of user interaction with the voice assistant, allowing users to intuitively select or generate video themes, thus enhancing the convenience and accuracy of video generation.
Smart Images

Figure CN119938965B_ABST
Abstract
Description
[0001] This application claims priority to the Chinese Patent Application No. 202311418278.6, filed on October 27, 2023, and entitled "A Video Production Method Based on a Large Model and an Electronic Device", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] Embodiments of the present application relate to the technical field of terminal technology, and in particular to a theme recommendation method, an electronic device, and a storage medium. BACKGROUND
[0003] With the popularity of electronic devices (such as mobile phones, tablet computers, etc.), users can store a large number of images in electronic devices, such as a large number of images stored in a gallery. When the electronic device generates a video based on the large number of images stored in the gallery, the electronic device cannot quickly determine the theme of the generated video due to the large number of images in the gallery. SUMMARY
[0004] Embodiments of the present application provide a theme recommendation method, an electronic device, and a storage medium, which can quickly recommend a theme that a user expects to generate a video in the interface of a voice assistant, improving the interaction efficiency between the user and the voice assistant.
[0005] To achieve the above-mentioned purpose, the embodiments of the present application adopt the following technical solutions:
[0006] In a first aspect, a theme recommendation method is provided, applied to an electronic device including a voice assistant. The method can include:
[0007] receiving a first operation of a user; in response to the first operation, displaying a first interface of the voice assistant, the first interface including at least one candidate theme; the candidate theme being generated by the voice assistant according to an image recognition result of the voice assistant on images in a gallery of the electronic device.
[0008] It can be understood that, in the case where the voice assistant determines that the gallery has completed the image recognition process, and the voice assistant can recommend at least one candidate theme according to the image recognition result of the gallery, the first interface (i.e., the intelligent interaction interface) displayed in response to the first operation of the user includes at least one candidate theme. Thus, the voice assistant can quickly recommend a candidate theme that the user expects to generate a video.
[0009] In a possible case of the first aspect, the above-mentioned displaying the first interface of the voice assistant in response to the first operation can include:
[0010] In response to the first operation, a second interface of the voice assistant is displayed, and the second interface includes a first control for triggering the gallery to start the image recognition processing; in response to the operation of the first control by the user, first prompt information is displayed on the second interface, and the first prompt information is used to prompt that the gallery is performing the image recognition; after the image recognition of the gallery is completed, the first interface of the voice assistant is displayed.
[0011] In one scenario, in the case where the gallery is not performing the image recognition processing, the voice assistant can guide the user to trigger the gallery to perform the image recognition processing in the second interface. In the process of the image recognition of the gallery, information prompting the image recognition of the gallery can be displayed in the second interface of the voice assistant to prompt the user that the gallery is performing the image recognition. After the image recognition of the gallery is completed, at least one candidate theme generated by the voice assistant according to the image recognition result of the gallery is displayed.
[0012] In another possible case of the first aspect, in response to the first operation, the first interface of the voice assistant is displayed, which can include:
[0013] In response to the first operation, a second interface of the voice assistant is displayed, and the second interface includes a first control for triggering the gallery to start the image recognition processing; in response to the operation of the first control by the user, a third interface of the gallery is displayed, and the third interface includes the image recognition progress of the gallery; after the image recognition of the gallery is completed, the first interface of the voice assistant is displayed.
[0014] In another scenario, in the case where the gallery is not performing the image recognition processing, the voice assistant can guide the user to trigger the gallery to perform the image recognition processing in the second interface. In the process of the image recognition of the gallery, the image recognition interface of the gallery is jumped to. After the image recognition of the gallery is completed, at least one candidate theme generated by the voice assistant according to the image recognition result of the gallery is displayed.
[0015] In another possible case of the first aspect, the third interface includes an icon of the voice assistant, and in the process of displaying the third interface of the gallery, the above-mentioned theme recommendation method can further include:
[0016] The second operation of the user on the icon of the voice assistant is received; in response to the second operation, a fourth interface of the voice assistant is displayed, and the fourth interface includes second prompt information, and the second prompt information is used to prompt the image recognition progress of the gallery.
[0017] It can be understood that, in the process of the image recognition of the gallery and the display of the image recognition interface of the gallery, the image recognition interface of the gallery can return to the fourth interface of the voice assistant (such as the intelligent interaction interface 908 shown in (d) of FIG. 8) in response to the triggering operation of the user. Figure 9
[0018] In another possible case of the first aspect, in the process of the image recognition of the gallery, the second interface displays third prompt information, and the third prompt information is used to prompt the image recognition duration of the gallery.
[0019] This can be understood as, by displaying the image recognition time in the image gallery on the intelligent interactive interface, users can intuitively determine the image recognition time and decide whether to wait for the recognition to finish on this interface. For example, Figure 12 The intelligent interactive interface 1203 shown in (b) displays a prompt message 1205, which is used to indicate the image recognition time of the image library.
[0020] In another possible scenario of the first aspect, after the image recognition is completed, the first interface of the voice assistant is displayed, including:
[0021] After the image recognition in the gallery is completed, if the voice assistant generates at least one candidate topic based on the image recognition results, a first interface including at least one candidate topic will be displayed.
[0022] For example, in Figure 9 After the image recognition process shown in (c) is completed, the voice assistant recommends three candidate topics based on the image recognition results from the image library. Figure 9 The intelligent interactive interface 907 (i.e., the first interface) shown in (e) is shown in the middle.
[0023] In another possible scenario of the first aspect, the aforementioned topic recommendation method may also include:
[0024] After the image recognition in the gallery is completed, if the voice assistant does not generate any candidate topics based on the image recognition results, the fifth interface of the voice assistant will be displayed. The fifth interface includes a fourth prompt message, which is used to inform the voice assistant that no candidate topics have been generated based on the image recognition results.
[0025] For example, in Figure 9 After the image recognition process shown in (c) is completed, the voice assistant does not recommend any candidate topics based on the image recognition results from the image library, and displays... Figure 9 The intelligent interactive interface 910 (i.e., the fifth interface) shown in (f) prompts the user that no candidate topics have been generated.
[0026] In another possible scenario of the first aspect, the aforementioned topic recommendation method may also include:
[0027] The system receives a third operation from a user on a target topic among at least one candidate topic; in response to the third operation, it displays at least one image corresponding to the target topic on a first interface, wherein the at least one image is an image in the gallery of the electronic device; it receives a fourth operation from a user to trigger the generation of a video; in response to the fourth operation, it displays a thumbnail of the target video on the first interface, wherein the target video is generated based on at least one image corresponding to the target topic.
[0028] It can be understood that after the voice assistant finds the image corresponding to the target theme from the gallery, the voice assistant can generate the video corresponding to the target theme based on the image corresponding to the target theme. For example, Figure 10 The thumbnail 1006 of the video shown in (c) is a thumbnail of a video generated based on the image corresponding to the candidate theme 1002.
[0029] In another possible implementation of the first aspect, the theme recommendation method can further include:
[0030] receiving a fifth operation of the user indicating to generate a video of a first theme; the first theme being different from the candidate theme; in response to the fifth operation, displaying fifth prompt information in the first interface, the fifth prompt information being used to prompt that the gallery does not include an image corresponding to the first theme.
[0031] It can be understood that after the voice assistant receives an operation of the user indicating to generate a video of a theme different from the candidate theme, the voice assistant cannot find an image corresponding to the first theme, resulting in the voice assistant being unable to generate a video corresponding to the first theme. The voice assistant can display prompt information in the smart interaction interface to prompt the user to take more images.
[0032] In another possible implementation of the first aspect, the theme recommendation method can further include:
[0033] in response to the fifth operation, displaying sixth prompt information in the first interface, the sixth prompt information being used to prompt the user to generate a video corresponding to a theme other than the first theme.
[0034] It can be understood that when the voice assistant determines that it is unable to generate a video corresponding to the first theme, the voice assistant can recommend a theme for which a video can be generated in the smart interaction interface.
[0035] In another possible implementation of the first aspect, the candidate theme can include a character theme, and the method further includes:
[0036] receiving a sixth operation of the user on the character theme in the at least one candidate theme; in response to the sixth operation, displaying a plurality of candidate characters in the first interface; receiving a seventh operation of the user on a target character in the plurality of candidate characters; in response to the seventh operation, displaying at least one portrait corresponding to the target character in the first interface, the at least one portrait being an image in a gallery of the electronic device; receiving an eighth operation of the user triggering generation of a character video; and in response to the eighth operation, displaying a thumbnail of the character video in the first interface, the character video being generated according to the at least one portrait corresponding to the target character.
[0037] It can be understood that, in the case where the voice assistant determines that the same or similar objects corresponding to the candidate theme are multiple according to the keywords of the candidate theme, the images of the multiple objects can be displayed in the intelligent interaction interface, and the voice assistant can generate a video of the target object in response to a triggering operation of the user. Here, the object is a person, but the object can also be a pet, a building, etc., which is not limited here.
[0038] In another possible implementation of the first aspect, after receiving the first operation of the user, the method further includes:
[0039] When the voice assistant determines that the gallery has not performed the image recognition processing, in response to the first operation, a sixth interface of the voice assistant is displayed, and the sixth interface includes at least one preset theme.
[0040] It can be understood that, in the case where the gallery has not performed the image recognition processing on the images in the gallery, the mobile phone receives the first operation of the user, and in response to the first operation, at least one preset theme is displayed in the sixth interface of the voice assistant.
[0041] In another possible implementation of the first aspect, the theme recommendation method can further include:
[0042] The ninth operation of the user on a second theme in the at least one preset theme is received; in response to the ninth operation, seventh prompt information is displayed in the sixth interface, the seventh prompt information is used to prompt that the gallery is performing the image recognition; after the image recognition of the gallery is completed, a seventh interface is displayed; the seventh interface includes at least one image found by the voice assistant according to the image recognition result of the gallery.
[0043] It can be understood that, after the voice assistant receives the triggering operation of the user on the second theme, the gallery performs the image recognition processing. Here, when the gallery performs the image recognition processing, the gallery can perform the image recognition processing only on the images related to the second theme, or the gallery can also perform the image recognition processing on all the images in the gallery.
[0044] For example, assuming that the second theme is “generate a video of the trip last weekend”, the gallery can perform the image recognition processing only on the images in the gallery with the time identifier of last weekend, so as to improve the efficiency of generating the video by means of targeted image recognition.
[0045] In another possible implementation of the first aspect, the theme recommendation method can further include:
[0046] After the image recognition of the gallery is completed, eighth prompt information is displayed in the sixth interface, the eighth prompt information is used to prompt that no image corresponding to the second theme is found, and the seventh prompt information is not displayed in the sixth interface.
[0047] For example, the voice assistant determines that the gallery is performing the image recognition, and displays Figure 15The intelligent interaction interface 1503 shown in (c) displays the seventh prompt information 1505, prompting that the gallery is in the process of image recognition. After the voice assistant determines that the gallery completes the image recognition, the eighth prompt information 1506 is displayed in the intelligent interaction interface shown in (d). Figure 15 The intelligent interaction interface 1503 shown in (c) displays the seventh prompt information 1505, prompting that the gallery is in the process of image recognition. After the voice assistant determines that the gallery completes the image recognition, the eighth prompt information 1506 is displayed in the intelligent interaction interface shown in (d).
[0048] In another possible implementation of the first aspect, the subject recommendation method can further include:
[0049] In the process of image recognition of the gallery, if there is an abnormality in image recognition, an abnormality prompt information is displayed on the sixth interface of the electronic device, and the abnormality prompt information is used to prompt that there is an abnormality in image recognition of the gallery.
[0050] That is, in the process of image recognition of the gallery, there can be a situation that the battery is insufficient, a cloned image, or the like, resulting in an abnormality in image recognition.
[0051] In another possible implementation of the first aspect, the subject recommendation method can further include:
[0052] When the voice assistant determines that there are new images in the gallery that have not been subjected to image recognition, the ninth prompt information is displayed on the first interface, and the ninth prompt information is used to prompt the user to perform image recognition on the new images in the gallery that have not been subjected to image recognition.
[0053] It can be understood that, after the voice assistant determines that the gallery completes the image recognition process, the voice assistant determines that there are new images in the gallery that have not been subjected to image recognition, and the voice assistant can prompt the user to trigger the gallery to perform image recognition again.
[0054] In another possible implementation of the first aspect, the first operation is an operation of triggering an icon of the voice assistant to trigger the intelligent albuming function, or the first operation is an operation of triggering a desktop card of the electronic device to trigger the intelligent albuming function, the desktop card including an entry for triggering the intelligent albuming function, or the first operation is an operation of triggering an entry of the first interface provided by the gallery to trigger the intelligent albuming function.
[0055] In a second aspect, the present application provides an electronic device, comprising: one or more processors; a memory; wherein the memory stores one or more computer programs, and the one or more computer programs include instructions that, when executed by the electronic device, cause the electronic device to perform the subject recommendation method of any one of the above first aspects.
[0056] In a third aspect, the present application provides a computer-readable storage medium, which stores instructions, and the instructions, when executed on an electronic device, cause the electronic device to perform the subject recommendation method of any one of the first aspect.
[0057] In a fourth aspect, the present application provides a computer program product, which comprises computer instructions, when the computer instructions are executed on an electronic device, the electronic device is caused to perform the subject recommendation method according to any one of the first aspect.
[0058] It can be understood that the electronic device according to the second aspect, the computer storage medium according to the third aspect, and the computer program product according to the fourth aspect are all used to execute the corresponding method provided above, and thus the beneficial effects achieved thereby can refer to the beneficial effects of the corresponding method provided above, which will not be described herein again. BRIEF DESCRIPTION OF DRAWINGS
[0059] Figure 1 A structural schematic diagram of an electronic device provided by an embodiment of the present application;
[0060] Figure 2 A software structural diagram of an electronic device provided by an embodiment of the present application;
[0061] Figure 3 An example of user triggering to enter an intelligent interaction interface provided by an embodiment of the present application Figure 1 ;
[0062] Figure 4 An example of user triggering to enter an intelligent interaction interface provided by an embodiment of the present application Figure 2 ;
[0063] Figure 5 An example of user triggering to enter an intelligent interaction interface provided by an embodiment of the present application Figure 3 ;
[0064] Figure 6 An example of user triggering to enter an intelligent interaction interface provided by an embodiment of the present application Figure 4 ;
[0065] Figure 7 An example of user triggering to enter an intelligent interaction interface provided by an embodiment of the present application Figure 5 ;
[0066] Figure 8 An example of user triggering to enter an intelligent interaction interface provided by an embodiment of the present application Figure 6 ;
[0067] Figure 9 A subject recommendation example provided by an embodiment of the present application Figure 1 ;
[0068] Figure 10 A subject recommendation example provided by an embodiment of the present application Figure 2 ;
[0069] Figure 11 Subject recommendation examples provided for embodiments of the present application Figure 3 ;
[0070] Figure 12 Subject recommendation examples provided for embodiments of the present application Figure 4 ;
[0071] Figure 13 Subject recommendation examples provided for embodiments of the present application Figure 5 ;
[0072] Figure 14 Subject recommendation examples provided for embodiments of the present application Figure 6 ;
[0073] Figure 15 Subject recommendation examples provided for embodiments of the present application Figure 7 ;
[0074] Figure 16 Subject recommendation examples provided for embodiments of the present application Figure 8 ;
[0075] Figure 17 Subject recommendation examples provided for embodiments of the present application Figure 9 . DETAILED DESCRIPTION
[0076] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B; "and / or" in this document only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which means that there are three cases of A alone, A and B together, and B alone.
[0077] Hereinafter, the terms "first" and "second" are used only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "multiple" is two or more.
[0078] In the embodiments of the present application, the words "exemplary" or "for example" are used to mean serving as an example, instance, or illustration. Any embodiment or design presented as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0079] This application provides a topic recommendation method applied to an electronic device including a voice assistant. The method involves the electronic device receiving a first operation from a user and, in response to the first operation, displaying a first interface (i.e., an intelligent interactive interface) of the voice assistant. The first interface includes at least one candidate topic, which is generated by the voice assistant based on image recognition results from the electronic device's image library.
[0080] This can be understood as follows: after the voice assistant determines that the image library has completed the image recognition process, the voice assistant responds to the user's first action and displays a first interface that includes at least one candidate topic.
[0081] For example, the topic recommendation method provided in this application embodiment can be applied to electronic devices with displays such as mobile phones, tablets, personal computers (PCs), personal digital assistants (PDAs), smartwatches, netbooks, wearable electronic devices, augmented reality (AR) devices, virtual reality (VR) devices, in-vehicle devices, and smart cars. This application embodiment does not impose any limitations on this.
[0082] like Figure 1 As shown, Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0083] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0084] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than shown, or combine certain components, or split certain components, or different arrangement of components. The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0085] The processor 110 can include one or more processing units. For example, the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices or integrated into one or more processors.
[0086] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.
[0087] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can save instructions or data that the processor 110 has just used or repeatedly uses. If the processor 110 needs to use the instructions or data again, it can be directly called from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thereby improving the efficiency of the system.
[0088] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0089] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a limitation on the structure of the electronic device 100. In some other embodiments of the present application, the electronic device 100 can also use different interface connection methods or a combination of multiple interface connection methods.
[0090] The charging management module 140 is configured to receive a charging input from a charger. The charger can be a wireless charger or a wired charger.
[0091] The power management module 141 is configured to connect the battery 142 and the charging management module 140. The power management module 141 receives the input of the battery 142 and / or the charging management module 140 to supply power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160, etc. The power management module 141 can also be configured to monitor parameters such as battery capacity, battery cycle number, battery health status (leakage, impedance), etc.
[0092] The wireless communication function of the electronic device 100 can be realized by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc.
[0093] Antennas 1 and 2 are used for transmitting and receiving electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization of antennas. For example, antenna 1 can be multiplexed as a diversity antenna for wireless local area networks. In some other embodiments, antennas can be used in combination with tuning switches.
[0094] Mobile communication module 150 can provide solutions for wireless communication including 2G / 3G / 4G / 5G, etc. applied on electronic device 100. Mobile communication module 150 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. Mobile communication module 150 can receive electromagnetic waves by antenna 1, and perform filtering, amplification, etc. on the received electromagnetic waves, and transmit the processed signals to a modem processor for demodulation. Mobile communication module 150 can also amplify signals modulated by the modem processor, and convert the signals into electromagnetic waves radiated by antenna 1. In some embodiments, at least part of the functional modules of mobile communication module 150 can be arranged in processor 110. In some embodiments, at least part of the functional modules of mobile communication module 150 can be arranged in the same device as at least part of the modules of processor 110.
[0095] Wireless communication module 160 can provide solutions for wireless communication including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. applied on electronic device 100. Wireless communication module 160 can be one or more devices integrated with at least one communication processing module. Wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and transmits the processed signals to processor 110. Wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification on the signals, and convert the signals into electromagnetic waves radiated by antenna 2.
[0096] The electronic device 100 implements a display function through a GPU, a display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.
[0097] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diode (QLED), etc. In some embodiments, the electronic device 100 can include 1 or N display screens 194, N being a positive integer greater than 1.
[0098] In the embodiments of the present application, the display screen 194 is used to display the display interface of the voice assistant. When the user interacts with the voice assistant to control the voice assistant to perform a certain function, the electronic device 100 receives the user's operation on the display interface of the voice assistant, and in response to the user's operation, the display screen 194 displays the corresponding execution status.
[0099] The external memory interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external storage card communicates with the processor 110 through the external memory interface 120 to implement a data storage function. For example, music, video, etc. Files are saved in the external storage card.
[0100] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 performs various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program (such as a sound playing function, an image playing function, etc.) required by a function, etc. The data storage area can store data (such as audio data, a phone book, etc.) created during the use of the electronic device 100, etc. In addition, the internal memory 121 can include a high-speed random access memory, and can further include a non-volatile memory such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0101] The electronic device 100 can realize an audio function through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, the application processor, etc. For example, music playing, recording, etc.
[0102] The audio module 170 is used to convert digital audio information into an analog audio signal output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode an audio signal. In some embodiments, the audio module 170 can be disposed in the processor 110, or part of the functions of the audio module 170 can be disposed in the processor 110.
[0103] The speaker 170A, also known as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or listen to a hands-free call through the speaker 170A.
[0104] The receiver 170B, also known as a "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the voice can be heard by placing the receiver 170B close to the ear.
[0105] The microphone 170C, also known as a "microphone", "sound collector", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can make a sound by placing the mouth close to the microphone 170C, and input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, in addition to collecting sound signals, it can also realize a noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C, which can realize the collection of sound signals, noise reduction, and can also identify the sound source, realize the directional recording function, etc.
[0106] For example, after the microphone 170C of the electronic device receives voice information indicating starting the voice assistant by user voice, the electronic device starts the voice assistant. During the voice interaction between the user and the voice assistant, the microphone 170C can receive voice information indicating that the voice assistant performs a certain function by user voice.
[0107] The earphone interface 170D is used to connect a wired earphone. The earphone interface 170D can be a USB interface 130, or a 3.5mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0108] The keys 190 include a power-on key, a volume key, and the like. The keys 190 can be mechanical keys. Alternatively, the keys 190 can be touch keys. The electronic device 100 can receive key input and generate key signal input related to user settings and function control of the electronic device 100.
[0109] The motor 191 can generate a vibration prompt. The motor 191 can be used for incoming call vibration prompts, or for touch vibration feedback. For example, touch operations on different applications (such as taking pictures, playing audio, and the like) can correspond to different vibration feedback effects. Touch operations on different regions of the display screen 194 can also correspond to different vibration feedback effects. Different application scenarios (such as time reminders, receiving information, alarms, games, and the like) can also correspond to different vibration feedback effects. The touch vibration feedback effects can also be customizable.
[0110] Of course, it can be understood that the above Figure 1 The above description is only an example when the electronic device is in the form of a mobile phone. If the electronic device is in the form of a tablet computer, a handheld computer, a personal computer, a wearable device (such as a smart watch, a smart bracelet), or other devices, the structure of the electronic device can include fewer structures than those shown in the above Figure 1 , or can include more structures than those shown in the above Figure 1 , without limitation.
[0111] It can be understood that, in addition to hardware support, the implementation of the functions of the electronic device also requires the cooperation of software. The software system of the electronic device can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservice architecture, or a cloud architecture. The embodiment of the present application takes the Android system with a layered architecture as an example to illustrate the software structure of the electronic device.
[0112] Figure 2 A software structure diagram of an electronic device provided in an embodiment of the present application.
[0113] It can be understood that the layered architecture divides software into several layers, each of which has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system can include an application layer (referred to as an application layer), an application framework layer (referred to as a framework layer), a system library, and a kernel layer.
[0114] The above-mentioned application layer can include a series of application packages.
[0115] As shown in Figure 2 , the application package can include a system application. The system application refers to an application that is set in the electronic device before the electronic device is shipped. For example, the system application can include a voice assistant, a camera, a gallery, a calendar, music, a memo, and weather programs.
[0116] In an embodiment of the present application, the voice assistant receives the user's interactive operation with the voice assistant, and in response to the user's operation, displays the recommended candidate theme in the intelligent interaction interface of the voice assistant. The candidate theme includes at least one theme recommended by the voice assistant according to the image recognition result of the gallery.
[0117] The gallery stores a plurality of images, and the gallery can perform image recognition processing on the stored plurality of images. For example, the gallery performs clustering processing on the plurality of images, so that the voice assistant can generate at least one candidate theme according to the image recognition result of the gallery.
[0118] It should be explained that the present application generates at least one candidate theme according to the image recognition result of the gallery by the voice assistant only as an example. In the present scheme, any application or module that can obtain the picture stored by the electronic device can perform the image recognition processing process, which is not limited here.
[0119] The application package can also include a third-party application, which refers to an application installed by the user after downloading the installation package from the application store (or application market). For example, a map application, a take-out application, a reading application (such as an e-book), a social application, and a travel application.
[0120] The above-mentioned application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications of the application layer. The application framework layer includes some pre-defined functions.
[0121] As shown in Figure 2As shown, the application framework layer can include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc.
[0122] The window manager is used to manage windows. The window manager can acquire the size of the display screen, determine whether there is a status bar, lock the screen, and intercept the screen, etc.
[0123] The content provider is used to store and acquire data, and make the data accessible to the application. The data can include videos, images, audios, dialed and received calls, browsing history and bookmarks, phonebook, etc.
[0124] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc.
[0125] The phone manager is used to provide the communication function of the electronic device. For example, the management of the call state (including call connection, call hang-up, etc.).
[0126] The resource manager provides various resources for the application, such as localized strings, icons, pictures, layout files, video files, etc.
[0127] The notification manager enables the application to display notification information in the status bar, which can be used to convey notification type messages that can automatically disappear after a short stay without user interaction.
[0128] The system library can include a plurality of functional modules. For example: surface manager, media library, three-dimensional graphics processing library (for example: OpenGL ES), two-dimensional graphics engine (for example: SGL), etc.
[0129] The kernel layer is the layer between hardware and software. The kernel layer at least contains display driver, camera driver, audio driver and sensor driver.
[0130] In the embodiment of the present application, the electronic device receives a triggering operation of a user on the icon of the voice assistant, and in response to the triggering operation of the user, the display driver controls the display screen to display the display interface of the voice assistant.
[0131] Based on the above hardware architecture and software architecture, the following takes a mobile phone as an example to introduce the subject recommendation method provided in the embodiment of the present application.
[0132] In a possible case of the embodiments of the present application, the mobile phone receives an operation of triggering the user to start the voice assistant, and in response to the operation of the user, the mobile phone can start the voice assistant. After starting the voice assistant, the user can use various functions provided by the voice assistant. For example, after the mobile phone receives an operation of triggering the user to enter an intelligent interaction interface (i.e., a first operation), in response to the operation of the user, the mobile phone displays the intelligent interaction interface. The intelligent interaction interface is an interface of a wisdom film making function provided by the voice assistant. With the wisdom film making function of the voice assistant, the user can create a video. In some embodiments, at least one candidate theme recommended by the voice assistant of the mobile phone according to image recognition results of images in a gallery can be displayed in the intelligent interaction interface (a first interface).
[0133] In the gallery, there can be multiple images. The images in the gallery can be images captured by the user using the camera of the mobile phone, or images received by the mobile phone from other electronic devices in real time or in advance, or images downloaded by the mobile phone from the network in real time or in advance, and the like. In the embodiments of the present application, the images in the gallery are not limited to the above-mentioned acquisition manners, and the acquisition manner of the images in the gallery is not limited in the embodiments of the present application.
[0134] It can be understood that the gallery of the mobile phone includes multiple images. After the mobile phone receives an operation of triggering the user to enter an intelligent interaction interface, the voice assistant of the mobile phone can recommend at least one candidate theme according to the multiple images in the gallery, and display the at least one candidate theme in the intelligent interaction interface of the voice assistant for the user to create a video. For example, the gallery includes images of a weekend trip and images of a child dancing. The voice assistant can recommend two candidate themes "making a video of a weekend trip" and "making a video of a child dancing" according to the images of the weekend trip and the images of the child dancing in the gallery. The voice assistant can display the candidate themes "making a video of a weekend trip" and "making a video of a child dancing" in the intelligent interaction interface for the user to create a video.
[0135] For example, the mobile phone receives an operation of triggering the user to select a target theme from the at least one candidate theme. In response to the operation of the user, the voice assistant can obtain images corresponding to the target theme, and display at least one image corresponding to the target theme in the intelligent interaction interface. The mobile phone receives an operation of triggering the user to generate a video corresponding to the target theme. In response to the operation of the user, the voice assistant can generate a video corresponding to the target theme based on the at least one image corresponding to the target theme. The voice assistant can also display a thumbnail of the video corresponding to the target theme in the intelligent interaction interface of the voice assistant.
[0136] In the embodiments of the present application, the user can trigger the mobile phone to enter the intelligent interaction interface in various ways.
[0137] In an embodiment, the user can wake up the voice assistant by voice, long-pressing the power button, etc. After the voice assistant is woken up, the phone can display the display interface of the voice assistant. The user can trigger the smart interaction interface through the display interface of the voice assistant. For example, the phone displays the main interface. After the phone receives the user's voice to wake up the voice assistant, the phone displays the interface 300 shown in (a) of FIG. 3. For example, after the phone receives the user's voice information "Xiao Y, Xiao Y" to wake up the voice assistant, the phone displays the interface 300 shown in (a) of FIG. 3. Alternatively, after the phone receives the user's operation of long-pressing the power button 301, the phone displays the interface 300 shown in (a) of FIG. 3 in response to the user's long-pressing operation. Then, after the phone receives the user's voice information "generate a video of children dancing from small to old", the phone can recognize the user's voice information and display the recognized voice information, such as the interface 302 shown in (b) of FIG. 3. After the phone determines that the user's voice information input is complete, the phone can display the interface 303 shown in (c) of FIG. 3. Figure 3 Figure 3 Figure 3 Figure 3 Figure 3
[0138] It can be understood that the voice assistant can provide a plurality of functions for the user, and these functions include the video generation function of the present application, such as the above-mentioned smart video generation function. The display interfaces of different functions of the voice assistant are different. In addition, the voice assistant can provide the user with an entry to the display interface of the corresponding function through a menu to trigger the phone to display the interface of the corresponding function.
[0139] For example, after the user triggers the interface of the voice assistant by voice, the phone displays the interface of the conversation function of the voice assistant by default, such as the above-mentioned interface 303. The interface 303 includes a menu (such as the area 306 shown in (c) of FIG. 3), and the menu includes the function topics of the functions, such as "conversation", "recommendation", "smart video generation", "text creation", etc. After the voice assistant receives the user's left or right sliding operation or voice instruction, the voice assistant can switch to display the interface of each function. Figure 3
[0140] In some implementations, when the phone displays the interface 303 of the voice assistant, such as when the voice assistant of the phone receives the user's operation of sliding the interface 303 to the left, the phone displays the interface 305 shown in (e) of FIG. 3 in response to the user's sliding operation. For example, after the voice assistant of the phone receives the user's triggering operation on the card in the interface 305, the phone displays the interface 306 shown in (f) of FIG. 3 in response to the user's triggering operation. Figure 3 Figure 3 The smart interaction interface 304 shown in (d). Also for example, after the voice assistant of the mobile phone receives the operation of the user sliding the interface 303 to the left, the voice assistant displays the smart interaction interface 304 in response to the sliding operation of the user. Figure 3 The smart interaction interface 304 shown in (d).
[0141] It needs to be explained that, due to the limited size of the display screen, Figure 3 The entry of each function in the area 306 shown in (c) is only an example, and the user can display the entry of other functions through the operation of sliding left and right.
[0142] In addition, Figure 3 The interface displayed by the mobile phone and the interfaces of the voice assistant shown in (a) to (c) are only described by way of example, and the triggering operation of the user is also only an example, which is not limited in the embodiments of the present application. For example, assuming that the interface 303 shown in (c) is as follows: Figure 3 The smart interaction interface 304 shown in (d). Also for example, after the voice assistant of the mobile phone receives the operation of the user sliding the interface 303 to the left, the voice assistant displays the smart interaction interface 304 in response to the sliding operation of the user. Figure 3 The interface 305 shown in (e).
[0143] In another embodiment, the user can start the voice assistant by triggering the icon of the voice assistant on the desktop of the mobile phone, and then trigger the smart interaction interface through the display interface of the voice assistant. For example, referring to Figure 4 (a), the mobile phone displays the main interface 401, and the main interface 401 includes the icon 402 of the voice assistant. After the mobile phone receives the triggering operation of the user on the icon 402 of the voice assistant, the mobile phone displays the recommended interface 403 in response to the triggering operation of the user. Figure 4 (a), the mobile phone displays the main interface 401, and the main interface 401 includes the icon 402 of the voice assistant. After the mobile phone receives the triggering operation of the user on the icon 402 of the voice assistant, the mobile phone displays the recommended interface 403 in response to the triggering operation of the user. Figure 4 (c), when the mobile phone receives the triggering operation of the user on any card in the interface 404, the mobile phone displays the smart interaction interface 405 in response to the triggering operation of the user. Figure 4 (c), when the mobile phone receives the triggering operation of the user on any card in the interface 404, the mobile phone displays the smart interaction interface 405 in response to the triggering operation of the user.
[0144] Each card cover displayed in the interface 404 corresponds to at least one image, and the cover of the card displayed in the interface 404 is any image in the image gallery corresponding to the card. The number of cards displayed in the interface 404 can be pre-set, for example, the number of cards can be set to 20, 24, etc. Due to the limited size of the display screen, all the cards are not completely displayed in the interface 404. The voice assistant receives the up and down sliding operations of the user on the interface 404, and in response to the sliding operation of the user, other cards can be displayed on the display screen.
[0145] In one case, the order of the plurality of cards displayed in the interface 404 can be sorted by the voice assistant according to the recognition result of the photo album in the case that the photo album recognition is completed. For example, the voice assistant can sort the generated cards according to the number of images that can generate the cards in the recognition result. For example, assuming that the card 406 and the card 407 are both cards generated according to the images in the photo album, the card 406 corresponds to 10 images, and the card 407 corresponds to 6 images, the card 406 is arranged in front of the card 407.
[0146] In another case, in the case that the photo album does not perform recognition on the images in the photo album, the plurality of cards displayed in the interface 404 can be randomly sorted, or can also be sorted according to a preset rule. For example, the preset rule can be that the card corresponding to the portrait image is sorted in front, and the card corresponding to the landscape image and the pet image is sorted in back, and the like.
[0147] In another embodiment, the user can enter the intelligent interaction interface by triggering the desktop card. The desktop card can be generated by the voice assistant based on the recognition result of the images in the photo album. That is, the desktop card includes an entrance of the intelligent photo function. For example, referring to FIG. 5(a), in the case that the mobile phone displays the interface 501, the mobile phone receives the triggering operation of the user on the desktop card 502 in the interface 501, and in response to the triggering operation of the user, the intelligent interaction interface 503 shown in FIG. 5(b) is displayed. Figure 5 Figure 5
[0148] It can be understood that since the desktop card in the interface 501 is generated by the voice assistant based on the recognition result of the images in the photo album, the desktop card includes the entrance of the candidate theme "park scenery" and the "intelligent photo function" (such as the "create video" control in the interface 501), and the mobile phone can display the intelligent interaction interface in response to the triggering operation of the user on the "create video" control in the desktop card 501. In the intelligent interaction interface, the images found by the voice assistant are displayed. Figure 5 The triggering of the display of the intelligent interaction interface through the entrance of the "intelligent photo function" in the desktop card is only an example.
[0149] In the process that the mobile phone displays the intelligent interaction interface, the voice assistant can also return to the main interface of the mobile phone in response to the operation of the user. For example, the voice assistant receives the triggering operation of the user on the "back" control 504, or receives the downward sliding of the user on the intelligent interaction interface 503, and in response to the triggering operation of the user, the intelligent interaction interface 503 is exited, and the interface 501 is displayed, that is, the interface shown in FIG. 5(a) is returned to the interface shown in FIG. 5(b). Figure 5 Figure 5
[0150] For example,Figure 5 The smart interaction interface 503 shown in (b) also provides a function (i.e. icon 505) of returning to the recommendation interface. The voice assistant can return to the recommendation interface in response to the operation of the user. For example, the voice assistant receives the triggering operation of the user on the icon 505 in the smart interaction interface 503, and displays the recommendation interface 506 shown in (c) in response to the triggering operation of the user. Figure 5 The phone receives the left and right sliding operation of the user on the recommendation interface 506, and can switch the display interface.
[0151] In another embodiment, the user can enter the smart interaction interface through the search mode in the negative one screen. For example, referring to (a) in FIG. 6, the phone searches the smart album in the phone in response to the input operation of the user, and displays the smart album card 602 in the interface 601 according to the search result. The phone receives the triggering operation of the user on the smart album card 602, and displays the smart interaction interface 603 shown in (b) in response to the triggering operation of the user. Figure 6 The phone receives the left and right sliding operation of the user on the recommendation interface 506, and can switch the display interface. Figure 6 In another embodiment, the user can enter the smart interaction interface through the search mode in the negative one screen. For example, referring to (a) in FIG. 6, the phone searches the smart album in the phone in response to the input operation of the user, and displays the smart album card 602 in the interface 601 according to the search result. The phone receives the triggering operation of the user on the smart album card 602, and displays the smart interaction interface 603 shown in (b) in response to the triggering operation of the user.
[0152] In another embodiment, the user can enter the smart interaction interface through the search mode in the negative one screen. For example, referring to (a) in FIG. 6, the phone searches the smart album in the phone in response to the input operation of the user, and displays the smart album card 602 in the interface 601 according to the search result. The phone receives the triggering operation of the user on the smart album card 602, and displays the smart interaction interface 603 shown in (b) in response to the triggering operation of the user. Figure 7 For example, referring to (a) in FIG. 7, when the phone displays the creation interface 701 of the gallery, the phone, such as the gallery of the phone, receives the triggering operation of the user on the “smart album” control 702, and displays the smart interaction interface 703 shown in (b) in response to the triggering operation of the user. Figure 7 The voice assistant receives the triggering operation of the user on the return control 704, or receives the downward sliding operation of the user on the smart interaction interface 703, and exits the smart interaction interface 703 and returns to the creation interface 701 in response to the operation of the user, i.e. returns to (a) in (b). Figure 7 The voice assistant receives the triggering operation of the user on the return control 704, or receives the downward sliding operation of the user on the smart interaction interface 703, and exits the smart interaction interface 703 and returns to the creation interface 701 in response to the operation of the user, i.e. returns to (a) in (b). Figure 7 The voice assistant receives the triggering operation of the user on the return control 704, or receives the downward sliding operation of the user on the smart interaction interface 703, and exits the smart interaction interface 703 and returns to the creation interface 701 in response to the operation of the user, i.e. returns to (a) in (b).
[0153] In another embodiment, the user can enter the smart interaction interface through the search mode in the negative one screen. For example, referring to (a) in FIG. 6, the phone searches the smart album in the phone in response to the input operation of the user, and displays the smart album card 602 in the interface 601 according to the search result. The phone receives the triggering operation of the user on the smart album card 602, and displays the smart interaction interface 603 shown in (b) in response to the triggering operation of the user. Figure 8 For example, referring to (a) in FIG. 7, when the phone displays the creation interface 701 of the gallery, the phone, such as the gallery of the phone, receives the triggering operation of the user on the “smart album” control 702, and displays the smart interaction interface 703 shown in (b) in response to the triggering operation of the user. Figure 8The smart interaction interface 803 shown in (b) is displayed. Similarly, the mobile phone receives the triggering operation of the user on the control 804, or receives the user sliding the smart interaction interface 803, and in response to the triggering operation of the user, the smart interaction interface 803 is exited and the main interface 801 is returned to, that is, the first interface 801 shown in (a) is displayed. Figure 8 The smart interaction interface 803 shown in (b) is displayed. Similarly, the mobile phone receives the triggering operation of the user on the control 804, or receives the user sliding the smart interaction interface 803, and in response to the triggering operation of the user, the smart interaction interface 803 is exited and the main interface 801 is returned to, that is, the first interface 801 shown in (a) is displayed. Figure 8 The smart interaction interface 803 shown in (b) is displayed. Similarly, the mobile phone receives the triggering operation of the user on the control 804, or receives the user sliding the smart interaction interface 803, and in response to the triggering operation of the user, the smart interaction interface 803 is exited and the main interface 801 is returned to, that is, the first interface 801 shown in (a) is displayed.
[0154] The above is described by taking the case that the gallery has completed the image recognition and recommended the candidate theme for the images in the gallery as an example. If the voice assistant determines that the gallery has not performed the image recognition processing in response to the user triggering the entering of the smart interaction interface, or the voice assistant does not recommend the candidate theme according to the image recognition result of the gallery after the gallery performs the image recognition processing on the images, the mobile phone can display the recommendation interface of the voice assistant.
[0155] It can be understood that when the gallery performs the image recognition processing on the multiple images in the gallery, the mobile phone can display at least one candidate theme recommended by the voice assistant according to the image recognition result of the gallery in the smart interaction interface, such as Figure 8 As shown in (b), three candidate themes are displayed in the smart interaction interface 803. However, when the voice assistant does not recommend the candidate theme according to the image recognition result of the gallery, or the voice assistant determines that the gallery has not performed the image recognition processing on the images, the mobile phone displays Figure 8 the recommendation interface 805 shown in (c).
[0156] The above image recognition processing can be clustering, semantic recognition and the like on the multiple images in the gallery, and the specific processing process is not described in detail here.
[0157] The theme recommendation method of the embodiment of the present application is described in detail below by taking the case that the user triggers the entering of the smart interaction interface through the creation interface in the gallery as an example.
[0158] In the embodiment of the present application, the case that the mobile phone receives the user triggering the entering of the smart interaction interface through the creation interface in the gallery before the gallery performs the image recognition processing on the images in the gallery is taken as an example.
[0159] In an embodiment, the mobile phone receives the first operation of the user, and in response to the first operation, displays the second interface of the voice assistant. The second interface includes a first control for triggering the starting of the intelligent image recognition. The voice assistant receives the triggering operation of the user on the first control, and in response to the triggering operation of the user, displays the third interface of the gallery. The third interface includes the image recognition progress of the gallery. After the voice assistant determines that the image recognition of the gallery is completed, the first interface of the voice assistant is displayed.
[0160] For example, refer to Figure 9In the case of the phone displaying the creation interface 901 of the gallery, the phone receives a triggering operation of the user on the "intelligent photo" control 902, and in response to the triggering operation (first operation) of the user, the voice assistant of the phone displays the intelligent interaction interface 903 (second interface) shown in (b). Figure 9 In the case of the phone displaying the creation interface 901 of the gallery, the phone receives a triggering operation of the user on the "intelligent photo" control 902, and in response to the triggering operation (first operation) of the user, the voice assistant of the phone displays the intelligent interaction interface 903 (second interface) shown in (b).
[0161] It can be understood that, since the gallery has not performed the image recognition processing on the images in the gallery before, the intelligent interaction interface 903 can display guiding information or guiding controls to guide the user to trigger the gallery to perform the image recognition processing. For example, the intelligent interaction interface 903 includes a "start intelligent image recognition" control 904 (first control) for triggering the gallery to perform the image recognition processing.
[0162] During the image recognition processing of the gallery, the phone can display an interface for the user to view the image recognition progress. For example, the voice assistant of the phone receives a triggering operation of the user on the "start intelligent image recognition" control 904, and in response to the triggering operation of the user, the phone jumps to the image recognition interface (third interface) of the gallery to facilitate the user to view the image recognition progress. For example, the phone displays the image recognition interface 905 of the gallery shown in (c). Figure 9 During the image recognition processing of the gallery, the image recognition interface 905 of the gallery can display the image recognition progress of the gallery, for example, Figure 9 The image recognition progress bar 906 is displayed in (c).
[0163] In some embodiments, during the image recognition processing of the gallery, the phone can exit the image recognition interface 905 of the gallery and display other interfaces, or the phone can start other applications. For example, the phone can start a video application to play a video, or can also start a news application to play news, and the like. Then, the phone receives a triggering operation of the user to enter the image recognition interface of the gallery, and in response to the triggering operation of the user, the phone enters the image recognition interface of the gallery again. For example, during the image recognition processing of the gallery, the gallery receives a triggering operation of the user to exit the image recognition interface 905 of the gallery, and in response to the triggering operation of the user, the phone exits the image recognition interface 905 of the gallery. The phone receives a triggering operation of the user to enter the image recognition interface of the gallery again, and in response to the triggering operation of the user, the phone displays the image recognition interface 905 of the gallery shown in (c) again. Figure 9 During the image recognition processing of the gallery, the image recognition interface 905 of the gallery can display the image recognition progress of the gallery, for example,
[0164] During the image recognition processing of the gallery, the voice assistant receives other triggering operations of the user, and in response to the triggering operation of the user, the voice assistant can perform corresponding tasks. When the voice assistant receives a user instruction to perform other tasks, for example, the user instructs the voice assistant to set an alarm, query information, and the like through a voice instruction, the voice assistant can normally perform the corresponding task.
[0165] For example, it is assumed that the phone displays the image recognition interface 905 of the gallery shown in (c). Figure 9During the image recognition process shown in (c), the voice assistant receives a user's voice instruction to set an alarm. In response to the user's trigger, the voice assistant sends an alarm-setting command to the clock application, enabling the clock application to set the alarm. While the voice assistant interacts with the alarm application, the image recognition process continues, allowing the voice assistant to perform other tasks without interrupting the image recognition process, thus improving the interaction performance between the voice assistant and the user.
[0166] In some embodiments, during the image recognition process of the image library, the image library may sequentially recognize the images in the image library according to the time identifier of the images in the image library, or the image library may randomly recognize the images in the image library, or the image library may recognize the images in the image library according to other specific orders, which are not limited here.
[0167] After the image search in the gallery is complete, the phone can switch from the gallery to the voice assistant, redisplaying the voice assistant's intelligent interactive interface. For example, it can display... Figure 9 The intelligent interactive interface 907 (first interface) shown in (e) displays three candidate topics generated by the voice assistant based on the image recognition results from the image library.
[0168] In another embodiment, the mobile phone receives a first operation from the user and, in response to the first operation, displays a second interface of the voice assistant. The voice assistant receives a trigger operation from the user on a first control and, in response to the user's trigger operation, displays a first prompt message on the second interface to indicate that the image library is performing image recognition. After the voice assistant determines that the image library image recognition is complete, it displays the first interface of the voice assistant.
[0169] In other words, the voice assistant receives the user's trigger operation on the "Start Smart Image Recognition" control 904. In response to the user's trigger operation, the phone can continue to display the voice assistant's interface, such as the smart interactive interface, and can display a prompt message (the first prompt message) to inform the user that the image library is processing image recognition. For example, displaying... Figure 9 The intelligent interactive interface 908 shown in (d) includes a prompt message 909 to inform the user that the image library is performing image recognition processing. Similarly, after the image recognition process is complete, the phone can display... Figure 9 The intelligent interactive interface 907 (first interface) is shown in (e). That is to say, during the image recognition process in the gallery, the mobile phone does not actively jump to the image recognition interface of the gallery, but instead displays prompts in the intelligent interactive interface to provide image recognition prompts.
[0170] It should be explained that the number of candidate topics displayed in the intelligent interactive interface is a preset value. For example, the preset value may be a maximum of three candidate topics. When the voice assistant determines that the number of candidate topics generated based on the image recognition results exceeds the preset number, the intelligent interactive interface may display the preset number of candidate topics based on the number of images corresponding to each candidate topic.
[0171] The above example illustrates the situation where the user's mobile phone receives a notification from the gallery before the gallery has processed the images in the gallery, and the gallery recommends candidate themes after the image recognition is triggered.
[0172] In some other embodiments, after the image recognition in the gallery is completed, if the voice assistant does not generate candidate topics based on the image recognition results, a fifth interface of the voice assistant is displayed. The fifth interface includes a fourth prompt message, which is used to prompt the voice assistant that no candidate topics have been generated based on the image recognition results.
[0173] This can be understood as follows: if the phone has not processed the images in the gallery before the user triggers entry into the smart interaction interface through the creation interface in the gallery, and no candidate themes are recommended after triggering image recognition, then the phone will display a prompt message in the smart interaction interface. For example, in Figure 9 After the image recognition process shown in (c) is completed, if the voice assistant does not recommend any candidate topics based on the image recognition results, the phone can display... Figure 9 The intelligent interactive interface 910 shown in (f) displays a prompt message 911 to indicate that no candidate topics have been recommended.
[0174] In some other embodiments, if the image library has completed image recognition processing and can recommend candidate themes based on the image recognition results before the mobile phone receives a trigger to enter the smart interactive interface, the mobile phone can display the candidate themes recommended by the voice assistant based on the image recognition results of the image library in the smart interactive interface. For example, if the mobile phone's image library receives a user's trigger operation on the "Smart Image" control 902, it directly displays the image in response to the user's trigger operation. Figure 9 The intelligent interactive interface 907 is shown in (e).
[0175] When the intelligent interactive interface displays at least one candidate topic recommended by the voice assistant based on the image recognition results of the image library, if the voice assistant receives a trigger operation from the user on any candidate topic (the candidate topic selected by the user can be called the target topic), in response to the user's trigger operation (the third operation), the voice assistant finds the image material corresponding to the target topic from the image library, and generates the corresponding video based on the image material corresponding to the target topic.
[0176] For example, refer to Figure 10In Figure (a), the mobile phone displays a smart interactive interface 1001 for a voice assistant, which includes three candidate topics. In some embodiments, the voice assistant receives a user's trigger operation on a candidate topic 1002 (target topic). In response to the user's trigger operation (third operation), the voice assistant can obtain image materials corresponding to the target topic and then display them. Figure 10 The intelligent interactive interface 1003 shown in (b) includes image materials 1004 corresponding to the target topic retrieved by the voice assistant based on the user's trigger operation on the target topic. The intelligent interactive interface 1003 also displays a "View Photos" control, the number of image materials (e.g., 12), and a "Generate Video" control.
[0177] It needs to be explained that, Figure 10 The image materials corresponding to the target topic shown in (b) are only an example. The materials found by the voice assistant for the target topic can be either image materials or video materials; this step is limited here. Furthermore, the number of image materials corresponding to the target topic is only an example. When the voice assistant finds too many image materials corresponding to the target topic, the intelligent interactive interface may only display a portion of the images, for example, Figure 10 The target theme shown in (b) has 12 images, but the intelligent interactive interface 1003 only displays 9 images.
[0178] After the voice assistant receives the user's trigger action by clicking the "View More" control, it can display more images in the intelligent interactive interface in response to the user's trigger action.
[0179] The voice assistant receives user input to the "Generate Video" control (fourth operation), or receives a user's voice command to generate a video (fourth operation). Based on the multiple images included in image material 1004, the voice assistant generates a video corresponding to the target theme. It can also display... Figure 10 The intelligent interactive interface 1005 is shown in (c). The intelligent interactive interface 1005 displays a thumbnail 1006 of the target video corresponding to the generated target theme. The voice assistant receives a user's trigger operation on the video thumbnail 1006 and, in response to the user's trigger operation, plays the content corresponding to the video. The video thumbnail can be one of the images in the image material or a predefined image; this embodiment of the application does not impose any limitations.
[0180] In other embodiments, the voice assistant receives a user instruction to generate another topic (a first topic different from the candidate topics). In response to the user's trigger operation (the fifth operation), the voice assistant does not acquire the image material corresponding to the first topic. (As shown...) Figure 10In the intelligent interactive interface 1007 shown in (d), the voice assistant receives a user's instruction to "create a video of Nanjing tourism." In response to the user's instruction, the voice assistant fails to find a corresponding image. The intelligent interactive interface 1007 displays a prompt message 1008 (fifth prompt message) to indicate that no image material for generating the video was found.
[0181] In some embodiments, when the voice assistant determines, based on image recognition results from the image library, that the keywords of the user-selected target topic correspond to multiple identical or similar objects (e.g., children, pets, etc.), the voice assistant can display multiple objects in the intelligent interactive interface for the user to choose from. After receiving the user's selection of the target object, the voice assistant responds to the user's selection by searching for image materials corresponding to the target object, and generates a video based on the retrieved image materials corresponding to the target object.
[0182] For example, let's take the example of a voice assistant identifying a child as the target subject (person subject) based on keywords. Figure 10 In step (a), during the display of the intelligent interactive interface 1001 of the voice assistant on the mobile phone, the voice assistant receives the user's trigger operation on the candidate topic 1002 (person topic). In response to the user's trigger operation (sixth operation), the voice assistant determines, based on the image recognition results of the image library, that there are multiple images of children (multiple candidate persons) in the image library. Based on this, it can display... Figure 11 The intelligent interactive interface 1101 (first interface) shown in (a) displays images 1102 of multiple children (multiple candidate figures) clustered by the voice assistant based on image recognition results from the image library. The voice assistant receives the user's selection operation on any child (e.g., child 1, hereinafter referred to as the target child or target figure), and in response to the user's trigger operation (seventh operation), displays... Figure 11 The intelligent interactive interface 1103 shown in (b) is as follows. The intelligent interactive interface 1103 displays a prompt message 1104 to inform the user that the voice assistant is searching for image materials corresponding to the target child.
[0183] After the voice assistant finds the image material corresponding to the target child, it displays... Figure 11 The intelligent interactive interface 1105 shown in (c) displays multiple images 1106 corresponding to the target child.
[0184] The voice assistant receives the user's trigger action on the "Generate Video" control (eighth operation), and in response, generates the target video (person video), such as... Figure 11 A thumbnail 1108 of the target video is displayed in the intelligent interactive interface 1107 shown in (d).
[0185] When the voice assistant determines that the gallery has not performed image recognition processing, the voice assistant receives a first operation of the user triggering entering the intelligent interaction interface, and in response to the first operation, displays a sixth interface of the voice assistant. The sixth interface includes at least one preset theme.
[0186] In some embodiments, when the gallery has not performed image recognition on the images in the gallery, the voice assistant receives an operation of the user indicating to generate a video corresponding to a second theme, and in response to the operation of the user, the voice assistant triggers the gallery to perform image recognition processing. After the gallery completes image recognition, the voice assistant determines that the gallery does not find the material corresponding to the second theme according to the image recognition result of the gallery. In this case, the voice assistant displays prompt information (eighth prompt information) in the intelligent interaction interface to prompt that the material corresponding to the target theme is not found, and more images can be taken. For example, during the process of displaying the intelligent interaction interface of the voice assistant on the mobile phone, the voice assistant receives an operation of the user indicating to generate a video of children's success, and in response to the operation of the user, the voice assistant finds multiple children, such as Figure 12 The interface 1201 displayed in (a) displays images of three children. The voice assistant receives an operation of the user selecting the first child, and in response to the triggering operation of the user, the voice assistant does not find the image of the child, and displays prompt information 1202 to prompt the user that the image of the child is not found, and more images can be taken.
[0187] It can be understood that after the voice assistant receives an operation of the user indicating to generate a target video, the gallery does not complete image recognition, or the gallery deletes the image of the first child after the image recognition is completed, so that the voice assistant does not find the image of the child. In another example, when the voice assistant determines that the gallery does not have image material of the object corresponding to the keyword of the target theme according to the image recognition result of the gallery, the voice assistant determines that the gallery has newly added images that have not been subjected to image recognition processing. In this case, the voice assistant can display prompt information (ninth prompt information) in the intelligent interaction interface to prompt the user to perform image recognition processing on the newly added images, and the user can trigger the gallery to perform image recognition processing on the newly added images. Therefore, after the gallery performs image recognition processing on the newly added images, the candidate theme that is more consistent with the user's expectation can be recommended in the intelligent interaction interface.
[0188] For example, as shown in Figure 12The prompt information 1204 displayed in the intelligent interaction interface 1203 shown in (b) is used to prompt the user to perform intelligent image recognition processing on the newly added image. The voice assistant receives the user's triggering operation of "starting intelligent image recognition", and in response to the user's triggering operation, the gallery can perform image recognition processing, and the intelligent interaction interface 1203 can display prompt information 1205 to prompt the user that the gallery is performing image recognition processing on the newly added image in the gallery, and the image recognition duration. After the voice assistant determines that the gallery completes the image recognition of the newly added image, the voice assistant displays Figure 12 The prompt information 1207 in the intelligent interaction interface 1206 shown in (c) is used to prompt that the gallery has completed the image recognition of the newly added image and has not recommended a candidate theme.
[0189] During the image recognition processing of the gallery, the phone provides a function of viewing the image recognition progress of the gallery, that is, during the image recognition processing of the gallery, the user can view the image recognition progress of the gallery.
[0190] In some embodiments, during the image recognition processing of the gallery, prompt information can be displayed in the intelligent interaction interface of the voice assistant to prompt the image recognition duration of the gallery, so that the user determines whether to wait for the end of the image recognition according to the image recognition duration displayed in the prompt information.
[0191] Optionally, when the voice assistant determines that the image recognition duration of the gallery is greater than a preset duration, the intelligent interaction interface can display prompt information to prompt the image recognition duration of the entire image recognition process. The preset duration is a pre-set duration, for example, the preset duration can be 10s, 15s, etc. When the voice assistant determines that the image recognition duration of the gallery for the image in the gallery is less than the preset duration, the intelligent interaction interface can not prompt the image recognition duration. For example, when the phone determines that the entire image recognition duration of the gallery is 8s, the prompt information that can be displayed in the intelligent interaction interface is "intelligent image recognition is being performed, please wait".
[0192] For example, as shown in the intelligent interaction interface 1203 in (b) of 12, the prompt information 1205 displays the image recognition duration.
[0193] The above embodiments are described by taking the gallery or the voice assistant prompting the image recognition progress for the user to view as an example. In other embodiments, during the image recognition processing of the gallery, the notification control can also be displayed through the control center interface of the phone to prompt the image recognition progress of the gallery. For example, as shown in (a), the control center interface of the phone displays a notification control 1301 for prompting the image recognition progress of the gallery. The phone receives the user's triggering operation of the notification control 1301, and in response to the user's triggering operation, the phone displays Figure 13 The prompt information 1204 displayed in the intelligent interaction interface 1203 shown in (b) is used to prompt the user to perform intelligent image recognition processing on the newly added image. The voice assistant receives the user's triggering operation of "starting intelligent image recognition", and in response to the user's triggering operation, the gallery can perform image recognition processing, and the intelligent interaction interface 1203 can display prompt information 1205 to prompt the user that the gallery is performing image recognition processing on the newly added image in the gallery, and the image recognition duration. After the voice assistant determines that the gallery completes the image recognition of the newly added image, the voice assistant displaysFigure 9 The gallery recognition interface 906 shown in (c) can also be displayed in the control center interface of the phone when the gallery stops the image recognition processing of the images in the gallery. For example, Figure 13 The notification control 1302 in the control center interface shown in (b) displays the progress of the image recognition of the gallery. The notification control 1302 displays the image recognition stop. The phone receives the triggering operation of the notification control 1302 by the user, and in response to the triggering operation of the user, the phone displays Figure 9 The gallery recognition interface 906 shown in (c).
[0194] When the gallery completes the image recognition of the images in the gallery, the control center interface can also display the image recognition completion prompt information. For example, Figure 13 The notification control 1305 in the control center interface shown in (c) displays the prompt information "Smart Recognition Completed". The phone receives the triggering operation of the notification control 1305 by the user, and in response to the triggering operation of the user, the phone displays Figure 13 The smart interaction interface 1306 shown in (d). The interface 1306 displays the message control 1307 "Recognition Completed". During the image recognition processing of the gallery, the phone receives the operation of the user to trigger the display of the smart interaction interface again, and in response to the operation of the user, the smart interaction interface of the voice assistant is displayed again. The smart interaction interface can display the prompt information that the gallery is in the image recognition.
[0195] In some embodiments, the third interface of the gallery can include an icon of the voice assistant. During the display of the third interface of the gallery by the phone, the phone receives the second operation of the user on the voice assistant icon in the third interface of the gallery, and in response to the second operation, the fourth interface of the voice assistant is displayed. The fourth interface includes the second prompt information to prompt the image recognition progress of the gallery.
[0196] For example, referring to Figure 14 In the process of displaying the gallery recognition interface 1401 (third interface) by the phone, the gallery receives the triggering operation of the voice assistant icon 1402 in the gallery recognition interface 1401 by the user, and in response to the triggering operation (second operation) of the user, the smart interaction interface 1403 (fourth interface) shown in (b) is switched to. Figure 14 The smart interaction interface 1403 (fourth interface) shown in (b). The smart interaction interface 1403 displays the prompt information 1404 (second prompt information) to prompt the user that the gallery is in the image recognition and the image recognition progress of the gallery.
[0197] It needs to be explained that in the process of displaying the voice assistant icon 1402 in any interface of the mobile phone, the mobile phone receives the triggering operation of the user on the voice assistant icon 1402, and in response to the triggering operation of the user, the smart interaction interface can be switched to. That is, no matter which interface of the mobile phone displays the voice assistant icon, the mobile phone receives the triggering operation of the user on the icon, and the dialogue with the voice assistant can be expanded or the dialogue with the voice assistant can be resumed.
[0198] The above-mentioned way of triggering the smart interaction interface of the voice assistant by triggering the icon of the voice assistant in the gallery recognition interface is only an example, and the user can also trigger the smart interaction interface of the voice assistant in other ways, which is not limited here.
[0199] In the embodiment of the application, in the case where the gallery does not perform image recognition processing on the images in the gallery, after the mobile phone receives the triggering of the user entering the smart interaction interface through the creation interface in the gallery, the smart interaction interface can display a pre-set theme. After the mobile phone receives the operation of the user indicating the generation of a target theme video, the gallery triggers the image recognition processing on the images in the gallery to determine whether to generate the video corresponding to the target theme according to the image recognition result. That is, in this case, the voice assistant receives the operation of generating the video, and then triggers the gallery to perform real-time image recognition processing.
[0200] For example, referring to Figure 15 (a), when the mobile phone displays the creation interface 1501 of the gallery, the mobile phone detects the triggering operation of the user on the smart shooting control, and displays the smart interaction interface 1502 shown in Figure 15 (b). The smart interaction interface 1502 displays three themes. Since the gallery has not performed image recognition processing on the images in the gallery, the themes displayed in the smart interaction interface 1502 are pre-set themes.
[0201] In the process that the mobile phone displays the smart interaction interface 1502 in Figure 15 (b), the mobile phone receives the operation of the user on the "generate today's Vlog video" control, or the mobile phone receives the operation of the user indicating "generate today's Vlog video" by voice, and the voice assistant determines that the time length of the gallery performing image recognition processing on the images in the gallery is less than the pre-set time length, and in response to the operation of the user, the smart interaction interface 1503 shown in Figure 15 (c) is displayed. Wherein, Figure 15 The smart interaction interface 1503 shown in
[0202] In one case, in the process that the mobile phone displays the smart interaction interface 1502 in Figure 15In the process of the intelligent interaction interface 1503 shown in (c), when the gallery performs image recognition processing on the images in the gallery, the gallery performs image recognition processing only on the images in the gallery that correspond to the semantics in the dialogue message 1504. For example, the gallery performs image recognition processing only on the images in the gallery that have a time identifier of today. For another example, the voice assistant receives the operation of the user indicating "generate a video of A city travel", and in response to the operation of the user, the voice assistant triggers the gallery to perform image recognition processing only on the images in the gallery of A city. Here, the gallery performs image recognition processing on the images in the gallery in a targeted manner instead of performing image recognition processing on all the images in the gallery, which is beneficial to improving the efficiency of obtaining the material corresponding to the theme and improving the efficiency of generating the video.
[0203] In another case, the mobile phone displays Figure 15 In the process of the intelligent interaction interface 1503 shown in (c), the gallery can perform image recognition processing on all the images in the gallery, so that in the subsequent process of generating the video, the gallery does not need to be triggered to perform image recognition.
[0204] When the gallery determines that there is no image corresponding to the theme in the gallery according to the image recognition result, the mobile phone displays Figure 15 the prompt information 1506 shown in (d) to prompt that no image corresponding to the theme is found.
[0205] When the gallery determines that there is at least one image corresponding to the theme according to the image recognition result, the mobile phone displays Figure 15 the control 1507 shown in (e), and the control 1507 displays a plurality of images corresponding to the target theme. Figure 15 The number of images displayed in the control 1507 in (e) is only an example. After the mobile phone detects the triggering operation of the user clicking the "view photo" control, the control 1507 can display more images.
[0206] In the embodiments of the present application, when the gallery determines that there are a plurality of images corresponding to the target theme according to the image recognition result, the mobile phone can perform similarity detection and / or aesthetic scoring on the plurality of images to filter out images whose similarity between any two images is less than a similarity threshold value and / or whose aesthetic score is greater than a score threshold value. In this way, the display effect of the video generated by the mobile phone according to the filtered images is better.
[0207] The voice assistant receives the triggering operation of the user on the "generate video" control or receives the triggering operation of the user indicating to generate the video by voice, and in response to the triggering operation of the user, the voice assistant generates a video corresponding to the theme according to the plurality of images corresponding to the theme in the control 1507, and displays Figure 15 the video 1508 shown in (f).
[0208] The voice assistant receives the triggering operation of the user on the "generate video" control or receives the triggering operation of the user indicating to generate the video by voice, and in response to the triggering operation of the user, the voice assistant generates a video corresponding to the theme according to the plurality of images corresponding to the theme in the control 1507, and displays Figure 16It can be known that the smart interaction interface of the mobile phone is displayed all the time during the whole process of generating the corresponding video in response to the trigger operation indicated by the user, and the mobile phone does not jump to other interfaces, so that the user can intuitively see the whole video generation process.
[0209] In another scenario, the voice assistant receives a trigger operation of the user on the "generate today's Vlog video" control in the smart interaction interface 1502, or receives an operation of the user indicating "generate today's Vlog video" by voice, and the voice assistant determines that the time length of the gallery for image recognition processing in the gallery is greater than the preset time length. The mobile phone can display prompt information in the smart interaction interface to prompt the user about the image recognition time length of the gallery. For example, the mobile phone displays the prompt information 1601 shown in (a) to prompt the image recognition time length of the whole image recognition process of the gallery. Figure 16
[0210] It can be understood that when the time length of the gallery for image recognition processing in the gallery is greater than the preset time length, the smart interaction interface of the mobile phone displays prompt information prompting the image recognition time length, and the user can determine whether to wait for the image recognition result in the interface according to the prompt information.
[0211] In the embodiment of the present application, after the voice assistant receives the operation of the user indicating to generate the video of the target theme, the voice assistant determines that the gallery does not perform image recognition processing on the images in the gallery, and then, during the image recognition processing of the gallery, the image recognition may be abnormal. The voice assistant can display prompt information to prompt the abnormal reason. For example, Figure 16 The smart interaction interface shown in (a) displays prompt information 1601 to prompt that the gallery is performing image recognition processing. During the display of (a), the mobile phone power is too low, causing the image recognition process to be interrupted, and the prompt information 1602 shown in (b) is displayed to prompt that the image recognition of the gallery is abnormal. Figure 16 Figure 16
[0212] For another example, Figure 16 Because the gallery is cloning pictures, the image recognition process is interrupted, and the smart interaction interface displays prompt information 1603. For another example, Figure 16 Because the image recognition process of the gallery is abnormal, the image recognition process is interrupted, and the smart interaction interface displays prompt information 1604.
[0213] It should be noted that the above Figure 17 The reasons for image recognition anomalies shown in (b) to (d) are merely examples. Image recognition may be interrupted for other reasons during the process. For instance, if the phone detects overheating during image recognition, the process may be interrupted. In this case, the smart interface can display the message "Image recognition incomplete. Device temperature too high. Please cool the device. It will automatically continue after the temperature drops." Another example is if the phone detects that the user has terminated the process, causing the image recognition to be interrupted. In this case, when the phone displays the smart interface again, it can display the message "Image recognition incomplete. Image recognition was interrupted. Please continue."
[0214] In this embodiment, after a user wakes up the voice assistant via voice or by long-pressing the power button, they can also interact with the voice assistant via voice to generate a video corresponding to the target topic within the voice assistant's dialogue interface. For example, when the phone is displaying any interface, after receiving the user's voice wake-up operation, the phone responds to the user's voice wake-up operation by displaying... Figure 17 The voice assistant's dialogue interface 1701 (i.e., the intelligent interaction interface) is shown in (a). While the phone displays the voice assistant's dialogue interface 1701, it receives the user's voice instruction to "generate a video of a child dancing." In response to the user's voice instruction, the voice assistant determines that the image library has not performed image recognition. In this case, the dialogue interface 1701 displays a "Start Smart Image Recognition" control. After receiving the user's voice instruction to start smart image recognition, the voice assistant, in response to the user's trigger operation, displays... Figure 17 Image recognition interface 1702 shown in (b) of the image library.
[0215] In another scenario, during the image recognition process in the gallery, instead of switching from the voice assistant's dialogue interface to the gallery's image recognition interface, a prompt message is displayed within the dialogue interface, such as... Figure 17 The dialog interface 1701 shown in (d) displays a prompt message 1703 to inform the user that the image library is recognizing images.
[0216] After the image recognition function in the photo library is completed, the phone displays... Figure 17 The prompt message 1704 shown in (c) prompts the user to select the child to be used in the video from among several children recommended based on the image recognition results. For example, if the voice assistant receives the user's selection of the first child, and in response to the user's selection, the voice assistant determines, based on the image recognition results of the image library, that there are not enough images of that child in the library to generate a video, it will display... Figure 17The prompt message 1705 shown in (e) indicates to the user that there are not enough usable images in the image library. For example, the voice assistant receives the user's selection of the first child image. In response to the user's selection, the voice assistant determines the image of that child in the image library based on the image recognition results, and the phone displays... Figures 3 to 17 The prompt message 1706 shown in (f) indicates to the user that the voice assistant is searching for the image corresponding to the child.
[0217] In another scenario, after the image library completes its image recognition process and recommends a child based on the results, the dialogue interface displays the child's image. The voice assistant receives the user's selection of the child and, in response, directly recommends the corresponding image.
[0218] In this embodiment, after the image corresponding to the target theme is displayed in the intelligent interactive interface of the voice assistant, the voice assistant receives a user's voice instruction to generate a video, or receives a user's trigger operation on the "Generate Video" control. In response to the user's operation, the voice assistant generates a corresponding video based on the image corresponding to the target theme and displays the generated video cover in the intelligent interactive interface. Thus, users can quickly generate videos on their desired themes through the intelligent interactive interface of the voice assistant.
[0219] It should be noted that the above The interface and content shown are merely examples and are not limited in this application. For instance, the icon for the voice assistant can be any icon different from those of existing applications; the image is only an example and is not limited here.
[0220] It is understood that the aforementioned electronic devices, etc., include hardware structures and / or software modules corresponding to the execution of each function in order to achieve the above-mentioned functions. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this invention.
[0221] The embodiments of the present application can divide the functional modules of the electronic device and the like according to the method examples described above. For example, each functional module can be divided according to each function, or two or more functions can be integrated in one processing module. The integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of the modules in the embodiments of the present application is illustrative, and is only a logical function division. In actual implementation, another division mode can be used.
[0222] In the case of dividing each functional module according to each function, a possible composition schematic diagram of the electronic device involved in the embodiments described above can include a display unit, a transmission unit, a processing unit and the like. It should be noted that all related contents of each step involved in the method embodiments can be referred to the function description of the corresponding functional module, and will not be described here.
[0223] The embodiments of the present application also provide an electronic device including one or more processors and one or more memories. The one or more memories are coupled to the one or more processors, and the one or more memories are configured to store computer program codes including computer instructions. When the one or more processors execute the computer instructions, the electronic device is caused to perform the steps of the related methods described above to implement the subject recommendation method in the embodiments described above.
[0224] The embodiments of the present application also provide a computer readable storage medium having computer instructions stored therein. When the computer instructions are run on an electronic device, the electronic device is caused to perform the steps of the related methods described above to implement the subject recommendation method in the embodiments described above.
[0225] The embodiments of the present application also provide a computer program product including computer instructions. When the computer instructions are run on an electronic device, the electronic device is caused to perform the steps of the related methods described above to implement the subject recommendation method in the embodiments described above.
[0226] In addition, the embodiments of the present application also provide a device, which can be a chip, a component or a module. The device can include a processor and a memory connected to each other. The memory is configured to store computer execution instructions. When the device is running, the processor can execute the computer execution instructions stored in the memory to cause the device to perform the subject recommendation method performed by the electronic device in the method embodiments described above.
[0227] The electronic device, the computer readable storage medium, the computer program product or the device provided by the embodiments of the present application are used to execute the corresponding methods provided above, and thus the beneficial effects achieved thereby can refer to the beneficial effects of the corresponding methods provided above, which will not be described here.
[0228] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of functional modules is exemplified, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0229] The functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0230] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: a flash memory, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk or an optical disk, and various media that can store program codes.
[0231] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any change or replacement within the technical scope disclosed in the present application should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A topic recommendation method, characterized in that, Applied to electronic devices including voice assistants, the method includes: Receive the user's first action; In response to the first operation, a first interface of the voice assistant is displayed, the first interface including at least one candidate topic; the candidate topic is generated by the voice assistant based on the image recognition results of the image in the image library of the electronic device; A video corresponding to the candidate topic is generated based on at least one image corresponding to the candidate topic, wherein the at least one image is an image from the image library of the electronic device.
2. The method according to claim 1, characterized in that, The first interface of the voice assistant, displayed in response to the first operation, includes: In response to the first operation, a second interface of the voice assistant is displayed. The second interface includes a first control, which is used to trigger the gallery to start image recognition processing. In response to the user's operation on the first control, a first prompt message is displayed on the second interface, the first prompt message being used to indicate that the image library is performing image recognition; After the image recognition in the image library is completed, the first interface of the voice assistant is displayed.
3. The method according to claim 1, characterized in that, The first interface of the voice assistant, displayed in response to the first operation, includes: In response to the first operation, a second interface of the voice assistant is displayed. The second interface includes a first control, which is used to trigger the gallery to start image recognition processing. In response to the user's operation on the first control, a third interface of the image library is displayed, the third interface including the image recognition progress of the image library; After the image recognition in the image library is completed, the first interface of the voice assistant is displayed.
4. The method according to claim 3, characterized in that, The third interface includes the icon of the voice assistant, and the method further includes the following steps during the process of displaying the third interface of the gallery: Receive the user's second operation on the icon of the voice assistant; In response to the second operation, a fourth interface of the voice assistant is displayed, the fourth interface including a second prompt message, the second prompt message being used to indicate the image recognition progress of the image library.
5. The method according to claim 3 or 4, characterized in that, While the image library is performing image recognition, the second interface displays a third prompt message, which is used to indicate the image recognition time of the image library.
6. The method according to claim 2 or 3, characterized in that, After the image recognition in the image library is completed, the first interface of the voice assistant is displayed, including: After the image recognition in the image library is completed, if the voice assistant generates at least one candidate topic based on the image recognition results in the image library, then a first interface including at least one candidate topic is displayed.
7. The method according to claim 2 or 3, characterized in that, The method further includes: After the image recognition in the image library is completed, if the voice assistant does not generate the candidate topic based on the image recognition results, the fifth interface of the voice assistant is displayed. The fifth interface includes a fourth prompt message, which is used to prompt the voice assistant that no candidate topic has been generated based on the image recognition results.
8. The method according to any one of claims 1-4, characterized in that, The method further includes: Receive a third operation from the user on the target topic among the at least one candidate topics; In response to the third operation, at least one image corresponding to the target topic is displayed on the first interface; Receive the fourth operation triggered by the user to generate a video; In response to the fourth operation, a thumbnail of the target video is displayed on the first interface, the target video being generated based on at least one image corresponding to the target theme.
9. The method according to any one of claims 1-4, characterized in that, The method further includes: Receive a fifth operation from the user instructing the generation of a video with a first theme, which is different from the candidate themes; In response to the fifth operation, a fifth prompt message is displayed on the first interface, which indicates that the image library does not include an image corresponding to the first theme.
10. The method according to claim 9, characterized in that, The method further includes: In response to the fifth operation, a sixth prompt message is displayed on the first interface, which prompts the user to generate a video corresponding to a theme other than the first theme.
11. The method according to any one of claims 1-4, characterized in that, The candidate topics include people-related topics, and the method further includes: Receive the user's sixth operation on the topic of people in the at least one candidate topic; In response to the sixth operation, multiple candidate objects are displayed on the first interface; Receive the user's seventh operation on the target person among the multiple candidate persons; In response to the seventh operation, at least one portrait of the target person is displayed on the first interface, the at least one portrait being an image from the gallery of the electronic device; The eighth operation, triggered by the user, is to generate a character video. In response to the eighth operation, a thumbnail of the character video is displayed on the first interface, the character video being generated based on at least one portrait corresponding to the target character.
12. The method according to any one of claims 1-4, characterized in that, After receiving the user's first operation, the method further includes: When the voice assistant determines that the image library has not undergone image recognition processing, in response to the first operation, the sixth interface of the voice assistant is displayed, which includes at least one preset theme.
13. The method according to claim 12, characterized in that, The method further includes: Receive the user's ninth operation on the second topic among the at least one preset topics; In response to the ninth operation, a seventh prompt message is displayed on the sixth interface, which is used to indicate that the image library is performing image recognition. After the image search in the image library is completed, a seventh interface is displayed; the seventh interface includes at least one image found by the voice assistant based on the image search results from the image library.
14. The method according to claim 13, characterized in that, The method further includes: After the image search in the image library is completed, an eighth prompt message is displayed on the sixth interface. The eighth prompt message is used to indicate that no image corresponding to the second theme was found. The seventh prompt message is not displayed on the sixth interface.
15. The method according to any one of claims 2-4, characterized in that, The method further includes: If an error occurs during the image recognition process, an error message will be displayed on the sixth interface of the electronic device to indicate that there is an error in the image recognition process.
16. The method according to any one of claims 1-4, characterized in that, The method further includes: When the voice assistant determines that there is a newly added image in the gallery that has not undergone image recognition processing, it displays a ninth prompt message on the first interface. The ninth prompt message is used by the user to perform image recognition processing on the newly added image in the gallery that has not undergone image recognition processing.
17. The method according to any one of claims 1-4, characterized in that, The first operation is that the user triggers the icon of the voice assistant to trigger the operation of entering the smart video creation function, or the first operation is that the user triggers the desktop card of the electronic device to trigger the operation of entering the smart video creation function, the desktop card includes an entry point for triggering the entry point for entering the smart video creation function, or the first operation is that the user triggers the entry point of the first interface provided by the gallery to trigger the operation of entering the smart video creation function.
18. An electronic device, characterized in that, include: One or more processors; Memory; The memory stores one or more computer programs, the one or more computer programs including instructions that, when executed by the electronic device, cause the electronic device to perform the topic recommendation method as described in any one of claims 1-17.
19. A computer-readable storage medium storing instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the topic recommendation method as described in any one of claims 1-17.
Citation Information
Patent Citations
Electronic photo album acquisition method and device, computer equipment and storage medium
CN111010611A
Theme video generation method and device, electronic equipment and readable storage medium
CN111669620A
Voice interaction method and electronic equipment
CN111724775A