An interaction method and electronic device
Patent Information
- Application Number
- CN202511009825.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-07-21
AI Technical Summary
[0020]第四方面,本申请提供一种计算机可读存储介质,包括计算机程序,当上述计算机程序在电子设备上运行时,使得上述电子设备执行如第一方面以及第一方面中任一可能的实现方式描述的方法。
Smart Images

Figure CN121614100B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminals, and more particularly to an interaction method and an electronic device. Background Technology
[0002] To avoid missing the perfect moment, users often use burst mode or manually click the shutter button multiple times in succession to take several photos with similar content. However, not all of these photos are necessarily what the user needs. Therefore, after shooting, users often review their photos and delete redundant ones to avoid wasting memory. However, deleting them one by one is time-consuming and inconvenient. Summary of the Invention
[0003] This application provides an interaction method and an electronic device.
[0004] Firstly, this application provides an interaction method. This method is applied to an electronic device. The electronic device includes, but is not limited to, mobile phones, tablet computers, and personal computers. The method includes: receiving a first voice command from a user; when the first voice command includes a first keyword, in response to the first voice command, displaying a first interface of a gallery, the first interface including a first option; in response to a user operation on the first option, displaying a second interface, the second interface displaying multiple sets of images, each set of images including multiple similar or identical images; when the first voice command includes a first keyword and a second keyword, the second keyword corresponds to the first option, and in response to the first voice command, displaying the second interface; in response to a user operation on a first control in the second interface, retaining one image from each set of images in the second interface.
[0005] The first keyword includes target action and target content. Target action refers to actions related to cleaning, such as cleaning, tidying, managing, and sweeping. Target content refers to images (including pictures (photos) and videos) or image-related content, such as a photo library. Photos include still photos composed of single frames and animated photos composed of multiple frames. The first interface is, for example... Figure 3A The image management interface shown has the following first option: Figure 3A The second interface shows "Duplicate Images", "Similar Images", and "Duplicate Videos". Figure 3B The diagram shows the duplicate image interface corresponding to "Duplicate Images," and the similar image interface corresponding to "Similar Images" (not shown in the attached diagram), and the duplicate video interface corresponding to "Duplicate Videos." The second keyword is the target type, for example... Figure 3A The terms "repeated" and "similar" are shown. The first control is, for example... Figure 3B The control shown is 302.
[0006] By implementing the method provided in the first aspect, users can trigger different types of image cleaning anytime and anywhere through different voice commands, and clean up a large number of redundant images in electronic devices with minimal interactive operations, thereby improving the user's cleaning efficiency and improving the effective utilization of electronic device memory.
[0007] In some embodiments, the first voice command further includes a third keyword, and the multiple images in the second interface are images that match the third keyword. Optionally, the third keyword includes one or more of the following: the subject being photographed, the time of shooting, the location of shooting, and the color.
[0008] This allows users to more effectively target and remove duplicate or similar images from their electronic devices, thereby improving their cleaning efficiency.
[0009] In some embodiments, the method further includes: when the first voice command includes a first keyword and a fourth keyword, wherein the fourth keyword corresponds to a second option in the first interface, and in response to the first voice command, displaying a third interface containing multiple blurred images, AI-enhanced images, or large videos; and in response to a user operation on a first control in the third interface, deleting an image from the third interface. The fourth keyword is a target type, for example... Figure 3A The terms "blurred", "AI enhanced", and "extra large" are shown.
[0010] This means that users can also directly obtain blurry, AI-enhanced, or ultra-large images from electronic devices, clean them up, improve cleaning efficiency, and increase the effective utilization of electronic device memory.
[0011] In some embodiments, the method further includes: when the first voice command does not include the first keyword, displaying a prompt message on the dialogue interface, the prompt message being used to instruct the user to re-enter the voice command.
[0012] In this way, users can re-enter voice commands that the electronic device can recognize based on the above prompts to clean up images.
[0013] In some embodiments, displaying a second interface in response to a first voice command includes: first displaying the first interface in response to the first voice command, and then automatically jumping to the second interface.
[0014] In this way, while quickly jumping to the second interface, users can also see the parent interface of the second interface, so that they can return to the parent interface to perform other cleanup operations.
[0015] In some embodiments, the image retained in each group of images is the image with the highest score in that group of images; wherein, the score of an image includes the image's score in N aspects, and the N aspects include one or more of the following: image quality sharpness, subject, object state, and composition.
[0016] In some embodiments, the highest-rated image in the set of images includes: the highest-rated image in the set of images carrying a first tag, the first tag indicating that the user liked the image and / or that the user edited the image.
[0017] In some embodiments, one image retained in each group of images is: an image obtained by fusing one or more images carrying a first tag in the group of images, the first tag indicating that the user likes the image and / or the user has edited the image.
[0018] In a second aspect, this application provides an electronic device including one or more processors and one or more memories; wherein the one or more memories are coupled to one or more processors, and the one or more memories are used to store a computer program, which, when executed by one or more processors, causes the electronic device to perform the method described in the first aspect and any possible implementation thereof.
[0019] Thirdly, embodiments of this application provide a chip system applied to an electronic device. The chip system includes one or more processors, which are used to invoke computer instructions to cause the electronic device to perform the methods described in the first aspect and any possible implementation thereof.
[0020] Fourthly, this application provides a computer-readable storage medium including a computer program that, when run on an electronic device, causes the electronic device to perform the method described in the first aspect and any possible implementation thereof.
[0021] Fifthly, this application provides a computer program product containing instructions that, when the computer program product is run on an electronic device, cause the electronic device to perform the method described in the first aspect and any possible implementation thereof.
[0022] Understandably, the electronic device provided in the second aspect, the chip system provided in the third aspect, the computer storage medium provided in the fourth aspect, and the computer program product provided in the fifth aspect are all used to execute the method provided in this application. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description
[0023] Figure 1 This is a flowchart of an interactive method for image management provided in an embodiment of this application;
[0024] Figure 2 This is a schematic diagram of a dialog interface provided in an embodiment of this application;
[0025] Figures 3A-3C These are schematic diagrams of a set of cleanup interfaces provided in the embodiments of this application;
[0026] Figure 4 This is a flowchart of another interactive method for image management provided in an embodiment of this application;
[0027] Figure 5 This is a schematic diagram of the user interface for displaying prompt information provided in an embodiment of this application;
[0028] Figure 6 This is a schematic diagram of the software architecture provided in the embodiments of this application;
[0029] Figure 7 This is another software architecture diagram provided in the embodiments of this application;
[0030] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0031] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be a limitation of this application.
[0032] When taking photos, to avoid missing the best moment, users often use burst mode or manually click the shutter button multiple times in succession, resulting in several photos with similar content. However, not all of these photos are what the user needs. Therefore, after shooting, users often review the photos they have taken and delete redundant ones to avoid consuming memory. However, deleting them one by one is time-consuming and laborious, causing inconvenience for users.
[0033] Therefore, embodiments of this application provide an interactive method for image management. This method can be applied to electronic devices such as mobile phones, tablets, and personal computers.
[0034] Figure 1 This is a flowchart of an interactive method for image management provided in an embodiment of this application.
[0035] S101, Received the user's voice command.
[0036] Electronic devices may be equipped with a microphone. This microphone can continuously collect sound signals from the surrounding environment. When a user speaks a voice command, the microphone of the electronic device can collect the sound signal containing the voice command, and then recognize the voice command.
[0037] S102. Display the dialogue interface and show the user's voice commands in the dialogue interface.
[0038] Optionally, upon receiving a user's voice command, the electronic device may display a dialog interface showing the user's voice command. The electronic device may switch from any previously displayed interface (such as a desktop, application interface, etc.) to the dialog interface to instruct the user to: recognize the user's voice command, respond to the user, and execute the operation instructed by the user's voice command.
[0039] Figure 2 This is a schematic diagram of a dialog interface provided in an embodiment of this application. For example... Figure 2 As shown, electronic devices (such as mobile phones) can display recognized user voice commands in the dialogue interface, such as: "Hi, YOYO, help me clean up the pictures" (first voice command). Here, "Hi, YOYO" is the wake-up word used to wake up the voice assistant YOYO and display YOYO's dialogue interface; "help me clean up the pictures" is the specific operation command.
[0040] S103. When the voice command includes the target action and the target content, the image management interface is displayed.
[0041] The target action is an action related to cleaning, such as cleaning, organizing, managing, and sweeping. The target content is an image (including pictures (pictures include photos) and video) or image-related content, such as a photo library. After receiving a user's voice command, the electronic device can parse the voice command to determine whether it includes the preset target action and target content. When the voice command includes the target action and target content (first keyword), the electronic device can display the image management interface (first interface), showing the images to be cleaned in the current device.
[0042] For example, from a user's voice command, "Help me clean up pictures," the electronic device can recognize the target action: "clean up" and the target content: "pictures." Then, in response to the voice command, the electronic device can display an image management interface.
[0043] Figure 3A This is a schematic diagram of an image management interface provided in an embodiment of this application. Figure 3A As shown, the image management interface may include an image module and a video module.
[0044] The image module includes several types of images to be cleaned, such as "Duplicate Images," "Similar Images," "Blurred Images," and "Artificial Intelligence (AI) Enhanced Images." Multiple images with completely identical content (i.e., two or more) constitute a group of duplicate images. "Duplicate Images" includes one or more groups of duplicate images. Multiple images with a content similarity exceeding a preset threshold (e.g., 0.9) constitute a group of similar images. "Similar Images" includes one or more groups of similar images. "Blurred Images" includes one or more images containing blur. This blurring includes, but is not limited to, motion blur, defocus blur, atmospheric turbulence blur, and compression blur. "AI Enhanced Images" includes one or more images edited using AI enhancement techniques. These images include static images composed of single frames and dynamic images composed of multiple frames.
[0045] The video module may also include multiple video types to be cleaned, such as "duplicate videos" and "extra-large videos." Multiple videos with completely identical content constitute a group of duplicate videos. Videos whose individual memory exceeds a preset threshold (e.g., 200MB) can be called extra-large videos. In some embodiments, the video module may also include "similar videos." Multiple videos with similar images, subjects, and audio content constitute a group of similar videos.
[0046] S104. After detecting any type of user operation in the image management interface, display the sub-interface corresponding to that type, show the image to be cleaned corresponding to that type in the sub-interface, and clean some or all of the images to be cleaned in that type according to the user operation.
[0047] refer to Figure 3A The electronic device can detect user interaction with the "repeating image" (first option), and in response to the aforementioned user interaction, the electronic device can display... Figure 3B The sub-interface that displays the repeating image shown is denoted as the repeating image interface (second interface).
[0048] like Figure 3BAs shown, the repeating image interface displays multiple sets of repeating images, such as G1 (including images 11-14), G2 (including images 21-25), etc. Each set of repeating images has a status control, such as control 303 and control 304. By default, when entering the repeating image interface, the status control corresponding to each set of repeating images is set to the selected state. Users can operate one or more status controls in the repeating image interface according to their needs to set the above status controls to the unselected state. The repeating image interface also includes controls 301 and 302. By default, when entering the repeating image interface, control 301 is set to the open state. After detecting user operation on control 302 (the first control), the electronic device can perform a cleanup operation to clean up the repeating images whose status controls are set to the selected state. Optionally, refer to Figure 3C Upon detecting a user action on control 302, the electronic device may first display pop-up window 306. After detecting a user action to delete a control in pop-up window 306, the electronic device then performs a cleanup operation to minimize accidental deletion.
[0049] When control 301 is enabled, for each set of duplicate images, the electronic device retains only one image from that set and deletes the others. The user can manipulate control 301 to disable it. When control 301 is disabled, the electronic device can delete all images in that set after detecting user interaction with control 302.
[0050] Similarly, upon detecting other types of user operations on the image management interface, in response to the aforementioned user operations, the electronic device may display the corresponding sub-interface and clear some or all of the images in the aforementioned sub-interface according to the user operations.
[0051] In some embodiments, when cleaning up similar images, the electronic device can obtain the ratings of each image in each group of similar images, determine the image with the highest rating, retain that image, and delete the other images in that group. The electronic device can determine the rating of each image based on factors such as image quality, subject, object state, and composition. A higher rating indicates that the user likes the image more. In other embodiments, when cleaning up similar images, the electronic device can also obtain other states (first tags) of each image in each group of similar images, such as whether it has been marked as liked by the user or whether it has been edited by the user. The electronic device can determine which image to retain in each group based on these other states. For example, the electronic device can prioritize identifying images marked as liked and / or edited images in each group of similar images. When a group of similar images includes multiple images marked as liked and / or edited images, the electronic device can further determine the image with the highest rating, retain that image, and delete the other images. Conversely, when a group of similar images has no images marked as liked or edited images, the electronic device can directly determine which image to retain based on the ratings of the images in that group.
[0052] In some embodiments, the electronic device may also identify images marked as favorites by the user and / or images that have been edited by the user in each group of images, and perform fusion processing on these images to obtain the corresponding fused image. During the cleanup operation, the electronic device may delete each image in each group of images, retaining only the corresponding fused image.
[0053] In some embodiments, after recognizing a voice command including a target action and target content, the electronic device can scan local images to identify duplicate images, similar images, blurred images, AI-enhanced images, duplicate videos, and oversized videos within the local images. In some embodiments, the electronic device can also periodically scan local images in the background to identify duplicate images, similar images, blurred images, AI-enhanced images, duplicate videos, and oversized videos within the local images. Thus, upon detecting any type of user operation on the image management interface, the electronic device can immediately display the corresponding sub-interface, showing the image to be cleaned for that type, without waiting for real-time scan results.
[0054] Figure 4 This is a flowchart of another interactive method for image management provided in an embodiment of this application.
[0055] S201, Received user's voice command.
[0056] S202. Display the dialogue interface and show the user's voice commands in the dialogue interface.
[0057] S203. When the voice command includes the target action, target type and target content, the sub-interface corresponding to the target type is directly displayed. The sub-interface displays the image to be cleaned corresponding to the type, and cleans part or all of the image to be cleaned according to the user's operation.
[0058] For example, an electronic device can receive a user's voice command: "Hi, YOYO, help me clean up duplicate images" (first voice command). In response to the voice command, the electronic device can display a dialog interface showing the voice command. In this embodiment, from the voice command, the electronic device can identify the target action: "clean up," the target content: "image" (first keyword), and the target type: "duplicate" (second keyword). Therefore, in response to the voice command, the electronic device can directly display... Figure 3B The duplicate image interface shown is the second interface. After detecting a user operation to delete a control (first control) in control 302 and / or pop-up window 306, the electronic device performs a cleanup operation to clear duplicate images whose status controls are set to the selected state. The specific process of the electronic device performing the cleanup operation can be found in the description of S104, and will not be repeated here.
[0059] The aforementioned direct display includes both direct jump displays without inserting other content and continuous jump displays that insert other content but require no user intervention. For example, a direct jump display without inserting other content might be shown in the context of... Figure 2 In the dialog interface shown, the electronic device can immediately display... Figure 3B The interface showing repeated images does not display the middle section. Figure 3A The image management interface shown above. The above-mentioned continuous jump display allows for the insertion of other content without user intervention, for example: in the display... Figure 2 In the dialog interface shown, the electronic device can first display Figure 3A The image management interface shown (first interface) will then be displayed automatically. Figure 3B The repeated image interface shown (second interface).
[0060] Similarly, when the user's voice command includes the type "similar" (second keyword), the electronic device can directly display a sub-interface (second interface) showing similar images, saving user interaction; when the user's voice command includes the type "blurred" (fourth keyword), the electronic device can directly display a sub-interface (third interface) showing blurred images, saving user interaction. These examples will not be listed here.
[0061] In some embodiments, after recognizing the target action "clean", the target content "image", and the target type "repeat", the electronic device can also directly display the above voice command. Figure 3BThe screen displays a duplicate image, and pop-up 306 is shown directly. Upon detecting a user action to delete a control in pop-up 306, the electronic device performs a cleanup operation, clearing duplicate images whose status controls are set to selected. This saves users from interactive operations and improves cleanup speed.
[0062] exist Figure 1 and Figure 4 In the method shown, the electronic device can also recognize qualifying words (third keywords) in the user's voice commands. These qualifying words include, but are not limited to, qualifying words for the subject of the photograph (e.g., "flower," "sea," "mountain," "cat," "dog," "food," etc.), qualifying words for the time of photograph (e.g., "year," "month," "date," "last month," "last day"), qualifying words for the location of the photograph (e.g., "city A," "district B," "scenic area C"), and qualifying words for the color (e.g., "white," "blue," "green," etc.). When scanning local images, the electronic device can further identify images that match the aforementioned image content qualifying words. Furthermore, when displaying any type of sub-interface, the electronic device can display only the images that match the aforementioned image content qualifying words in that sub-interface.
[0063] For example, the electronic device can receive a user's voice command: "Hi, YOYO, help me clean up similar images related to seascapes." In this embodiment, the electronic device can also recognize the qualifier "sea" from the aforementioned voice command. In response to the voice command, the electronic device can scan local images to identify similar images related to "sea." When displaying a similar image sub-interface, the electronic device only displays similar images related to "sea" in that sub-interface. Therefore, when performing the cleaning operation, the electronic device can only clean up similar images related to "sea," without cleaning up other similar images (e.g., similar images related to "flowers," similar images related to "cats," etc.).
[0064] For example, an electronic device can receive a user's voice command: "Hi, YOYO, help me clean up similar images from the last 6 months." In this embodiment, the electronic device can also recognize the qualifier "last 6 months" from the aforementioned voice command. In response to the voice command, the electronic device can scan local images to identify images similar to those captured within the last 6 months. When displaying the similar images sub-interface, the electronic device only displays similar images captured within the last 6 months in that sub-interface. Therefore, when performing the cleaning operation, the electronic device can only clean up similar images from the last 6 months, without cleaning up other similar images older than 6 months.
[0065] In the above method, when the "target action + target content" is not recognized, the electronic device can display a prompt message on the dialogue interface, instructing the user to re-enter the voice command.
[0066] Figure 5 This is a schematic diagram of the user interface for displaying prompt information provided in an embodiment of this application. For example... Figure 5 As shown, after receiving the user's voice command, "Hi, YOYO, help me clean up storage space," the electronic device can recognize the target action but cannot recognize the target content. At this point, the electronic device can display a prompt: "What exactly do you want to clean up? You can try saying 'Help me clean up my photo library!'" The user can then re-enter a voice command that the electronic device can recognize based on the prompt.
[0067] Figure 6 This is a schematic diagram of the software architecture provided in the embodiments of this application.
[0068] like Figure 6 As shown, electronic devices can have system applications installed: voice assistants (such as YOYO). The voice assistant can receive user voice commands. It can call an audio processing engine to process these commands, understand their intent, and then obtain operation instructions. For example, upon receiving the user's voice command, "Help me clean up the pictures," based on its understanding of the command's intent, the voice assistant can obtain the operation instruction: "Open the gallery cleaning tool and clean up the pictures." Wherein, such as Figure 6 As shown, preferably, the intent understanding model can be set up in the cloud. The audio computing engine can call the intent understanding model over the network to perform intent understanding, thereby avoiding occupying the storage space of electronic devices.
[0069] The voice assistant can send the operation commands obtained from understanding the intent to the agent management service. The agent management service can be viewed as a system-level application. It includes a gallery agent and an Advanced Gateway Interface (AGI) agent framework. The gallery agent includes a task engine and gallery cleanup tools, while the AGI agent framework includes a task orchestration module and a response generation module.
[0070] In response to the operation commands sent by the voice assistant, the agent management service can activate the corresponding agent. For example, in response to the operation command: "Activate the gallery cleanup tool," the agent management service can activate the gallery agent. The gallery agent's task engine can call the task orchestration module of the AGI agent framework to define tasks and task parameters (qualifiers, type, content). For example, for the operation command: "Activate the gallery cleanup tool and clean up pictures," the task engine can define task X1 (Gallery cleanup task, parameters (all, pictures)). For example, for the operation command: "Activate the gallery cleanup tool and clean up similar pictures from 1 year ago," the task engine can define task X2 (Gallery cleanup task, parameters (1 year ago, similar, pictures)). For example, for the operation command: "Activate the gallery cleanup tool and clean up large videos," the task engine can define task X3 (Gallery cleanup task, parameters (large, video)).
[0071] After defining the task, the gallery agent can invoke the gallery cleanup tool to execute the gallery cleanup task. The gallery cleanup tool can launch the gallery application. The gallery application can display the image management interface and specific sub-interfaces of the image management interface (such as duplicate image sub-interface, similar image sub-interface, etc.) according to the instructions of the gallery cleanup tool, and display the corresponding images to be cleaned. After receiving a user confirmation of deletion (such as a user operation to delete a control in control 302 and / or pop-up window 306), the electronic device can perform the cleanup operation to clean up the selected images / videos.
[0072] After the cleanup is complete, the task engine module can call the response module of the AGI proxy framework to generate a task response, such as "XX photos cleaned" or "XX similar photos from one year ago cleaned." The electronic device can then display the task response. Optionally, the electronic device can navigate to a dialog interface and display the above response there. Alternatively, the electronic device can return to the image management interface and display the above response there. If an error occurs during cleanup and the operation is not completed, the task engine module can also call the response module of the AGI proxy framework to generate a task response, such as "Cleanup failed, please try again," informing the user of the cleanup failure.
[0073] like Figure 6 As shown, preferably, a gallery proxy can also be set up on the cloud, referred to as the cloud-side gallery proxy. When defining tasks and task parameters, the task orchestration module on the electronic device's local machine can also call the task orchestration module of the cloud-side gallery proxy to define tasks and task parameters. When generating task responses, the response module on the electronic device's local machine can also call the response module of the cloud-side gallery proxy to generate task responses.
[0074] Figure 7 This is a schematic diagram of another software architecture provided in the embodiments of this application.
[0075] like Figure 7 As shown, a gallery cleanup tool can be set up in the voice assistant. After processing the user's voice command through the audio computing engine to obtain the operation command for cleaning images, the voice assistant can launch the gallery cleanup tool. The gallery cleanup tool calls the task orchestration module of the cloud-based gallery agent to define the task and task parameters. After defining the task, the gallery cleanup tool can send the task and task parameters to the UI agent. In response to the above task, the UI agent can launch the gallery application and send the above task and task parameters to the gallery application. Then, the gallery application can display the image management interface and specific sub-interfaces of the image management interface according to the task parameters, displaying the corresponding images to be cleaned. After receiving the user's confirmation of deletion, the electronic device can perform the cleanup operation, cleaning up the selected pictures / videos. Similarly, after the cleanup is completed, the gallery cleanup tool can call the response module of the cloud-based gallery agent to generate a task response message, and then display the task response message to notify the user.
[0076] Not limited to mobile phones, tablets, and personal computers, electronic devices can also be wearable devices (such as watches, bracelets, glasses, etc.), in-vehicle devices, smart home devices, and / or smart city devices. This application does not impose any special restrictions on the specific type of electronic device.
[0077] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0078] The electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0079] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0080] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.
[0081] An NPU (Neural Processing Unit) is a neural network (NN) computing processor that, by drawing inspiration from the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, rapidly processes input information and can continuously learn on its own. In the embodiments of this application, electronic devices can perform intelligent cognition through the NPU, such as image recognition, speech recognition, and text understanding.
[0082] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0083] For example, processor 110 can couple to touch sensor 180K, charger, flash, camera 193, etc., via I2C bus interface. Processor 110 can couple to audio module 170 via I2S bus interface to realize communication between processor 110 and audio module 170. Audio module 170 can transmit audio signals to wireless communication module 160 via PCM interface to realize the function of answering phone calls through Bluetooth headset. Processor 110 can communicate with Bluetooth module in wireless communication module 160 via UART interface to realize Bluetooth function. Audio module 170 can transmit audio signals to wireless communication module 160 via UART interface to realize the function of playing music through Bluetooth headset. MIPI interface can be used to connect processor 110 to peripheral devices such as display 194 and camera 193. MIPI interface includes camera serial interface (CSI), display serial interface (DSI), etc. Processor 110 and camera 193 can communicate through CSI interface to realize the shooting function of electronic device. The processor 110 and display screen 194 can communicate via the DSI interface to realize the display function of the electronic device. The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, display screen 194, wireless communication module 160, audio module 170, sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, I2S interface, UART interface, MIPI interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device, or to transfer data between the electronic device and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as AR devices.
[0084] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a limitation on the structure of the electronic device. In other embodiments of this application, the electronic device may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0085] The charging management module 140 receives charging input from the charger. While charging the battery 142, the charging management module 140 can also power electronic devices via the power management module 141. The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc.
[0086] The wireless communication function of electronic devices can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0087] Antennas 1 and 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 150 can provide wireless communication solutions for electronic devices, including 2G / 3G / 4G / 5G. Wireless communication module 160 can provide wireless communication solutions for electronic devices, including wireless local area networks (WLANs) (such as Wireless Fidelity (Wi-Fi) networks), Bluetooth (BT), Global Navigation Satellite System (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies.
[0088] In some embodiments, antenna 1 of the electronic device is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling the electronic device to provide other wireless communication technologies. These other wireless communication technologies may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), etc. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), BeiDou Navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS), etc.
[0089] In this embodiment, through the wireless communication functions provided by antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor, the electronic device can access the cloud, call the intent understanding model on the cloud to perform intent understanding, and call the image library agent on the cloud to define tasks and task parameters and generate task response messages.
[0090] Electronic devices implement display functions through GPUs, display screens 194, and application processors. A GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor, used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0091] Display screen 194 is used for display, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD). The display panel can also be manufactured using organic light-emitting diodes (OLEDs), active-matrix organic light-emitting diodes (AMOLEDs), flexible light-emitting diodes (FLEDs), miniled, microled, micro-oled, quantum dot light-emitting diodes (QLEDs), etc. In some embodiments, the electronic device may include one or N displays 194, where N is a positive integer greater than 1.
[0092] In this embodiment, the electronic device can display information through the display functions provided by the GPU, display screen 194, and application processor. Figure 2 , Figures 3A-3C , Figure 5 The user interface shown.
[0093] The electronic device can implement shooting functions through an ISP, a camera 193, a video codec, a GPU, a display 194, and an application processor. The electronic device may include one or N cameras 193, where N is a positive integer greater than 1. In some embodiments, the ISP may be located within the camera 193.
[0094] Internal memory 121 may include one or more random access memory (RAM) and one or more non-volatile memory (NVM).
[0095] RAM can include static random-access memory (SRAM), dynamic random-access memory (DRAM), synchronous dynamic random-access memory (SDRAM), and double data rate synchronous dynamic random-access memory (DDR SDRAM, such as fifth-generation DDR SDRAM, generally referred to as DDR5 SDRAM). NVM can include disk storage devices and flash memory. RAM can be directly read and written by the processor 110 and can be used to store executable programs (such as machine instructions) of the operating system or other running programs, as well as user and application data. NVM can also store executable programs and user and application data, which can be pre-loaded into RAM for direct read and write by the processor 110.
[0096] The executable program code implementing the interaction method described in this application embodiment can be stored in NVM. After the electronic device is powered on, the electronic device can load the aforementioned executable program code stored in NVM into RAM, run the aforementioned executable program code, and thus provide the user with... Figure 1 and Figure 4 The convenient gallery cleanup function is shown.
[0097] The external memory interface 120 can be used to connect to an external NVM to expand the storage capacity of electronic devices.
[0098] The electronic device can implement audio functions through an audio module 170, a speaker 170A (also called a "loudspeaker," used to convert audio electrical signals into sound signals), a receiver 170B (also called a "handpiece," used to convert audio electrical signals into sound signals, mostly used in incoming call scenarios), a microphone 170C (also called a "microphone" or "voice transducer," used to convert sound signals into electrical signals), a headphone jack 170D, and an application processor. In this embodiment, the electronic device can collect sound signals through the microphone 170C, and then use them to recognize user voice commands. In other embodiments, the electronic device can be equipped with multiple microphones 170C to achieve functions such as data acquisition, noise reduction, and directionality. After generating a task response, the electronic device can also play the task response through the speaker 170A to prompt the user.
[0099] A pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. The pressure sensor 180A is typically located on the display screen 194. When a touch operation is applied to the display screen 194, the electronic device detects the intensity of the touch operation based on the pressure sensor 180A. The electronic device can also calculate the touch position based on the detection signal from the pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. A gyroscope sensor 180B can be used to determine the three-axis (i.e., x, y, and z-axis) angular velocity of the electronic device. An accelerometer sensor 180E can be used to determine the three-axis acceleration of the electronic device. The electronic device can determine its own posture and motion state through the gyroscope sensor 180B and the accelerometer sensor 180E. A barometric pressure sensor 180C is used to measure barometric pressure. A magnetic sensor 180D includes a Hall effect sensor for detecting changes in the ambient magnetic field. A distance sensor 180F is used to measure distance. A proximity sensor 180G can be used to detect objects approaching, such as when the electronic device is held close to the ear for a call, in pocket mode, etc. The ambient light sensor 180L is used to sense ambient light intensity. The fingerprint sensor 180H is used to collect fingerprints. Electronic devices can utilize the collected fingerprint characteristics to achieve fingerprint unlocking, app access lock, fingerprint photography, fingerprint call answering, etc. The temperature sensor 180J is used to detect temperature. The touch sensor 180K can be placed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also known as a "touchscreen". The touch sensor 180K is used to detect touch operations applied to or near it. The touch sensor 180K can transmit the detected touch operation to the application processor to determine the touch event type, and then provide visual output related to the touch operation through the display screen 194. The bone conduction sensor 180M can acquire vibration signals from the human vocal cords, blood pressure signals, etc.
[0100] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch buttons. Electronic devices can receive button input via buttons 190, generating key signal inputs related to user settings and function control. Motor 191 can generate vibration alerts. Indicator 192 can be an indicator light, used to indicate charging status, battery level changes, messages, missed calls, notifications, etc. SIM card interface 195 is used to connect a SIM card.
[0101] The term "user interface (UI)" used in the specification, claims, and drawings of this application refers to the medium through which an application or operating system interacts and exchanges information with the user. It converts information from its internal form to a form acceptable to the user. The user interface of an application is source code written in a specific computer language such as Java or Extensible Markup Language (XML). This source code is parsed and rendered on the terminal device, ultimately presenting user-recognizable content such as images, text, and buttons. Controls, also known as widgets, are the basic elements of the user interface. Typical controls include toolbars, menu bars, text boxes, buttons, scroll bars, images, and text. The attributes and content of controls in the interface are defined using tags or nodes, such as XML tags. <textview> 、 <imgview> 、 <videoview>Nodes define the controls contained in the interface. A node corresponds to a control or property in the interface, and after parsing and rendering, the node is presented as the content visible to the user. In addition, many applications, such as hybrid applications, often contain web pages within their interfaces. A web page, also known as a webpage, can be understood as a special control embedded in the application interface. Web pages are source code written in a specific computer language, such as Hypertext Markup Language (HTML), Cascading Style Sheets (CSS), JavaScript (JS), etc. Web page source code can be loaded and displayed as user-readable content by a browser or a web page display component with browser-like functionality. The specific content contained in a webpage is also defined through tags or nodes in the webpage source code; for example, HTML uses tags or nodes to define the content. 、 、 <video> 、 <canvas>Used to define the elements and attributes of a webpage.
[0102] The most common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.
[0103] As used in the specification and appended claims of this application, the singular expressions "a," "an," "the," "the," "the," and "this" are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the listed items. As used in the above embodiments, depending on the context, the term "when" can be interpreted as meaning "if..." or "after..." or "in response to determining..." or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining..." or "in response to determining..." or "when (the stated condition or event) is detected" or "in response to detecting (the stated condition or event)."
[0104] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0105] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.< / canvas> < / video> < / videoview> < / imgview> < / textview>
Claims
1. An interaction method applied to an electronic device, characterized in that, The method includes: Receive the user's first voice command; When the first voice command includes a first keyword, in response to the first voice command, a first interface of the gallery is displayed, the first interface including a first option; in response to user operation on the first option, a second interface is displayed, the second interface displaying multiple sets of images, each set of images including multiple similar or identical images; When the first voice command includes the first keyword and the second keyword, the second keyword corresponds to the first option, and in response to the first voice command, the second interface is directly displayed; In response to a user operation on the first control in the second interface, one image is retained for each group of images in the second interface; The retained image is the highest-rated image in the group, or the highest-rated image carrying a first tag; the first tag indicates that the user likes the image and / or the user has edited the image; the rating of an image includes the image's rating in N aspects, which include one or more of the following: image quality clarity, subject, object state, and composition.
2. The method according to claim 1, characterized in that, The first voice command also includes a third keyword, and the multiple images in the second interface are images that match the third keyword.
3. The method according to claim 2, characterized in that, The third keyword includes one or more of the following: subject, shooting time, shooting location, and color.
4. The method according to claim 1, characterized in that, The method further includes: When the first voice command does not include the first keyword, a prompt message is displayed on the dialogue interface, which is used to instruct the user to re-enter the voice command.
5. The method according to claim 1, characterized in that, The images include photographs and videos, and the photographs include still photographs composed of single-frame images and moving photographs composed of multiple-frame images.
6. The method according to claim 1, characterized in that, The step of directly displaying the second interface in response to the first voice command includes: In response to the first voice command, the first interface is displayed first, and then the user automatically jumps to the second interface.
7. The method according to claim 1, characterized in that, Alternatively, the retained image is an image obtained by fusing one or more images carrying the first label from the group of images.
8. The method according to claim 1, characterized in that, The method further includes: When the first voice command includes the first keyword and the fourth keyword, the fourth keyword corresponds to the second option in the first interface. In response to the first voice command, a third interface is displayed, which contains multiple blurred images, AI-enhanced images, or ultra-large videos. In response to a user operation on a first control in the third interface, the image in the third interface is deleted.
9. An electronic device, characterized in that, It includes one or more processors and one or more memories; wherein the one or more memories are coupled to the one or more processors, and the one or more memories are used to store a computer program that, when the one or more processors execute the computer program, causes the method as described in any one of claims 1-8 to be performed.
10. A chip system applied to an electronic device, the chip system comprising one or more processors, characterized in that, The processor is used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1-8.
11. A computer program product containing instructions, characterized in that, When the computer program product is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1-8.
12. A computer-readable storage medium comprising a computer program, characterized in that, When the computer program is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Picture cleanup method, picture cleanup device and terminal device
CN104809198A
Picture processing method and mobile terminal
CN106649759A
Photo album management method, storage medium and electronic equipment
CN110516083A