Image processing method and device, electronic equipment and storage medium
By invoking a virtual assistant in any application and utilizing an image processing model for image processing, the problem of low image processing efficiency in existing technologies is solved, achieving a more efficient and user-friendly image processing experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, image processing needs to be performed in specific applications, resulting in inefficiency and a poor user experience.
By invoking a virtual assistant in any application, the system processes the original image using an image processing model to generate the target image, including visual positioning, image extraction, and restoration models, supporting users to perform image processing operations within any interface.
It improves image processing efficiency and user experience, simplifies operation processes, and meets users' diverse image processing needs.
Smart Images

Figure CN121807205A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal equipment technology, and in particular to an image processing method, apparatus, electronic device and storage medium. Background Technology
[0002] With the rapid development of electronic devices (such as smartphones and tablets), image processing technology has made significant progress and become increasingly intelligent. Image processing can be used for tasks such as background removal or removing minor objects from images. For example, image processing tools can be used to extract specific objects from an image or remove unwanted minor objects, making the image clearer and more professional.
[0003] Currently, electronic devices typically require specific applications or software to perform image cutout or removal of minor objects, resulting in low image processing efficiency. Summary of the Invention
[0004] To address the aforementioned issues, this application provides an image processing method that allows electronic devices to invoke a virtual assistant in any application. By processing the original image, the method improves image processing efficiency and ensures a positive user experience.
[0005] To achieve the above objectives, firstly, a first interface of a first application is displayed, the first interface including an original image;
[0006] The virtual assistant is invoked on the first interface to display the first control corresponding to the virtual assistant; in response to the user's first operation on the first control, the first interface is updated to obtain the updated second interface; wherein, the second interface includes a first target image; the first target image is the updated image of the original image.
[0007] The image processing method provided in this application embodiment allows an electronic device to invoke a virtual assistant in a first interface of a first application to provide recommendation services. Subsequently, if the user performs a first operation in the first interface, the electronic device will process the original image to generate a first target image. In this way, the electronic device can invoke the first interface containing the virtual assistant in any application, eliminating the need for the user to switch between applications multiple times to restrict image processing to a specific program, thereby improving image processing efficiency and ensuring a better user experience.
[0008] In one feasible implementation, in response to a user's first operation on a first control, updating the first interface to obtain an updated second interface includes: in response to a user's first operation on an input control within the first control, updating the first interface to obtain an updated second interface; or, in response to a user's first operation on a selection control within the first control, updating the first interface to obtain an updated second interface. In this way, the electronic device can operate on different controls in the first interface, thereby processing the original image to obtain a first target image, ensuring the diversity of user choices.
[0009] In one feasible implementation, in response to a user's first operation on an input control in a first control, updating a first interface to obtain an updated second interface includes: in response to the user's first operation on the input control, processing the original image using an image processing model corresponding to the second operation to generate a first target image; updating the original image in the first interface to the first target image to obtain the second interface. Thus, when the user triggers the second operation, the electronic device will capture the original image in the first interface and display the first interface containing the first control corresponding to the virtual assistant, providing recommendation services to the user.
[0010] In one feasible implementation, in response to a user's first operation on an input control, an image processing model is used to process the original image corresponding to the first operation to generate a first target image. This includes: in response to the user's first operation on the input control, obtaining instruction information; and using the image processing model to process the original image corresponding to the instruction information to generate the first target image. In this way, electronic devices can invoke a virtual assistant in any application to process the original image using instruction information, improving image processing efficiency and ensuring a positive user experience.
[0011] In one feasible implementation, an image processing model is used to process the original image according to the instruction information to generate a first target image. This includes: using a visual positioning model and a first image extraction model to process the original image according to the instruction information to generate a first target image with at least one first target subject; or using a second image extraction model and an image restoration model to process the original image according to the instruction information to generate a first target image without at least one second target subject. In this way, the electronic device can respond to different instructions input by the user in the input controls on the first interface, providing different image processing scenarios to meet the diverse needs of the user and thus improve the user experience.
[0012] In one feasible implementation, a visual positioning model and a first image extraction model are used to process the original image in accordance with the instruction information to generate a first target image having at least one first target subject. This includes: processing the original image using the visual positioning model in accordance with the instruction information to generate a first subject region; the first subject region is the area enclosed by a straight-line bounding box representing the edge of the first target subject; and processing the first subject region using the first image extraction model to generate the first target image having at least one first target subject. In this way, image processing using the visual positioning model can more accurately identify the position of the first subject, thereby avoiding the influence of non-first subject parts on subsequent processing.
[0013] In one feasible implementation, a visual positioning model is used to process the original image according to the instruction information to generate a first subject region. This includes: extracting instruction features corresponding to the instruction information; extracting at least one image feature corresponding to the original image; determining a first image feature from the at least one image feature based on the instruction features; and determining the first subject region corresponding to the instruction feature in the original image based on the first image feature. This allows for flexible adjustment of the recognized image features according to different instruction features, enabling it to adapt to diverse application scenarios and ensuring a good user experience.
[0014] In one feasible implementation, determining a first image feature from at least one image feature based on instruction features includes: determining the similarity between each image feature and the instruction feature based on the image features and the instruction feature; and determining the first image feature corresponding to the instruction feature based on the similarity. In this way, by calculating the similarity between the instruction feature and the image feature, the image region pointed to by the user's intent can be accurately located, thereby improving the accuracy of recognition.
[0015] In one feasible implementation, a second image extraction model and an image inpainting model are used to process the original image according to the instruction information to generate a first target image that does not have at least one second target subject. This includes: when the amount of information in the instruction information meets a first preset threshold, using a visual positioning model to process the original image according to the instruction information to generate an output image of the bounding box of the second target subject; using the second image extraction model to process the output image to generate a second target image; and using the image inpainting model to process the second target image and the original image to generate a first target image that does not have at least one second target subject. Thus, when the amount of information in the instruction information is relatively rich, the user only needs to provide the instruction information to obtain a result image with the specific target subject removed, simplifying the user operation process.
[0016] In one feasible implementation, when the amount of information in the instruction information meets a first preset threshold, a visual positioning model is used to process the original image in accordance with the instruction information to generate an output image of the bounding box of the second target subject. This includes: extracting instruction features from the instruction information and at least one image feature from the original image, provided the amount of information in the instruction information meets the first preset threshold; determining a second image feature from the at least one image feature based on an image clustering library and the instruction features; and obtaining a second subject region in the original image corresponding to the instruction feature based on the second image feature. The image clustering library includes a correspondence between instruction features and preset image features. Thus, based on the image clustering library, the electronic device can recognize instruction information with pronoun representations, allowing users to express themselves using more natural language without worrying about the electronic device's inability to understand.
[0017] In one feasible implementation, determining a second image feature from at least one image feature based on an image clustering library and instruction features includes: if the image clustering library includes preset image features corresponding to the instruction features, determining the second image feature from at least one image feature based on the preset image features. In this way, preset image features corresponding to the instruction features are predefined in the image clustering library, allowing for direct matching using these preset features and accelerating the processing speed.
[0018] In one feasible implementation, determining a second image feature from at least one image feature based on an image clustering library and instruction features includes: if the image clustering library does not include a preset image feature corresponding to the instruction feature, displaying a third interface, the third interface including an image control corresponding to at least one image feature; and determining the second image feature from at least one image feature in response to a third operation by the user on the image control. Thus, when the user's requirements exceed the scope of the image clustering library, the electronic device provides an additional third interface, allowing the user to manually select or adjust, thereby increasing the flexibility of the electronic device and enabling it to meet more diverse user needs.
[0019] In one feasible implementation, a second image extraction model and an image inpainting model are used to process the original image in accordance with the instruction information to generate a first target image that does not have at least one second target subject. This includes: when the amount of information in the instruction information does not meet a first preset threshold, using the second image extraction model to process the original image in accordance with the instruction information to generate a second target image; and using the image inpainting model to process the second target image and the original image to generate a first target image that does not have at least one second target subject. In this way, even when the amount of instruction information is limited, the user does not need to provide detailed instruction information, and the electronic device can still perform basic image processing tasks, making operation simpler.
[0020] In one feasible implementation, in response to a user's first operation on a selection control in a first control, an updated second interface is obtained, including: based on the original image, determining at least one target subject bounding box corresponding to at least one first target subject in the original image, and at least one non-target subject bounding box corresponding to at least one second target subject; if the number of target subject bounding boxes is less than a third preset threshold and the number of non-target subject bounding boxes is greater than the third preset threshold, displaying a first control corresponding to the virtual assistant; wherein the first control includes an input control. In this way, the electronic device can not only more accurately identify user intent but also provide a better user experience.
[0021] In one feasible implementation, in response to a user's first operation on a selection control in a first control, the first interface is updated to obtain an updated second interface, including: based on the original image, determining at least one target subject bounding box corresponding to at least one first target subject in the original image, and at least one non-target subject bounding box corresponding to at least one second target subject; if the number of target subject bounding boxes is greater than a third preset threshold and the number of non-target subject bounding boxes is less than the third preset threshold, displaying the first control corresponding to the virtual assistant; wherein the first control includes an input control. In this way, the electronic device can not only more accurately identify user intent but also provide a better user experience.
[0022] In one feasible implementation, in response to a user's first operation on a selection control within a first control, updating the first interface to obtain an updated second interface includes: in response to a user's second operation on the selection control, obtaining control instructions; using an image processing model to process the original image corresponding to the control instructions to generate a first target image; and updating the original image in the first interface to the first target image to obtain the second interface. In this way, electronic devices can invoke a virtual assistant in any application, processing the original image through control instructions, improving image processing efficiency and ensuring a positive user experience.
[0023] In one feasible implementation, an image processing model is used to process the original image in accordance with control instructions to generate a first target image. This includes: processing the original image using the image processing model in accordance with control instructions to generate a first target image having at least one first target subject; or processing the original image using an image recognition model and an image restoration model in accordance with control instructions to generate a first target image not having at least one second target subject. In this way, the electronic device can respond to different instructions selected by the user in the selection controls on the first interface, providing different image processing scenarios to meet the diverse needs of the user and thus improving the user experience.
[0024] In one feasible implementation, an image processing model is used to process the original image in accordance with control instructions to generate a first target image having at least one first target subject. This includes: using a first image extraction model to process the original image in accordance with control instructions to generate a first target image having at least one first target subject. In this way, by processing the image using the first image extraction model, the position of the first subject can be identified more accurately, thereby avoiding the influence of non-first subject parts on subsequent processing.
[0025] In one feasible implementation, activating a virtual assistant on a first interface to display a first control corresponding to the virtual assistant includes: processing the original image using a third image extraction model to generate a third subject region in the original image, wherein the third subject region is the area enclosed by the curved bounding box of the edge of the first target subject; and displaying the first control corresponding to the virtual assistant when the third subject region meets a fourth preset threshold; wherein the first control includes a selection control. In this way, by using a third image extraction model to generate a third subject region, the edges of the target subject can be captured more accurately, even irregularly shaped objects can be accurately identified. Simultaneously, when the third subject region occupies a sufficiently large proportion of the original image, a first interface that conforms to the user's preferences and habits can be displayed, improving the user experience.
[0026] In one feasible implementation, an image recognition model and an image restoration model are used to process the original image in accordance with control instructions to generate a first target image that does not have at least one second target subject. This includes: using the image recognition model to process the original image in accordance with control instructions to generate a second target image; and using the image restoration model to process the second target image and the original image to generate a first target image that does not have at least one second target subject. In this way, by using control instructions to guide the operation of the image recognition model and the image restoration model, the second target subject that the user wants to remove can be identified and processed more accurately, improving the user experience.
[0027] In one feasible implementation, an image recognition model is used to process the original image in accordance with control instructions to generate a second target image. This includes: using an image segmentation model to process the original image in accordance with control instructions to generate a second subject region corresponding to the second target subject; and using a second extraction model to process the second subject region to determine the second target image. In this way, electronic devices can improve the accuracy and efficiency of image processing, ensuring a better user experience.
[0028] In one feasible implementation, an image segmentation model is used to process the original image according to the control instructions, generating a second subject region corresponding to the second target subject. This includes: extracting text features from the control instructions; extracting at least one image feature from the original image; the image features include the second target subject image features; and obtaining the second subject region in the original image corresponding to the instruction features based on the correlation between each image feature and the second target subject image features. In this way, based on the correlation between subjects in the original image, the electronic device can more accurately identify unwanted secondary objects (passersby), improving the accuracy of image processing by the electronic device.
[0029] In one feasible implementation, activating a virtual assistant on a first interface to display a first control corresponding to the virtual assistant includes: determining, based on the original image, at least one target subject bounding box corresponding to a first target subject and at least one non-target subject bounding box corresponding to a second target subject in the original image; displaying the first interface when the number of target subject bounding boxes is greater than a third preset threshold and the number of non-target subject bounding boxes is greater than a second preset threshold; wherein the first interface includes a first control corresponding to the virtual assistant, and the first control includes a selection control. This provides a user-friendly interaction method, not only improving the efficiency and accuracy of image processing but also enhancing the user experience, making complex image processing tasks simpler and easier.
[0030] In one feasible implementation, based on the original image, determining at least one target bounding box corresponding to a first target subject and at least one non-target bounding box corresponding to a second target subject in the original image includes: processing the original image using a subject recognition model to generate a third target image; the third target image includes the bounding box corresponding to the first target subject; processing the third target image using an instance segmentation model to generate a fourth target image; the fourth target image includes at least one target bounding box corresponding to the first target subject and at least one non-target bounding box corresponding to the second target subject. Thus, by first generating the third target image using a subject recognition model and then generating the fourth target image using an instance segmentation model, the various target subjects in the image can be identified and segmented more accurately, even if these subjects are close to or overlap each other, improving the accuracy of image processing.
[0031] To achieve the above objectives, in a second aspect, this application provides an image processing apparatus that has the function of implementing the electronic device behavior in the image processing method of the first aspect. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described function.
[0032] Thirdly, this application provides an electronic device, including: a display screen, a memory, and one or more processors; the display screen, the memory, and the processors are coupled; wherein the memory stores computer program code, the computer program code including computer instructions, and when the computer instructions are executed by the processor, the electronic device performs the image processing method provided in the first aspect above.
[0033] Fourthly, this application provides a computer-readable storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform the image processing method provided in the first aspect above.
[0034] Fifthly, this application provides a computer program product that, when run on a computer, causes the computer to perform the image processing method provided in the first aspect above.
[0035] It is understood that the beneficial effects that the technical solutions provided in the second to fifth aspects described above can be achieved by referring to the beneficial effects of the first aspect and any of its optional embodiments, which will not be repeated here. Attached Figure Description
[0036] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This application provides an embodiment of an image comparison diagram between an original image and a target image;
[0038] Figure 2 This is a schematic flowchart of an image processing method provided in an embodiment of this application;
[0039] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0040] Figure 4 This is a schematic diagram of the layered architecture of the software system of the electronic device provided in the embodiments of this application;
[0041] Figure 5 This is a schematic flowchart of an image processing method provided in an embodiment of this application;
[0042] Figure 6 This is a schematic diagram illustrating how a first operation is triggered on a first interface, as provided in an embodiment of this application.
[0043] Figure 7 This is a schematic diagram of a first interface provided in an embodiment of this application;
[0044] Figure 8A This is a flowchart illustrating an interface for acquiring a first target image, provided in an embodiment of this application.
[0045] Figure 8B This is another flowchart illustrating an interface for acquiring a first target image, as provided in this application embodiment.
[0046] Figure 8C This is yet another flowchart illustrating an interface for acquiring a first target image, provided in an embodiment of this application.
[0047] Figure 9A This is a schematic diagram of a process for content extraction from an image provided in an embodiment of this application;
[0048] Figure 9B This is a schematic diagram illustrating an embodiment of obtaining instruction information provided in this application;
[0049] Figure 10 This is a schematic diagram of the structure of a visual positioning model provided in an embodiment of this application;
[0050] Figure 11 This is a schematic diagram of an image cropping interface provided in an embodiment of this application;
[0051] Figure 12 This is a schematic diagram of a process for background purification of an original image provided in an embodiment of this application;
[0052] Figure 13 This is a schematic diagram of a structure for background purification of an original image provided in an embodiment of this application;
[0053] Figure 14 This is another schematic diagram of the structure of a visual positioning model provided in an embodiment of this application;
[0054] Figure 15 This is a schematic diagram illustrating a method for determining second image features according to an embodiment of this application;
[0055] Figure 16 This is another schematic diagram illustrating the determination of a second image feature provided in an embodiment of this application;
[0056] Figure 17 This is another schematic diagram of a process for background purification of an image provided in an embodiment of this application;
[0057] Figure 18 This is another structural schematic diagram of an image background purification method provided in an embodiment of this application;
[0058] Figure 19This is a flowchart illustrating an interface for displaying a first control corresponding to a virtual assistant, provided in an embodiment of this application.
[0059] Figure 20 This application provides a schematic diagram of determining a target subject frame according to an embodiment;
[0060] Figure 21 This application provides a schematic diagram for determining a target subject frame and a non-target subject frame according to an embodiment of the present application;
[0061] Figure 22A This is a schematic diagram of a structure for content extraction from an image provided in an embodiment of this application;
[0062] Figure 22B This is a schematic diagram of a recommended selection control provided in an embodiment of this application;
[0063] Figure 23 This is another schematic flowchart illustrating a first interface for displaying a first control corresponding to a virtual assistant, provided in an embodiment of this application.
[0064] Figure 24 This is another structural diagram illustrating background purification of an original image provided in an embodiment of this application;
[0065] Figure 25 This is a schematic diagram illustrating the acquisition of a second subject region according to an embodiment of this application;
[0066] Figure 26 This is yet another flowchart illustrating the display of a first interface provided in an embodiment of this application;
[0067] Figure 27 This is another schematic flowchart of an image processing method provided in an embodiment of this application;
[0068] Figure 28 This is another schematic flowchart of an image processing method provided in an embodiment of this application;
[0069] Figure 29 This is another schematic flowchart of an image processing method provided in an embodiment of this application;
[0070] Figure 30 This application provides a structural block diagram of a chip system. Detailed Implementation
[0071] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are all within the protection scope of this application.
[0072] In the following description, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0073] Furthermore, in this application, directional terms such as "upper," "lower," "inner," and "outer" are defined relative to the indicated placement of the components in the accompanying drawings. It should be understood that these directional terms are relative concepts, used for relative description and clarification, and can change accordingly depending on the placement of the components in the accompanying drawings.
[0074] To facilitate understanding of the technical solutions of the embodiments of this application by those skilled in the art, the technical terms involved in the embodiments of this application will be explained below.
[0075] A virtual assistant is a software tool that helps users complete various tasks and provides support by simulating the behavior of a human assistant. Virtual assistants can communicate with users through different interaction methods, such as text, voice, and graphical interface clicks. Virtual assistants can be widely used on electronic devices. They may include service recommendations, assistant functions, intelligent agents, and visibility functions.
[0076] Service recommendation is a system that uses data analysis and algorithms to provide personalized suggestions to users based on the recognition of content on the interface. It typically recommends products, content, or services based on the user's historical behavior, interests, or preferences. Examples include YOYO byside and YOYO inside.
[0077] Virtual assistants, through voice recognition and natural language understanding, allow users to converse with them, ask questions, or make requests, making interaction more natural and convenient. Examples include YOYO Assistant.
[0078] Intelligent agents continuously learn from machine learning algorithms to analyze user feedback, constantly optimizing their recommendations and services to generate one-click intelligent recommendations. Examples include AI agents.
[0079] Virtual Vision is a service that makes recommendations based on content in a preview stream. For example, when an electronic device invokes a virtual assistant in the preview stream interface, it can capture a frame from the preview stream as the original image. Alternatively, the electronic device can automatically capture a frame as the original image after the preview stream has stabilized. In addition, the original image can also be captured through a screenshot. Examples include YOYO Vision and HONOR Lens.
[0080] Image matting, or cutout, is a technique in image processing that refers to the process of separating a specific object (usually the foreground or subject) from the background of an image. The main purpose of cutout is to extract an object from the original image and then place it in another background or perform further editing operations.
[0081] The bottom toolbar provides shortcuts for quick access to frequently used functions or pages. In this embodiment, the virtual assistant's recommendation service can be activated by long-pressing the bottom toolbar.
[0082] Image processing models are a class of machine learning models specifically designed for analyzing, processing, and generating images. These models learn the features and patterns of large amounts of image data to perform a variety of complex image processing tasks, such as image classification, object detection, image segmentation, and style transfer.
[0083] Natural Language Processing (NLP) is a method of computer-to-human language interaction. Its goal is to enable electronic devices to understand, interpret, and generate natural language, allowing them to process and analyze language data like humans do.
[0084] A slot is a concept in Natural Language Processing (NLP) used to extract key information from user input. Each slot typically corresponds to a specific data field or variable. Slot extraction refers to the process of identifying and extracting this key information from the user's natural language input.
[0085] Intern Visual-Language Model 2 (InternVL2) is an advanced visual-language model designed to handle various tasks between images and text, such as image-text matching.
[0086] The information content of instruction information generally refers to the level of detail or specificity of the content or description contained in the instruction. For instruction information in image processing tasks, the amount of information directly affects the model's ability to understand and execute the task. In this embodiment, the amount of instruction information can include low instruction information and high instruction information. For example, low instruction information is "extract the girl," and high instruction information is "extract the long-haired girl from the image."
[0087] Image inpainting models are techniques used to restore or fill in missing or unwanted parts of an image. These models are typically used to remove distracting elements from an image, repair damaged images, or fill in occluded areas. In this application embodiment, they are used to remove unwanted subjects from the original image.
[0088] The embodiments of this application will now be described with reference to the accompanying drawings.
[0089] Images, as a carrier of information, have come to occupy a pivotal position in the digital world. They are not only used to express emotions and convey information, but have also become an important tool for creativity and content production.
[0090] With the widespread use of electronic devices such as smartphones, tablets, and laptops, users can now easily obtain raw images through various applications.
[0091] For example, users can obtain raw images from system utility applications on electronic devices, such as gallery applications. Users can also obtain raw images from social applications on electronic devices. While images have become an important part of the digital world, in real-world applications, the raw images obtained by users are often not ideal or fully meet their needs. This phenomenon is usually limited by various scenarios, such as shooting environment, lighting, composition, and background, which may cause the captured or downloaded images to not meet the user's expectations.
[0092] For example, Figure 1 This application provides an embodiment of an image comparison diagram between an original image and a target image.
[0093] like Figure 1 As shown in (a) in the figure, Figure 1 The image shown in (a) is the original image acquired by the user. The original image includes multiple image objects, such as the subject 10 looking at the camera, the pedestrian 11 hurrying along, and the scenery. The original image includes the pedestrian 11 hurrying along, which is not an image that fully meets the user's needs.
[0094] like Figure 1 As shown in (b) in the figure, Figure 1The image shown in (b) is the target image (background cleansing) required by the user, which means that the target image does not include the secondary object "pedestrian 11 who is on the way".
[0095] like Figure 1 As shown in (c) in the figure, Figure 1 The image shown in (c) is the target image (content extraction) required by the user, that is, the target image is an object that only contains "the subject 10 looking at the camera".
[0096] It is evident that when the acquired original image is not ideal and does not fully meet the requirements, further processing of the original image is necessary to generate the target image required by the user and improve the user experience.
[0097] Figure 2 This is a flowchart illustrating an image processing method provided in an embodiment of this application.
[0098] Users can obtain the original image from any application; let's take a social application as an example.
[0099] like Figure 2 As shown, Figure 2 (a) shows the main interface 21 of the electronic device. The main interface 21 can display icons for multiple applications, such as a clock application, a calendar application, a memo application, and a social networking application 211. The electronic device can launch the application in response to the user's click on its icon for easy use. The main interface 21 can also display a status bar, which may include one or more signal strength indicators for mobile communication signals, one or more signal strength indicators for Wi-Fi signals, a battery indicator, and a time indicator.
[0100] The electronic device responds to a user's click operation 01 on the social application 211 in the main interface 21, and jumps to the main social interface (not shown in the figure). The main social interface includes multiple settings, such as user chat settings, account and security settings, etc. The electronic device responds to a user's click operation on any user chat settings in the main social interface, and jumps to the chat interface (not shown in the figure). The chat interface includes image thumbnails.
[0101] It should be noted that the image thumbnail obtained in the social chat interface mentioned above can also be obtained in other settings, and no specific limitation is made here.
[0102] The electronic device responds to the user's click on the image thumbnail in the chat interface and jumps to the first image interface 22 (e.g., ...). Figure 2(as shown in (b)). The first image interface 22 may include the original image corresponding to the image thumbnail and operation buttons, such as the download button 221, other buttons, etc.
[0103] In response to the user's click operation 02 on the download button 221 in the first image interface 22, the electronic device jumps to the second image interface 23 (e.g., Figure 2 (as shown in (c)). The second image interface 23 may include a prompt text box. The prompt text box provides information about the original image's save location and indicates whether the download was successful. In response to a user's left swipe operation 03 on the second image interface 23, the electronic device jumps to the main interface 21 (as shown in (c)). Figure 2 (as shown in (d)). The electronic device continues to respond to the user's click operation 04 on the gallery application 212 in the main interface 21, and jumps to the third image interface 24 (as shown in (d)). Figure 2 (as shown in (e)).
[0104] Furthermore, in response to a user's long-press operation 05 on the subject 241 in the original image on the third image interface 24, the electronic device jumps to the fourth image interface 25 (e.g., Figure 2 (as shown in (f)). In the fourth image interface 25, including the subject 251 separated from the original image background, the target image (content extraction) is obtained. Further, in response to the user's click operation 06 on the edit button 242 in the original image in the third image interface 24, the electronic device performs corresponding image processing to obtain the target image, as shown in (f). Figure 2 The fifth image interface 26 is shown in (g).
[0105] In summary, electronic devices require specific applications (gallery applications) to process the original images. Therefore, if the original image is not in such an application, it needs to be stored there first before further processing, such as image cutout or removal of minor objects. This requires users to repeatedly switch pages, resulting in low image processing efficiency.
[0106] To address the aforementioned issues, this application provides an image processing method that eliminates the need for users to process raw images within specific applications, thereby improving image processing efficiency and ensuring a superior user experience.
[0107] The image processing method provided in this embodiment can be applied to electronic devices. In some embodiments, the electronic device may be a mobile phone, tablet computer, handheld computer, personal computer (PC), ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), augmented reality (AR) device, virtual reality (VR) device, artificial intelligence (AI) device, wearable device, etc. This application embodiment does not impose any special limitations on the specific type of electronic device.
[0108] For example, Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, with a mobile phone as an example.
[0109] like Figure 3 As shown, the mobile phone may include a processor 310, an external memory interface 320, an internal memory 321, a universal serial bus (USB) interface 330, a charging management module 340, a power management module 341, a battery 342, antenna 1, antenna 2, a mobile communication module 350, a wireless communication module 360, a sensor module 380, a display screen 393, a subscriber identification module (SIM) card interface 394, and a camera 395, etc. The sensor module 380 may include a pressure sensor 380A, a fingerprint sensor 380B, an ambient light sensor 380C, a gyroscope sensor 380D, a temperature sensor 380E, etc.
[0110] Processor 310 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. The different processing units may be independent devices or integrated into one or more processors.
[0111] The controller can be the nerve center and command center of a mobile phone. Based on the instruction opcode and timing signals, the controller generates operation control signals to control the fetching and execution of instructions.
[0112] The processor 310 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 310 is a cache memory. This memory can store instructions or data that the processor 310 has just used or that are used repeatedly. If the processor 310 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 310, and thus improves the efficiency of the system.
[0113] In some embodiments, processor 310 may include one or more interfaces.
[0114] The external memory interface 320 can be used to connect to external non-volatile memory, thereby expanding the phone's storage capacity. The external non-volatile memory communicates with the processor 310 through the external memory interface 320 to perform data storage functions. For example, music, video, and other files can be saved in the external non-volatile memory.
[0115] Internal memory 321 may include one or more random access memory (RAM) and one or more non-volatile memory (NVM).
[0116] The charging management module 340 is used to receive charging input from a power supply device (such as a charger, laptop power supply, etc.). The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 340 can receive charging input from the wired charger via a USB interface 330. In some wireless charging embodiments, the charging management module 340 can receive wireless charging input via the wireless charging coil of a mobile phone.
[0117] While charging the battery 342, the charging management module 340 can also supply power to the mobile phone through the power management module 341. Specifically, the battery 342 can be composed of multiple batteries connected in series. The power management module 341 connects the battery 342, the charging management module 340, and the processor 310.
[0118] The power management module 341 connects the battery 342, the charging management module 340, and the processor 310. The power management module 341 receives input from the battery 342 and / or the charging management module 340, providing power to the processor 310, internal memory 321, display screen 393, camera 395, and wireless communication module 360. The power management module 341 can also monitor parameters such as battery voltage, current, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 341 may also be located within the processor 310.
[0119] The wireless communication function of a mobile phone can be achieved through antenna 1, antenna 2, mobile communication module 350, wireless communication module 360, modem, and baseband processor.
[0120] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in a mobile phone can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization.
[0121] The mobile communication module 350 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use in mobile phones. The mobile communication module 350 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 350 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 350 can be housed in the processor 310. In some embodiments, at least some functional modules of the mobile communication module 350 and at least some modules of the processor 310 can be housed in the same device.
[0122] The wireless communication module 360 can provide solutions for wireless communication applications in mobile phones, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 360 can be one or more devices integrating at least one communication processing module. The wireless communication module 360 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signal, and sends the processed signal to processor 310. The wireless communication module 360 can also receive signals to be transmitted from processor 310, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0123] In some embodiments, the sensor module 380 may include a pressure sensor 380A, a fingerprint sensor 380B, an ambient light sensor 380C, a gyroscope sensor 380D, a temperature sensor 380E, etc.
[0124] The pressure sensor 280A is used to sense pressure signals and convert them into electrical signals. In some embodiments, the pressure sensor 280A can be disposed on the display screen 293. There are many types of pressure sensors 280A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to the pressure sensor 280A, the capacitance between the electrodes changes. The mobile phone 200 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to the display screen 293, the mobile phone 200 detects the intensity of the touch operation based on the pressure sensor 280A. The mobile phone 200 can also calculate the touch position based on the detection signal from the pressure sensor 280A. In this embodiment, the pressure sensor 280A is used to detect the user's operation on the first control, and then perform image processing on the original image.
[0125] The fingerprint sensor 380B is used to collect fingerprints.
[0126] The 380C ambient light sensor is used to detect ambient light intensity.
[0127] The gyroscope sensor 380D can be used to determine the motion attitude of an electronic device. In some embodiments, the gyroscope sensor 380D can be used to determine the angular velocity of the electronic device about three axes (i.e., the x, y, and z axes). The gyroscope sensor 380D can be used for image stabilization.
[0128] The 380E temperature sensor is used to detect temperature.
[0129] In some embodiments, a mobile phone may include one or N cameras 395, where N is a positive integer greater than 1. The type of camera 395 can be distinguished based on hardware configuration and physical location. In this embodiment, the cameras can be used to capture images to obtain raw images.
[0130] The mobile phone uses a GPU, a display screen (393), and an application processor to achieve its display function.
[0131] The ISP (Image Signal Processor) is used to process data fed back from the camera 395. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise and brightness. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 395. The camera 395 is used to capture still images or videos. In the embodiments of this application, the ISP can be used to adjust parameters in the shooting scene.
[0132] The display screen 393 is used to display images, videos, etc. In this embodiment, the display screen 393 can be used to display pages required by the mobile phone (e.g., a first interface, etc.).
[0133] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a limitation on the structure of the mobile phone. In other embodiments of this application, the mobile phone may also adopt different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0134] Of course, this is understandable. Figure 3 The illustration shown is merely an example of an electronic device in the form of a mobile phone. If the electronic device is a tablet, handheld computer, PC, PDA, wearable device (such as a smartwatch, smart bracelet), or other device form factor, the structure of the electronic device may include more advanced features. Figure 3 The fewer structures shown can also include more than Figure 3 The structures shown are not limited here.
[0135] It is understandable that, generally speaking, the implementation of electronic device functions requires not only hardware support but also software cooperation. The software system of electronic devices can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application's embodiment uses a layered architecture... Taking the system as an example, the software structure of the electronic device is illustrated.
[0136] Figure 4 This is a schematic diagram of the layered architecture of the software system of the electronic device provided in the embodiments of this application.
[0137] Layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces (such as APIs).
[0138] In some examples, refer to Figure 4 As shown in this embodiment, the software of the electronic device is divided into five layers, from top to bottom: the application layer, the framework layer (or application framework layer), the system library and Android runtime, the HAL layer (hardware abstraction layer), and the driver layer (or kernel layer). The system library and Android runtime can also be referred to as the native framework layer or the native layer.
[0139] The application layer can include a series of applications. For example... Figure 4 As shown, the application layer can include applications (APPs) such as camera, gallery, calendar, map, WLAN, Bluetooth, music, SMS, and calls.
[0140] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer.
[0141] The application framework layer includes some predefined functions or services. For example, the application framework layer may include an activity manager, window manager, content provider, audio service, view system, phone manager, resource manager, notification manager, package manager, and system service manager, etc., and this application embodiment does not impose any limitations on this.
[0142] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0143] Content providers store and retrieve data, making that data accessible to applications. This data can include videos, images, audio, phone calls made and received, browsing history and bookmarks, phone books, etc.
[0144] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views.
[0145] A phone manager is used to provide communication functions for electronic devices.
[0146] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0147] The notification manager allows applications to display notification information in the status bar. It can be used to convey informational messages and can disappear automatically after a short time without user interaction.
[0148] Package manager in The system is used to manage application packages. It allows applications to obtain detailed information about installed applications and their services, permissions, etc. The package manager is also used to manage events such as application installation, uninstallation, and upgrades. The system service manager is responsible for managing and running system services to provide various system functions to applications and system components. In this embodiment, it can be used to receive user operation instructions for applications at the application layer.
[0149] The system library can include multiple functional modules, such as the surface manager and media libraries. The surface manager manages the display subsystem and provides 2D and 3D layer blending for multiple applications. The media libraries support playback and recording of various common audio and video formats, as well as still image files. The media libraries support multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0150] The Hardware Abstraction Layer (HAL) is an interface layer located between the operating system kernel and the hardware circuitry. Its purpose is to abstract the hardware. It hides the hardware interface details of a specific platform, providing the operating system with a virtual hardware platform that is hardware independent and portable across multiple platforms.
[0151] The driver layer is the layer between hardware and software. The driver layer includes at least display drivers, camera module drivers, audio drivers, sensor drivers, battery drivers, etc., but this application is not limited to these. Specifically, the sensor driver can include the driver for each sensor included in the electronic device, such as an ambient light sensor driver. For example, the ambient light sensor driver can, in response to an indication or instruction from the sensor module to acquire detection data, promptly send the detection data from the ambient light sensor to the sensing module. In this embodiment, the camera module driver and display driver are used to capture images and acquire raw images.
[0152] exist In this system, applications are typically written in Java. Each application can include one or more class files. Each application can run its own class files as a process within the application layer. When a user interacts with an application, the application can call the relevant application programming interface (API) or service in the application framework layer to interact with system libraries or the kernel layer and implement the functionality corresponding to the user's actions.
[0153] The following is combined Figure 5 The flow of the image processing method provided in the embodiments of this application will be described.
[0154] Figure 5 This is a schematic flowchart of an image processing method provided in an embodiment of this application.
[0155] Combination Figure 5 As shown, this embodiment includes the following steps S501-S505:
[0156] In step S501, the electronic device responds to the user-triggered start operation and displays the first interface.
[0157] Optionally, the initial operation can be performed by the user in any application. For example, the electronic device responds to user actions such as clicking or typing in a social application.
[0158] The first interface is the interface currently displayed on the electronic device, which can be the interface displayed by any application.
[0159] In step S502, the electronic device responds to the second operation triggered by the user by detecting whether the original image exists on the first interface.
[0160] Alternatively, the electronic device can respond to user actions in any application to acquire the raw image to be processed.
[0161] Optionally, the second operation may include a long press or a voice activation operation, as described below. Figure 6 The second operation is illustrated by example.
[0162] Figure 6 This is a schematic diagram illustrating how a second operation is triggered on a first interface, as provided in an embodiment of this application.
[0163] In one example, combined Figure 6 As shown in (a) in the figure, Figure 6 (a) in the diagram represents a first interface 60 displayed by a social application, which includes a bottom toolbar 601. In response to a user's long-press operation 06 on the bottom toolbar 601 within the first interface 60, the electronic device performs content recognition on the first interface 60 to detect the presence of an original image 602.
[0164] In another example, combined Figure 6 As shown in (b), in response to a user’s voice activation operation on the first interface 60, the electronic device performs content recognition on the first interface 60 and detects whether the original image 602 exists.
[0165] Optionally, the second operation may also include double-clicking the power button or long-pressing the interface with two fingers, such as for 3 seconds; this is not a limiting factor. Optionally, the original image 602 typically has a specific format, such as RAW, DNG, JPG, GIF, NEF, etc. The electronic device can detect the resource files loaded on the first interface 60, check whether the image file format exists in the resource files, and thus determine whether the first interface 60 contains the original image 602.
[0166] For example, continue to combine Figure 6 As shown in (a), the electronic device detects the resource files loaded on the first interface 60, checks for image file formats within the resource files, and then determines that the first interface 60 includes the original image 602 that needs to be processed. It should be noted that the method by which the electronic device determines whether the current interface includes the original image 602 is not specifically limited; it can be implemented using the method described above, or by using functions such as virtual assistants, intelligent agents, virtual vision, or service recommendations to obtain the original image.
[0167] In step S503, if the original image exists in the first interface, the electronic device displays a first interface including displaying a first control corresponding to the virtual assistant. Optionally, the first control may include at least one of an input control and at least one selection control.
[0168] For example, the selection control can provide recommendations for services supported by the virtual assistant, as well as recommendations for services supported by third-party services. In this scenario, the electronic device can obtain the original image of the current first interface through system controls, or it can obtain the original image of the current first interface through a screenshot.
[0169] Alternatively, system controls are designed based on function calls to local resources or network downloads.
[0170] System controls can access locally stored resources, such as raw images, through predefined functions. System controls can also download resources, such as raw images, from remote servers using network interfaces.
[0171] Specifically, Table 1 provides an example. Table 1 shows the whitelist stored by the electronic device, which includes scenarios, applications, activity names, preset conditions, and extraction targets.
[0172] Table 1
[0173]
[0174] For example, the electronic device can obtain the activity corresponding to the current interface application, and further determine whether the current activity is in a preset whitelist (Table 1). If it is in Table 1, the corresponding original image address is obtained through the preset function (such as getImageResourceId) of the image control (such as ImageView control). If the original image address is a local address, the original image is obtained through the local interface; if the original image address is not a local address, the original image is obtained through the network interface. If it is not in the whitelist, the original image can also be obtained through user interaction (such as taking a screenshot).
[0175] Optionally, the input control can provide a virtual assistant dialogue mode supported by the virtual assistant, enabling the user to realize their intentions through interaction. In this scenario, the electronic device can obtain the original image selected by the user through user-to-user interaction, or by dragging the original image into the virtual assistant dialogue interface.
[0176] After detecting the presence of the original image 602 in the first interface 60, the electronic device updates the first interface 60, generates a first interface 61 including the first control corresponding to the virtual assistant, and displays it.
[0177] It should be noted that the first interface 61 can be the interface of any application, such as a camera application. Electronic devices can capture a frame from the preview stream as the original image, and no specific limitation is made here.
[0178] The following is combined Figure 7The first interface 61 is described by way of example.
[0179] Figure 7 This is a schematic diagram of a first interface provided in an embodiment of this application.
[0180] In one example, combined Figure 7 As shown in (a), when the electronic device detects the presence of an original image 601 in the first interface 60, it updates the first interface 60 and displays the first interface 61a. The first interface 61a includes an input control 611.
[0181] In another example, combined Figure 7 As shown in (b), when the electronic device detects the presence of the original image 601 in the first interface 60, it updates the first interface 60 and displays the first interface 61b. The first interface 61b includes at least one selection control 612.
[0182] In yet another example, combining Figure 7 As shown in (c), when the electronic device detects the presence of an original image 601 in the first interface 60, it updates the first interface 60 to obtain a first interface 61c, which includes at least one selection control 612 and an input control 611.
[0183] In step S504, the electronic device responds to the user's first operation on the first control by using an image processing model to process the original image in accordance with the first operation, thereby generating a first target image.
[0184] Optionally, the first operation may include a click operation or an input operation.
[0185] Clicking refers to a user briefly pressing and releasing a control on the interface; it is a way for users to interact with the interface.
[0186] For example, continue to combine Figure 7 As shown in (c), the user clicks on the selection control 612 in the first interface 61c.
[0187] Input operations refer to an interaction method in which users provide data to electronic devices through physical or voice means.
[0188] Input operations can include text input operations or voice input operations.
[0189] Text input refers to users entering text, numbers, symbols, etc., using a keyboard, touchscreen, stylus, etc. Voice input refers to users interacting with electronic devices using voice commands through a microphone or other voice input devices.
[0190] For example, continue to combine Figure 7As shown in (c), the user performs text input operations on the input control 611 in the first interface 61c, such as inputting "remove passersby".
[0191] Optionally, the electronic device responds to a second operation performed by the user on a first control on the first interface of the electronic device, thereby triggering a call to an image processing model, which is used to process the original image to obtain a first target image.
[0192] Optionally, the image processing model can be set on an electronic device or in the cloud, without any specific limitation.
[0193] Step S505: Update the original image in the first interface to the first target image to obtain the second interface.
[0194] like Figure 8A As shown in (a), the first interface 61a includes an input control 611.
[0195] In one example, the electronic device, in response to a user's text input operation (such as "remove passersby") on the input control 611 in the first interface 61a, jumps to the display interface 62 (e.g., ... Figure 8A As shown in (b) in the figure, the display interface 62 includes a first target image.
[0196] In another example, the electronic device, in response to a user's text input operation (such as "remove passersby") on the input control 611 in the first interface 61a, jumps to the virtual assistant dialog interface 63 (e.g., ... Figure 8A As shown in (c) above, the virtual assistant dialogue interface 63 includes a text prompt box. This text prompt box is used to inform the user that the original image is being processed and to ask them to wait. Further, when the original image processing is complete, the user is redirected to the completion interface 64 (as shown in [example interface]). Figure 8A As shown in (d) in the diagram, the completed interface 64 includes a thumbnail 641 corresponding to the first target image. In response to a user's click operation 07 on the thumbnail 641 corresponding to the first target image, the electronic device jumps to the display interface 62, which includes the first target image. The display interface 62 may include share controls, save controls, and reset controls. The electronic device can perform further processing on the first target image based on these controls; specific limitations are not specified here.
[0197] It should be noted that the above exemplary description of obtaining the first target image through the first interface 61a can also be described using the first interface 61b and the first interface 61c, which will not be repeated here.
[0198] In another example, such as Figure 6 and Figure 8B As shown, when a user activates the voice assistant, they enter the first interactive interface 65 (e.g., ...). Figure 8B As shown in (a)), the first interactive interface 65 includes an input control 611, which includes an image add button 6110. In response to a user's click on the image add button 6110 in the first interactive interface 65, the electronic device jumps to the image interface to obtain the original image (not shown in the figure), and then jumps to the first interactive interface 66 (as shown in (a)). Figure 8B As shown in (b)), the first interactive interface 66 includes the original image to be sent. In response to a user's click operation on the control 611 input in the first interactive interface 66, the electronic device jumps to the first interactive interface 67 (as shown in [example]). Figure 8B As shown in (c), the first interactive interface 67 waits for the removal of passersby from the original image. After the processing is completed, it jumps to the first interactive interface 68 (as shown in (c)). Figure 8B As shown in (d) in the figure, the first interactive interface 68 includes a thumbnail 641 corresponding to the first target image. In response to a user's click operation 08 on the thumbnail 641 corresponding to the first target image, the electronic device jumps to the display interface 62 (as shown in the figure). Figure 8A (b) of the above, the display interface 62 includes the first target image.
[0199] It should be noted that the first interactive interface 66 may also include a selection control 612 located below the original image, such as "Remove Passersby". In response to the user's click operation on the first interactive interface 66 to select control 612, the electronic device jumps to the first interactive interface 67. On the first interactive interface 67, the process of removing passersby from the original image is waited for. After processing is completed, the device jumps to the first interactive interface 68, which includes a thumbnail 641 corresponding to the first target image. In response to the user's click operation 07 on the thumbnail 641 corresponding to the first target image, the electronic device jumps to the display interface 62 (e.g., ...). Figure 8A (b) of the above, the display interface 62 includes the first target image.
[0200] In another example, such as Figure 8B and Figure 8C As shown, when a user activates the voice assistant, they enter the first interactive interface 65 (e.g., ...). Figure 8B As shown in (a) above, the first interactive interface 65 includes selection buttons, such as an agent button. In response to a user's click on the agent button in the first interactive interface 65, the electronic device jumps to the first agent interface 69 (e.g., ...). Figure 8C As shown in (a)), the first intelligent agent interface 69 includes multiple settings, such as "AI Edit Intelligent Agent" 691. In response to a user's click on "AI Edit Intelligent Agent" 691 in the first intelligent agent interface 69, the electronic device jumps to the second intelligent agent interface 70 (e.g., ...). Figure 8CAs shown in (b)), the second intelligent agent interface 70 includes scene buttons, such as a pedestrian removal button 701. In response to the user's click on the pedestrian removal button 701 in the second intelligent agent interface 70, the electronic device jumps to the third intelligent agent interface 71 (e.g., ...). Figure 8C (as shown in (c)). In response to a user's click on the image add button 6110 in the input control 611, the electronic device jumps to the image interface to obtain the original image (not shown in the figure), and then jumps to the fourth intelligent agent interface 72 (as shown in (c)). Figure 8C As shown in (d) in the diagram, the original image sent is included in the fourth agent interface 72. In response to a user's click operation on the control 611 input in the fourth agent interface 72, the electronic device jumps to the fifth agent interface 73 (as shown in the diagram). Figure 8C As shown in (e), the fifth agent interface 73 waits for the original image to be processed by one-click image cutout. After the processing is completed, it jumps to the sixth agent interface 74 (as shown in (e)). Figure 8C As shown in (f) in the diagram, the sixth intelligent agent interface 74 includes a thumbnail 741 corresponding to the first target image. In response to a user's click on the thumbnail 741 corresponding to the first target image, the electronic device jumps to the display interface 62 (as shown in (f)). Figure 8A (b) of the above, the display interface 62 includes the first target image.
[0201] It should be noted that the electronic device can directly acquire the original image in response to the user's click operation on the image addition button 6110 in the input control 611 in the second intelligent agent interface 70, so as to subsequently acquire the first target image. No specific example is given here.
[0202] In summary, electronic devices can invoke virtual assistants within any application to process raw images, improving image processing efficiency and ensuring a better user experience.
[0203] When the electronic device displays the first interface 61a, step S504 will be explained by way of example in combination with the specific scenario.
[0204] In one embodiment, content extraction can be performed on the original image to obtain the first target image, as shown below with reference to Figure 9 and... Figure 10 The processing method shown in the figure will be used to explain this embodiment.
[0205] Figure 9A This is a schematic diagram of a process for content extraction from an image provided in an embodiment of this application.
[0206] Combination Figure 9A As shown, this embodiment includes the following steps S901-S903:
[0207] In step S901, the electronic device responds to the user's second operation on the input control 611 in the first interface 61a and obtains instruction information.
[0208] The second operation will be explained using text input as an example. For instance, the user enters "cut out the long-haired girl" in input control 611.
[0209] It should be noted that the above instructions can be input based on the features of the person in the original image, or based on the position of the object in the original image. For example, a user can input "cut out the girl wearing sunglasses on the left" in the input control 611.
[0210] In addition, combined Figure 9B As shown, Figure 9B This is a schematic diagram illustrating an embodiment of obtaining instruction information provided in this application.
[0211] Combination Figure 9B As shown, the electronic device can input the original image into the instance segmentation model to obtain multiple objects segmented from the original image, and label each object as the basis for providing input instructions to the user.
[0212] For example, the instruction information entered by the user in the input control 611 is "remove the person with the number 5".
[0213] Alternatively, electronic devices can use natural language processing (NLP) to extract user intent from instruction information and classify that intent into specific operations.
[0214] For example, if an electronic device fails to extract the user's intent on the first attempt, it may further confirm the user's intent through multiple conversations. For instance, it might send the user a message such as, "The image includes two long-haired girls; please describe them in detail," to further confirm the user's intent.
[0215] Furthermore, the electronic device extracts slot information from the instruction information to determine detailed information related to instruction execution. Slot extraction may include aspects such as object type, feature description, and operation type.
[0216] Taking the instruction "cut out the girl with long hair" as an example, the slots are extracted, and the slots corresponding to the object type are "girl", the feature description is "long hair", and the operation type is "cut out".
[0217] Optionally, after determining the type of task the instruction information is used to indicate, the electronic device transmits the user's instruction information to the appropriate processing module or service for processing.
[0218] For example, given the instruction "cut out the girl with long hair", the electronic device can determine that the current scene is a "cutout" scene.
[0219] In step S902, the electronic device uses the visual grounding model 12 to process the original image in accordance with the instruction information to generate the first main body region.
[0220] Optionally, the visual localization model 12 is obtained by fine-tuning the instructions of the Intern Visual-Language Model 2 (InternVL2). It should be noted that the visual localization model can also be obtained by fine-tuning other algorithms, such as the Visual-Language Model (VisualBERT), the Multi-Modal Object Detection Model (Multi-ModalDETR, MDETR), the pre-trained model (Grounded Language Image Pre-training, GLIP), etc., without being specifically limited here.
[0221] Specifically, step S902 above can be implemented by the following method:
[0222] Figure 10 This is a schematic diagram of the structure of a visual positioning model provided in an embodiment of this application.
[0223] Combination Figure 10 As shown, the original image and instruction information are input into the visual positioning model 12 to obtain the output image. The electronic device crops the bounding box of the first subject in the output image to obtain the first subject region.
[0224] The output image includes the bounding box of the first subject, and the first subject region is the area enclosed by the straight bounding box of the edge of the first target subject.
[0225] For example, the process of using the visual positioning model 12 to process the original image in accordance with the instruction information may include the following steps:
[0226] The first step is to input the instruction information into the text encoder 121 for feature extraction to obtain instruction features.
[0227] For example, inputting the phrase "cut out the long-haired girl" into a Generative Pre-trained Transformer (GPT) generates instruction features.
[0228] The second step is to input the original image into the image encoder 122 for feature extraction to obtain at least one image feature.
[0229] For example, the original image is input into a Vision Transformer (ViT) for feature extraction, resulting in the first image feature, the second image feature, the third image feature, and so on.
[0230] Third, the electronic device matches the instruction features with at least one image feature to determine an output image that includes a bounding box of the first subject.
[0231] In one implementation, the electronic device first determines a first image feature among at least one image feature based on instruction features and at least one image feature.
[0232] The electronic device determines the similarity between the instruction feature and each image feature. If the similarity meets a similarity threshold, it determines the first image feature corresponding to the instruction feature. In other words, the closer the similarity is to the similarity threshold, the better the match between the instruction feature and the image feature.
[0233] For example, if the similarity threshold is set to 1, the similarity between the instruction feature and the first image feature is 0.5, the similarity between the instruction feature and the second image feature is 0.9, and the similarity between the instruction feature and the third image feature is 0.3. The instruction feature matches the second image feature, and further, the electronic device determines that the second image feature is the first image feature.
[0234] Alternatively, the electronic device may use cosine similarity or Euclidean distance to calculate the similarity between instruction features and each image feature.
[0235] Furthermore, the electronic device determines an output image that includes a bounding box of the first target subject based on the first image features.
[0236] The bounding box is a rectangular area used to mark the location of the first subject in the image.
[0237] In another implementation, the electronic device can input input features and image features into a contrastive learning network to obtain an input image that includes a bounding box of a first target subject.
[0238] For example, contrastive learning networks mainly rely on self-attention mechanisms, contrastive learning, cross-modal alignment, generative adversarial networks and other methods to align image features with text features and obtain the output image of the bounding box of the first subject.
[0239] The following is combined Figure 11 The cropping of the bounding box of the first subject in the output image is illustrated by way of example.
[0240] Figure 11 This is a schematic diagram of an image cropping interface provided in an embodiment of this application.
[0241] like Figure 11 As shown in (a) in the figure, Figure 11 (a) is the output image including the bounding box of the first subject. Figure 11 As shown in (b), the output image is input into the object detection model to determine the coordinates of the bounding box of the first subject, such as... Figure 11 The intermediate image shown in (b) has the coordinates of point A, point B, point C, and point D. Further, as... Figure 11 As shown in (c), the first main body region is obtained by cropping based on the straight line area enclosed by these coordinate points (e.g., ...). Figure 11 (as shown in (d)).
[0242] For example, if the input image size is 800 pixels × 600 pixels, the output image is input into the object detection model, which identifies the coordinates of the top-left corner A (300, 300), the bottom-left corner B (300, 0), the top-right corner C (500, 300), and the bottom-right corner D (500, 0). Based on the straight-line region enclosed by the coordinates of points A, B, C, and D, a first main body region is obtained (e.g., ...). Figure 11 (as shown in (a)).
[0243] Optionally, the first main area can be an RGB image.
[0244] In step S903, the electronic device processes the first subject region using the first image extraction model to generate a first target image having at least one first target subject.
[0245] The first image extraction model is used to extract the first subject region from the original image.
[0246] Optionally, the first image extraction model can use a convolutional neural network (CNN), or other models, without specific restrictions.
[0247] For example, continue to combine Figure 11 As shown, Figure 11 (d) in the text represents the first main region (e.g., Figure 11 The first target image obtained after extraction by the first image extraction model (as shown in (c)).
[0248] For example, the first image extraction model is based on a matting algorithm.
[0249] Optionally, since the first subject area is an RGB image, meaning it only contains three color channels: red, green, and blue, it's necessary to make the background transparent in some image scenarios, such as "image cutout scenarios." For example, when embedding a subject into another background image, a transparent background allows the subject to blend perfectly with the new background.
[0250] Continue to combine Figure 11 As shown in (c), if the target subject in the first subject area is directly separated from the background, the edges of the subject may appear too sharp and abrupt. In "cutout" scenarios, the edges are often the most prone to artifacts or imperfections. Therefore, it is necessary to further apply a gradient processing to the edges of the subject.
[0251] For example, the first step is to obtain the main body region within the first target region. For instance, electronic devices can extract the main body region from an image using methods such as Canny edge detection or the Laplacian operator. For example, a binary mask is used to represent the main body region and the background region, with the pixel value of the main body region set to 1 and the pixel value of the background region set to 0.
[0252] The second step is to expand the edges of the main body region to obtain the first edge region. For example, electronic devices can use distance transformation to obtain the expanded edge region, i.e., the first edge region, to enhance the boundary of the main body region.
[0253] The third step is to perform a gradient processing on the first edge region to obtain the gradient result corresponding to the Alpha channel.
[0254] The alpha channel is used to control the transparency of an image. In addition to the common RGB (red, green, blue) color channels, the alpha channel is introduced as a fourth channel, specifically to describe the transparency (or opacity) of each pixel in the image. Together with the RGB channels, the image forms an RGBA format.
[0255] In other words, after processing the first edge region with the Alpha channel, the main body region gradually transitions to the background region from opaque to transparent.
[0256] For example, the pixel value of the main body region of the first main body region is set to 1, representing complete opacity, the pixel value of the background region of the first main body region is set to 0, representing complete transparency, and the first edge region is the boundary region between the main body region and the background region, with the region value of the first edge region being [0, 1].
[0257] The fourth step is to obtain the first target image based on the gradient result and the first subject region.
[0258] Optionally, the first target image is an RGBA image, which visually presents a smooth transition between the subject and the background.
[0259] It should be noted that the first main area can also be the area obtained by content recognition and other processing of objects and text in the original image, which will not be shown one by one here.
[0260] In summary, electronic devices can invoke virtual assistants in any application, allowing users to extract content from raw images through input controls, thereby improving image processing efficiency and ensuring a better user experience.
[0261] In another embodiment, the original image can be cleaned of its background to obtain the first target image, as described below. Figure 12 and Figure 13 The processing method shown in the figure will be used to explain this embodiment.
[0262] Figure 12 This is a schematic diagram of a process for background purification of an original image provided in an embodiment of this application.
[0263] Figure 13 This is a schematic diagram of a structure for background purification of an original image provided in an embodiment of this application.
[0264] Combination Figure 12 and Figure 13 As shown, this embodiment includes the following steps S121-S125:
[0265] In step S121, the electronic device responds to the user's second operation on the input control 611 in the first interface 61a and obtains instruction information.
[0266] The specific content of step S121 can be referred to step S901 above, and is not specifically limited here.
[0267] In step S122, when the amount of information in the instruction information meets the first preset threshold, the electronic device uses the visual positioning model 12 to process the original image in accordance with the instruction information to generate an output image of the bounding box of the second target subject.
[0268] Optionally, the first preset threshold is an information content threshold.
[0269] For example, the first preset threshold can be set to include at least three key elements, which may be the target object, the operation type, and a detailed feature description. For instance, if the instruction information is "remove passersby", the information content of the instruction information does not meet the first preset threshold; if the instruction information is "remove the girl with her hair tied up and wearing long sleeves in the image", the information content of the instruction information meets the first preset threshold.
[0270] It's important to note that image processing models are generalizable. They enhance their generalization ability by learning from a wide range of scenes and expressions, ensuring correct task execution across various image types and expressions. For example, the instruction "remove passersby" is not a fixed instruction. If the user expresses the instruction "remove passersby" in different ways, the model should have sufficient semantic understanding to execute the task. Similarly, expressions like "remove background people" or "remove distracting elements" should trigger the same operation. Figure 14 This is another schematic diagram of the structure of a visual positioning model provided in an embodiment of this application.
[0271] like Figure 14 As shown, when the amount of information in the instruction information meets the first preset threshold, the process of using a visual positioning model to process the original image in accordance with the instruction information may include the following steps:
[0272] The first step is to input the instruction information into the text encoder 121 for feature extraction when the amount of information in the instruction information meets the first preset threshold, so as to obtain the instruction features.
[0273] For example, the first preset threshold is set to the target object, operation type, and detailed feature description. If the instruction information is "Remove a girl with her hair tied up and wearing long sleeves from the image," the information content of this instruction information includes the target object: girl, operation type: removal, and detailed feature description: tied-up hair and wearing long sleeves. The information content of this instruction information meets the first preset threshold. When the information content of the instruction information meets the first preset threshold, "Remove a girl with her hair tied up and wearing long sleeves from the image" is input into the Generative Pre-trained Transformer (GPT) to generate instruction features.
[0274] The second step is to input the original image into the image encoder 122 for feature extraction to obtain at least one image feature.
[0275] For example, the original image is input into the image encoder 122 for feature extraction, resulting in the fourth image feature, the fifth image feature, the sixth image feature, and so on.
[0276] The second step above can be referred to the above. Figure 10 The specific details of the second step in the embodiment shown will not be repeated here.
[0277] The third step involves inputting the instruction features and at least one image feature into the feature matching module 123 to obtain an output image that includes the bounding box of the second target subject.
[0278] Optionally, the third step may include the first sub-step and the second sub-step.
[0279] In the first sub-step, the electronic device determines a second image feature from at least one image feature based on instruction features and at least one image feature.
[0280] In one implementation, if the instruction information is as in the example above, "remove the girl with her hair tied up and wearing long sleeves from the image", the electronic device determines the similarity between the instruction feature and each image feature. If the similarity meets the similarity threshold, the electronic device determines the second image feature corresponding to the instruction feature.
[0281] The above example can be referred to above. Figure 10 The specific details of the third step in the illustrated embodiment will not be repeated here.
[0282] In another implementation, if the instruction information includes ambiguous pronouns, the electronic device needs to attempt to understand the object referred to by the pronoun in order to perform the corresponding task. For example, "Remove my cousin from the image".
[0283] Alternatively, the electronic device may determine a second image feature from at least one image feature based on an image clustering library and instruction features.
[0284] Figure 15 This is a schematic diagram illustrating how to determine a second image feature according to an embodiment of this application. The following is a detailed explanation. Figure 15 An illustrative example is provided.
[0285] Combination Figure 15 As shown, firstly, after obtaining the instruction features of the instruction information, the image clustering library is used to determine whether it includes preset image features corresponding to the instruction features.
[0286] The image clustering library includes the correspondence between instruction features and preset image features.
[0287] For example, if the image clustering library includes a first preset image feature and a first preset instruction feature corresponding to the first preset image feature, a second preset image feature and a second preset instruction feature corresponding to the second preset image feature, a third preset image feature and a third preset instruction feature corresponding to the third preset image feature, etc.
[0288] Furthermore, if the image clustering library includes preset image features corresponding to instruction features, the preset image features are matched with at least one image feature to determine a second image feature among the at least one image features.
[0289] For example, based on the instruction information "remove the cousin from the image", the corresponding instruction feature is obtained. The electronic device searches for a preset instruction feature that is the same as the instruction feature in the image clustering library, and then searches for the preset image feature corresponding to the preset instruction feature. If the preset image feature that matches the instruction feature is the first preset image feature, the first preset image feature is matched with multiple image features. If the fourth image feature matches the first preset image feature, then the fourth image feature is determined to be the second image feature.
[0290] It should be noted that, in addition to the image clustering library mentioned above acquiring the objects referred to by ambiguous pronouns, electronic devices can learn models through neural networks to identify the objects referred to by ambiguous pronouns in different scenarios and perform corresponding learning. For example, in a chat scenario, the electronic device learns, by analyzing the user's dialogue, that "cousin" refers to a girl (e.g.,...). Figure 15 (The girl on the right in the original image).
[0291] Figure 16 This is another schematic diagram of determining a second image feature provided in the embodiments of this application, which is described below in conjunction with... Figure 16 An illustrative example is provided.
[0292] Combination Figure 16 As shown, after obtaining the instruction features of the instruction information, the image clustering library is used to determine whether it includes preset image features corresponding to the instruction features.
[0293] The above content can be referred to above. Figure 15 The contents of the corresponding embodiments will not be repeated here.
[0294] Furthermore, if the image clustering library does not include preset image features corresponding to the instruction features, a third interface is displayed.
[0295] Optionally, the third interface includes at least one image control corresponding to an image feature.
[0296] For example, continue to combine Figure 16 As shown, Figure 16 (a) refers to the third interface 61a. Entering "remove cousin from image" in the input control of the first interface will redirect to... Figure 16 (b) in the text refers to the third interface, which includes multiple image controls, such as a first image control, a second image control, and a third image control 161.
[0297] Optionally, the image control can be multiple face recognition boxes.
[0298] The electronic device, in response to a third operation by a user on an image control on a third interface, determines a second image feature from at least one image feature.
[0299] Optionally, the third action includes a click action.
[0300] Based on the above example, the electronic device responds to the user's click operation 091 on the third image control 161 and determines that the image feature corresponding to the third image control is the second image feature.
[0301] The second sub-step involves determining an output image that includes the bounding box of the second target subject, based on the second image features.
[0302] The second sub-step described above can refer to the above-described step embodiment to determine the specific content of the output image including the bounding box of the first target subject, which will not be repeated here.
[0303] Optionally, the bounding box of the second target body can be a curved bounding box or a straight bounding box.
[0304] Step S123: The output image is processed using the second image extraction model 13 to generate the second target image.
[0305] Optionally, the second image extraction model extracts the image region in the output image excluding the bounding box of the second subject.
[0306] Continue to combine Figure 13 For example, the bounding box of the second target in the output image is the curved bounding box of the second target.
[0307] Optionally, the output image including the bounding box of the second target subject curve is input into the second image extraction model 13 for extraction to obtain the second target image.
[0308] Step S124: The image inpainting model 14 is used to process the second target image and the original image to generate a first target image that does not have at least one second target subject.
[0309] Optionally, continue to combine Figure 13 As shown, the image inpainting model 14 is used to repair the blank area in the second target image after the removal of the second target subject, generating a first target image that does not have at least one second target subject.
[0310] In summary, electronic devices can invoke virtual assistants in any application, allowing users to perform background purification and scene processing on original images through input controls, thereby improving image processing efficiency and ensuring a better user experience.
[0311] In yet another embodiment, the original image can be cleaned of its background to obtain the first target image, which will be described below in conjunction with... Figure 17 and Figure 18 The processing method shown in the figure will be used to explain this embodiment.
[0312] Figure 17 This is another schematic diagram of a process for background purification of an image provided in an embodiment of this application.
[0313] Figure 18 This is another schematic diagram of a structure for background purification of an image provided in an embodiment of this application.
[0314] Combination Figure 17 and Figure 18 As shown, this embodiment includes the following steps S1701-S1703.
[0315] In step S1701, the electronic device responds to the user's second operation on the input control 611 in the first interface 61a and obtains instruction information.
[0316] The specific content of step S1701 can be referred to step S121 above, and is not specifically limited here.
[0317] In step S1702, if the amount of information in the instruction information does not meet the first preset threshold, the second image extraction model 13 is used to process the original image in accordance with the instruction information to generate the second target image.
[0318] The specific content of step S1702 can be referred to step S131 above, and is not specifically limited here.
[0319] For example, the instruction message is "Remove the girl from the image". The information content of this instruction message includes the target object: girl, operation type: removal, and detailed feature description: none. The information content of this instruction message does not meet the first preset threshold.
[0320] Step S1703: The second target image and the original image are processed using an image restoration model to generate a first target image that does not have at least one second target subject.
[0321] The specific content of step S1703 can be referred to step S124 above, and is not specifically limited here.
[0322] It should be noted that if the amount of information in the instruction does not meet the first preset threshold, the electronic device may also use a multi-turn dialogue method to further determine the object to be removed, without making specific limitations here.
[0323] In summary, electronic devices can invoke virtual assistants in any application, allowing users to perform background purification and scene processing on original images through input controls, thereby improving image processing efficiency and ensuring a better user experience.
[0324] Optionally, in embodiments of background purification scenarios, combined with Figure 19 This is an exemplary description of displaying the first interface 61a.
[0325] Figure 19 This is a flowchart illustrating an interface for displaying a first control corresponding to a virtual assistant, provided in an embodiment of this application.
[0326] like Figure 19 As shown, the steps for displaying interface 61a, which includes the first control corresponding to the virtual assistant, include:
[0327] Step S191: Based on the original image, determine the target subject bounding box corresponding to at least one first target subject and the non-target subject bounding box corresponding to at least one second target subject in the original image.
[0328] Optionally, step S191 may include steps S911-S912.
[0329] Step S911: The original image is processed using a subject recognition model to obtain a first processed image including a target subject bounding box corresponding to at least one first target subject.
[0330] Alternatively, the subject recognition model can employ a deep learning-based object detection model.
[0331] For example, the electronic device uses a deep learning-based object detection model to process the original image and identify the target bounding box corresponding to the first target subject in the original image.
[0332] The following is combined Figure 20 An illustrative explanation is provided for determining the target subject bounding box in the original image.
[0333] Figure 20 This is a schematic diagram of determining a target subject frame according to an embodiment of this application.
[0334] Combination Figure 20 As shown in (a), the target subject bounding box can be a single bounding box that includes multiple subjects. Combined with... Figure 20 As shown in (b) in the figure, the target subject bounding box can be multiple subject bounding boxes of multiple subjects.
[0335] Optionally, the electronic device can determine the target subject bounding box based on a preset subject bounding box threshold. When the number of target subjects is greater than the preset subject bounding box threshold, the target subject bounding box can be a single bounding box including multiple subjects; when the number of target subjects is less than the preset subject bounding box threshold, the target subject bounding box can be multiple subject bounding boxes of multiple subjects.
[0336] The preset edge threshold can be set to 5.
[0337] In summary, by merging multiple subjects into a single bounding box, the computational load of subsequent processing (such as feature extraction and classification) can be significantly reduced.
[0338] Step S912: The image is processed using an image instance segmentation model to obtain a second processed image that includes a non-target subject bounding box corresponding to at least one second target subject.
[0339] Alternatively, the image instance segmentation model can employ a region convolutional neural network (Mask R-CNN).
[0340] Step S192: If the number of target subject boxes meets the third preset threshold and the number of non-target subject boxes meets the third preset threshold, display the first control corresponding to the virtual assistant.
[0341] Optionally, the first interface includes a first control corresponding to the virtual assistant, and the first control includes an input control. For example, first interface 61a.
[0342] Specifically, step S192 above can be achieved by the following method:
[0343] Figure 21 This is a schematic diagram illustrating the determination of a target subject frame and a non-target subject frame according to an embodiment of this application.
[0344] Step S921: When the number of target subject frames is less than or equal to a third preset threshold, and the number of non-target subject frames is greater than or equal to a third preset threshold, the first interface 61b, including the first control corresponding to the virtual assistant, is displayed.
[0345] Optionally, the third preset threshold can be set to 0.
[0346] exist Figure 21 In the diagram, the target body outline can be represented by a rectangular dashed frame, while non-target body outlines can be represented by circular dashed frames. For example, as shown... Figure 21 As shown in (a), if the number of target body boxes is 0 and the number of non-target body boxes is 2, the first interface 61a is displayed.
[0347] In step S922, if the number of target subject boxes is greater than a third preset threshold and the number of non-target subject boxes is less than or equal to the third preset threshold, the first interface 61b, including the first control corresponding to the virtual assistant, is displayed.
[0348] For example, such as Figure 21 As shown in (b), if the number of target body boxes is 3 and the number of non-target body boxes is 0, the first interface is displayed.
[0349] In another embodiment, content extraction can be performed on the original image to obtain the first target image. This embodiment will be described below with reference to the processing method shown in Figure 22.
[0350] Figure 22A This is a schematic diagram of a structure for content extraction from an image, provided in an embodiment of this application.
[0351] Combination Figure 22A As shown, the electronic device responds to the user's second operation of selecting control 612 on the first interface 61b and obtains the control instruction.
[0352] Optionally, the second action is a click action.
[0353] Selection controls are controls that provide recommendation services to users in the first interface 61b. Examples include "one-click background removal," "portrait enhancement," and "removing passersby."
[0354] For example, combined Figure 22A As shown, Figure 23 (a) in the diagram represents a first interface 61b containing the original image 231. The first interface 61b includes selection controls such as "one-click cutout," "portrait enhancement," and "remove passersby." The electronic device responds to the user's click operation 08 on the first interface 61b for "one-click cutout" and obtains control instructions.
[0355] Alternatively, the electronic device can perform semantic understanding and recognition on the original image to predict the user's intended intent.
[0356] For example, Figure 22B This is a schematic diagram of a recommended selection control provided in an embodiment of this application.
[0357] Assuming the original image contains three objects, such as "a girl," "a trash can," and "a boy," the electronic device can generate recommended selection controls based on semantic understanding of the original image, such as "extract the girl wearing glasses" or "remove the trash can." Furthermore, the electronic device uses a first image extraction model 15 to process the original image according to the control instructions, generating a first target image.
[0358] Continue to combine Figure 22A As shown, the electronic device extracts the first target image by combining the original image with the control instructions input into the first image extraction model 15.
[0359] In summary, electronic devices can invoke virtual assistants in any application, allowing users to select controls to perform image cutout processing on the original image, thereby improving image processing efficiency and ensuring a better user experience.
[0360] Optionally, in the above embodiments of the content extraction scenario, combined with Figure 23 This is an exemplary description of displaying the first interface 61b.
[0361] Figure 23 This is another schematic flowchart illustrating a first interface for displaying a first control corresponding to a virtual assistant, provided in an embodiment of this application.
[0362] like Figure 23 As shown, the steps for displaying the first interface 61b, which includes the first control corresponding to the virtual assistant, include:
[0363] The original image is processed using the third image extraction model 16 to generate the third subject region in the original image.
[0364] Optionally, the third image extraction model is a lightweight extraction model. For example, the computational cost of the third image extraction model is less than that of the first image processing model.
[0365] The third main body region is the area enclosed by the curved boundary box of the edge of the first target main body.
[0366] For example, assume that the area of the third subject region is 30,000 pixels squared.
[0367] It should be noted that the third main body region of the generated original image can also be obtained using the first image extraction model 15.
[0368] Furthermore, if the third main area meets the fourth preset threshold, the first interface is displayed.
[0369] Optionally, the fourth presets a specific threshold standard to measure the proportion of a certain region in the entire image.
[0370] For example, the fourth preset threshold is 5%, but it can also be other percentages, which are not limited here.
[0371] For example, if the original image size is 800 pixels × 600 pixels, the area of the original image is 480,000 pixels squared. If the fourth preset threshold is 5%, then the area occupied by the fourth preset threshold is 24,000 pixels squared. If the area of the third main body region is 30,000 pixels squared, and the area of the third main body region is greater than or equal to the area occupied by the fourth preset threshold, then the first interface (such as...) is displayed. Figure 7 (As shown in 61b).
[0372] Optionally, the first interface includes a first control corresponding to the virtual assistant, and the first control includes a selection control.
[0373] In yet another embodiment, the original image can be cleaned of its background to obtain the first target image, which will be described below in conjunction with... Figure 24 The processing method shown in the figure will be used to explain this embodiment.
[0374] Figure 24 This is another structural diagram of a method for background purification of an original image provided in an embodiment of this application.
[0375] Combination Figure 24 As shown, the electronic device responds to the user's second operation of selecting control 612 on the first interface 61b and obtains the control instruction. Based on the control instruction and the original image, the image recognition model 17 processes the original image in accordance with the control instruction to generate the second target image.
[0376] The image recognition model may include a second image extraction model 13 and an image segmentation model 18.
[0377] Image segmentation models are used to divide images into different regions for further analysis or processing.
[0378] In one implementation, firstly, the electronic device uses an image segmentation model 18 to process the original image in accordance with the control instructions, generating a second subject region including the second target subject.
[0379] The first step above can be achieved through the following method:
[0380] The electronic device extracts text features from instruction controls and at least one image feature from the original image.
[0381] The above content can be referred to Figure 10 The contents of the corresponding embodiments will not be repeated here.
[0382] Optionally, the image features include image features corresponding to the third target subject.
[0383] For example, in combination Figure 25The second main area is obtained as an example.
[0384] Figure 25 This is a schematic diagram illustrating the acquisition of a second subject area according to an embodiment of this application.
[0385] like Figure 25 As shown in (a), the original image includes three main objects: object A, object B, and object C. Assume that the image features extracted from the original image are A1, B1, and C1, respectively. Assume that the image feature of the second target object is B1.
[0386] Furthermore, the method for achieving the first step also includes: obtaining the second subject region in the original image corresponding to the instruction features based on the correlation between each image feature and the second target subject image features.
[0387] Relationships refer to the interactions, dependencies, and connections among multiple objects in an image.
[0388] Optionally, the relationships between main objects can be determined based on visual features of the original image, such as the pose and angle of each main object in the original image. Alternatively, the relationships between objects can be determined by reading the features of people in the original image and comparing them with data pre-stored in the electronic device.
[0389] Continue to combine Figure 25 As shown, in Figure 25 In (a), object A is moving toward object B. If object B, corresponding to the image feature of B1, is the third target subject, then object A and object B are related.
[0390] The second step is to determine the second main body region based on the image features (A1 image features, B1 image features) corresponding to object A and object B.
[0391] For example, continue to combine Figure 25 As shown, in Figure 25 In (b) of the above, the second main area includes object A and object B.
[0392] The second step is to use the second extraction model to process the second main body region and determine the second target image.
[0393] The third step mentioned above can refer to the specific content of step S123 above, and will not be specifically limited here.
[0394] Continue to combine Figure 24 As shown, the electronic device processes the second target image and the original image using the image inpainting model 14 to generate a first target image that does not have at least one second target subject.
[0395] The above steps can refer to the content of step S124 above, and are not specifically limited here.
[0396] In summary, electronic devices can invoke virtual assistants in any application, allowing users to remove passersby from original images through input controls, thereby improving image processing efficiency and ensuring a better user experience.
[0397] Optionally, in the above embodiments of the content extraction scenario, combined with Figure 26 This is an exemplary description of displaying the first interface 61b.
[0398] Figure 26 This is another flowchart illustrating the display of a first interface provided in an embodiment of this application.
[0399] like Figure 26 As shown, the steps for displaying the first interface 61b include:
[0400] Step S261: Based on the original image, determine the target subject bounding box corresponding to at least one first target subject and the non-target subject bounding box corresponding to at least one second target subject in the original image.
[0401] Optionally, step S261 may include steps S6101-S6102.
[0402] In step S6101, the electronic device processes the original image using a subject recognition model to generate a third target image. The third target image includes the subject bounding box corresponding to the first target subject.
[0403] The above content can be referred to in Figure 9 and the corresponding embodiments, and will not be repeated here.
[0404] Step S6102: Process the third target image using the instance segmentation model to generate the fourth target image.
[0405] The fourth target image includes at least one target subject bounding box corresponding to the first target subject, and at least one non-target subject bounding box corresponding to the second target subject.
[0406] Step S6102 can be referred to the specific content of step S912 above, and will not be repeated here.
[0407] Step S262: If the number of target subject frames is greater than the third preset threshold and the number of non-target subject frames is greater than the second preset threshold, the first interface 61a is displayed.
[0408] Optionally, the third preset threshold can be set to 0.
[0409] For example, such as Figure 21As shown in (c), if the number of target body boxes is 3 and the number of non-target body boxes is 2, the first interface is displayed.
[0410] The first interface includes the first control corresponding to the virtual assistant, and the first control includes a selection control.
[0411] Optionally, the above description is based on a first interface 61a that includes only the input control 611 and a first interface 61b that includes only the selection control. Alternatively, it can be a first interface 61c that includes both the input control 611 and at least one selection control 612. The specific content of the first interface 61c can be referred to the above embodiments, and will not be repeated here.
[0412] Optionally, the embodiments shown above are distinguished by different first interfaces, and different scenarios are described in different first interfaces. Below, we first determine the image processing scenario, and then provide an exemplary description of the displayed first interface. For example, the following embodiment uses a first interface 61c including input controls and selection controls as an example for description.
[0413] In the first embodiment, combined with Figure 27 The example shown illustrates a scenario where content extraction is performed on the original image.
[0414] Figure 27 This is another schematic flowchart of an image processing method provided in an embodiment of this application.
[0415] like Figure 27 As shown, this embodiment includes steps S2701-S2708.
[0416] In step S2701, the electronic device receives the first operation triggered by the user and displays the first interface.
[0417] The specific details of step S2701 can be found in step S501 above, and will not be repeated here.
[0418] In step S2702, the electronic device responds to the first operation triggered by the user and detects whether the original image exists on the first interface.
[0419] The specific details of step S2702 can be found in step S502 above, and will not be repeated here.
[0420] In step S2703, if the electronic device has an original image in the first interface, it detects whether there is a third subject region in the target image.
[0421] The specific details of step S2703 can be found in step S241 above, and will not be repeated here.
[0422] Step S2704: If a third subject region exists in the target image, the first interface 61c is displayed.
[0423] The specific details of step S2704 can be found in step S242 above, and will not be repeated here.
[0424] In step S2705, when the electronic device responds to the user's second operation on the selection control in the first interface, the original image is input into the first image extraction model to obtain the first target image.
[0425] The specific details of step S2705 can be found in step S202 above, and will not be repeated here.
[0426] In step S2706, when the electronic device responds to the user performing a second operation on the input control in the first interface, NLP intent understanding is performed to obtain instruction information.
[0427] The specific details of step S2706 can be found in step S901 above, and will not be repeated here.
[0428] Step S2707: Input the instruction information and the original image into the visual positioning model to obtain the first subject region.
[0429] The specific details of step S2707 can be found in step S902 above, and will not be repeated here.
[0430] Step S2708: Input the first subject region into the first image extraction model to obtain the first target image.
[0431] The specific details of step S2708 can be found in step S903 above, and will not be repeated here.
[0432] In summary, in image cutout processing scenarios, electronic devices can utilize virtual assistants to perform cutout processing in different applications, avoiding the need to switch between different applications to achieve the cutout process, thereby improving image processing efficiency and ensuring user experience.
[0433] In another embodiment, combined with Figure 28 The following illustration demonstrates a scenario where background purification is performed on the original image. Exemplarily, the following embodiments use a first interface 61c including an input control and a selection control, and a first interface 61a including an input control, as examples for explanation.
[0434] Figure 28 This is another schematic flowchart of an image processing method provided in an embodiment of this application.
[0435] like Figure 28 As shown, this embodiment includes steps S2801-S2818.
[0436] In step S2801, the electronic device receives the first operation triggered by the user and displays the first interface.
[0437] The specific details of step S2801 can be found in step S501 above, and will not be repeated here.
[0438] In step S2802, the electronic device responds to the first operation triggered by the user and detects whether the original image exists on the first interface.
[0439] The specific details of step S2802 can be found in step S502 above, and will not be repeated here.
[0440] Step S2803: If the electronic device has an original image in the first interface, it determines whether the original image includes the target subject frame corresponding to the first target subject.
[0441] The specific details of step S2803 can be found in step S241 above, and will not be repeated here.
[0442] Step S2804: If a target subject bounding box exists in the target image, determine whether the original image includes at least one non-target subject bounding box corresponding to the second target subject.
[0443] The specific details of step S2804 can be found in step S242 above, and will not be repeated here.
[0444] Step S2805: If a non-target subject bounding box exists in the target image, the first interface 61c is displayed.
[0445] Step S2806: If there is no non-target subject bounding box in the target image, the first interface 61a is displayed.
[0446] Step S2807: If there is no target subject bounding box in the target image, determine whether the original image includes at least one non-target subject bounding box corresponding to the second target subject.
[0447] Step S2808: If there is a non-target subject bounding box in the target image, the first interface 61a is displayed.
[0448] In step S2809, when the electronic device responds to the user's second operation on the input control, NLP intent understanding is performed to obtain instruction information.
[0449] Step S2810: Based on the instruction information, determine whether to uniformly eliminate pedestrians.
[0450] Step S2811: Under the condition of uniformly eliminating pedestrians, the instruction information and the original image are input into the image recognition model to obtain the second target image.
[0451] Step S2812: Input the second target image and the original image into the image restoration model to obtain the first target image.
[0452] In step S2813, the electronic device inputs the original image and instruction information into the visual positioning model to obtain the output image of the bounding box of the second target subject.
[0453] In step S2814, the electronic device inputs the output image into the second image extraction model to obtain the second target image.
[0454] Step S2815: Input the second target image and the original image into the image restoration model to obtain the first target image.
[0455] In step S2816, when the electronic device responds to the user's second operation on the selection control, NLP intent understanding is performed to obtain control instructions.
[0456] In step S2817, the electronic device inputs the control command and the original image into the image recognition model to obtain the second target image.
[0457] Step S2818: Input the second target image and the original image into the image restoration model to obtain the first target image.
[0458] In summary, in scenarios involving the removal of passersby, electronic devices can utilize virtual assistants to perform image cutout processing across different applications, avoiding the need to switch between different applications to achieve the cutout, thus improving image processing efficiency and ensuring a better user experience.
[0459] The following is combined Figure 29 An example illustration of image processing is provided.
[0460] Figure 29 This is another schematic flowchart of an image processing method provided in an embodiment of this application.
[0461] Combination Figure 29 As shown, the image processing method provided in this application embodiment may include the following steps:
[0462] Step S2901: Display the first interface of the first application, which includes the original image.
[0463] Step S2902: Invoke the virtual assistant on the first interface to display the first control corresponding to the virtual assistant.
[0464] Step S2903: In response to the user's first operation on the first control, update the first interface to obtain the updated second interface; wherein, the second interface includes the first target image; the first target image is the updated image of the original image.
[0465] Step S2904: The second interface is displayed.
[0466] In summary, electronic devices can invoke virtual assistants within any application to process raw images, improving image processing efficiency and ensuring a better user experience.
[0467] It should be noted that the interfaces shown in the embodiments of this application are all examples and are not specifically limited here.
[0468] It is understood that, in order to achieve the aforementioned functions, electronic devices include corresponding hardware structures and / or software modules for performing each function. Those skilled in the art should readily recognize that the image enhancement method steps described in conjunction with the embodiments disclosed in this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is implemented in hardware or by software-driven hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0469] This application also provides an image processing apparatus, characterized in that it is applied to an electronic device and includes: a detection module configured to: detect whether an original image exists on a first interface in response to a first operation triggered by a user;
[0470] The display module is configured to display the first interface when the original image exists in the first interface; wherein the first interface includes a first control corresponding to the virtual assistant.
[0471] The processing module is configured to: respond to a second operation by the user on the first control, use an image processing model to process the original image in accordance with the second operation, and generate a first target image;
[0472] The display module is also configured to display the first target image.
[0473] This application provides an electronic device that may include a display screen (such as a touchscreen or a non-touchscreen), a memory, and one or more processors. The display screen, memory, and processors are coupled. The memory stores computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device can perform various functions or steps performed by the electronic device in the above method embodiments. The structure of the electronic device can be referred to... Figure 3 The structure of the electronic device shown.
[0474] Figure 30 This application provides a structural block diagram of a chip system.
[0475] This application also provides a chip system 3000, such as... Figure 30 As shown, the chip system 3000 includes at least one processor 3001 and at least one interface circuit 3002. The processor 3001 and the interface circuit 3002 are interconnected via lines. For example, the interface circuit 3002 can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, the interface circuit 3002 can be used to send signals to other devices (e.g., the processor 3001 or the touchscreen of an electronic device). Exemplarily, the interface circuit 3002 can read instructions stored in the memory and send those instructions to the processor 3001. When the instructions are executed by the processor 3001, the electronic device can perform the steps in the above embodiments. Of course, the chip system may also include other discrete devices, which are not specifically limited in this application embodiment.
[0476] This application also provides a computer storage medium that includes computer instructions. When the computer instructions are executed on the electronic device, the electronic device performs various functions or steps performed by the electronic device in the above method embodiments.
[0477] This application also provides a computer program product that, when run on a computer, causes the computer to perform various functions or steps performed by the electronic device in the above method embodiments.
[0478] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0479] It is readily understood that, based on the several embodiments provided in this application, those skilled in the art can combine, split, or reorganize the embodiments of this application to obtain other embodiments, none of which exceed the protection scope of this application.
[0480] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0481] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0482] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0483] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks. It should be noted that those skilled in the art, after considering the specification and practicing the application disclosed herein, will readily conceive of other embodiments of this application. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary technical means in the art not disclosed in this application. The description and examples are to be considered exemplary only, and the true scope of this application is indicated by the claims.
[0484] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. An image processing method, characterized in that, Applied to electronic devices, including: Displaying a first interface of a first application, the first interface including the original image; The virtual assistant is invoked on the first interface to display the first control corresponding to the virtual assistant; In response to a user's first operation on the first control, the first interface is updated to obtain an updated second interface; wherein, the second interface includes a first target image; the first target image is the updated version of the original image; The second interface is displayed.
2. The method according to claim 1, characterized in that, The step of responding to a user's first operation on the first control and updating the first interface to obtain an updated second interface includes: In response to the user's first operation on the input control in the first control, the first interface is updated to obtain the updated second interface; or, In response to the user's first operation on the selection control in the first control, the first interface is updated to obtain the updated second interface.
3. The method according to claim 2, characterized in that, The step of responding to a user's first operation on an input control in the first control, updating the first interface, and obtaining the updated second interface includes: In response to a second operation by the user on the input control, the original image is processed using an image processing model in accordance with the first operation to generate a first target image; The original image in the first interface is updated to the first target image to obtain the second interface.
4. The method according to claim 3, characterized in that, The step of responding to the user's first operation on the input control by using an image processing model to process the original image in accordance with the first operation to generate a first target image includes: In response to the user's first operation on the input control, instruction information is obtained; The image processing model is used to process the original image in accordance with the instruction information to generate the first target image.
5. The method according to claim 4, characterized in that, The step of using the image processing model to process the original image in accordance with the instruction information to generate the first target image includes: The original image is processed according to the instruction information using a visual positioning model and a first image extraction model to generate a first target image with at least one first target subject; or, The original image is processed using a second image extraction model and an image restoration model in accordance with the instruction information to generate a first target image that does not have at least one second target subject.
6. The method according to claim 5, characterized in that, The step of processing the original image using a visual positioning model and a first image extraction model in accordance with the instruction information to generate a first target image having at least one first target subject includes: The original image is processed using the visual positioning model in accordance with the instruction information to generate a first subject region; the first subject region is the area enclosed by the straight-line bounding box of the edge of the first target subject. The first image extraction model is used to process the first subject region to generate the first target image having at least one first target subject.
7. The method according to claim 6, characterized in that, The step of processing the original image using the visual positioning model in accordance with the instruction information to generate a first main body region includes: Extract the instruction features corresponding to the instruction information; Extract at least one image feature corresponding to the original image; Based on the instruction features, a first image feature is determined from at least one of the image features; Based on the first image feature, the first subject region corresponding to the instruction feature in the original image is determined.
8. The method according to claim 7, characterized in that, The step of determining a first image feature from at least one of the image features based on the instruction features includes: Based on the image features and the instruction features, determine the similarity between each image feature and the instruction feature; Based on the similarity, the first image feature corresponding to the instruction feature is determined.
9. The method according to claim 5, characterized in that, The step of processing the original image using a second image extraction model and an image restoration model in accordance with the instruction information to generate a first target image that does not have at least one second target subject includes: When the amount of information in the instruction information meets the first preset threshold, the original image is processed using the visual positioning model in accordance with the instruction information to generate an output image of the bounding box of the second target subject. The output image is processed using the second image extraction model to generate a second target image; The image restoration model is used to process the second target image and the original image to generate the first target image that does not have at least one second target subject.
10. The method according to claim 9, characterized in that, When the information content of the instruction information meets a first preset threshold, the visual positioning model is used to process the original image in accordance with the instruction information to generate an output image of the bounding box of the second target subject, including: If the amount of information in the instruction information meets the first preset threshold, extract the instruction features of the instruction information and extract at least one image feature of the original image. Based on the image clustering library and the instruction features, determine at least one second image feature from the image features; Based on the second image feature, a second main region in the original image corresponding to the instruction feature is obtained; wherein, the image clustering library includes the correspondence between the instruction feature and the preset image feature.
11. The method according to claim 10, characterized in that, The step of determining a second image feature from at least one of the image features based on the image clustering library and the instruction features includes: If the image clustering library includes the preset image features corresponding to the instruction features, then based on the preset image features, at least one of the image features is determined as the second image feature.
12. The method according to claim 10, characterized in that, The step of determining a second image feature from at least one of the image features based on the image clustering library and the instruction features includes: If the preset image feature corresponding to the instruction feature is not included in the image clustering library, a third interface is displayed, the third interface including at least one image control corresponding to the image feature; In response to a third operation by the user on the image control, the second image feature is determined from at least one of the image features.
13. The method according to claim 5, characterized in that, The step of processing the original image using a second image extraction model and an image restoration model in accordance with the instruction information to generate a first target image that does not have at least one second target subject includes: If the amount of information in the instruction information does not meet the first preset threshold, the second image extraction model is used to process the original image in accordance with the instruction information to generate a second target image. The image restoration model is used to process the second target image and the original image to generate the first target image that does not have at least one second target subject.
14. The method according to any one of claims 10-13, characterized in that, The step of activating the virtual assistant on the first interface to display the first control corresponding to the virtual assistant includes: Based on the original image, a target subject bounding box corresponding to at least one first target subject and a non-target subject bounding box corresponding to at least one second target subject are determined in the original image. When the number of target subject frames is less than a third preset threshold and the number of non-target subject frames is greater than the third preset threshold, the first control corresponding to the virtual assistant is displayed; wherein, the first control includes the input control.
15. The method according to any one of claims 10-13, characterized in that, The step of activating the virtual assistant on the first interface to display the first control corresponding to the virtual assistant includes: Based on the original image, a target subject bounding box corresponding to at least one first target subject and a non-target subject bounding box corresponding to at least one second target subject are determined in the original image. When the number of target subject frames is greater than a third preset threshold and the number of non-target subject frames is less than the third preset threshold, the first control corresponding to the virtual assistant is displayed; wherein, the first control includes the input control.
16. The method according to claim 2, characterized in that, The step of responding to a user's first operation on the selection control in the first control, updating the first interface, and obtaining the updated second interface includes: In response to the user's first operation on the selection control, obtain the control instruction; The image processing model is used to process the original image in accordance with the control instructions to generate the first target image. The original image in the first interface is updated to the first target image to obtain the second interface.
17. The method according to claim 16, characterized in that, The step of using the image processing model to process the original image in accordance with the control instructions to generate the first target image includes: The original image is processed using an image processing model in accordance with the control instructions to generate a first target image having at least one first target subject; or, The original image is processed using an image recognition model and an image restoration model in accordance with the control instructions to generate a first target image that does not have at least one second target subject.
18. The method according to claim 17, characterized in that, The step of processing the original image using an image processing model in accordance with the control instructions to generate a first target image having at least one first target subject includes: The original image is processed using a first image extraction model in accordance with the control instructions to generate a first target image having at least one first target subject.
19. The method according to any one of claims 16-18, characterized in that, The step of activating the virtual assistant on the first interface to display the first control corresponding to the virtual assistant includes: The original image is processed using a third image extraction model to generate a third subject region in the original image, wherein the third subject region is the region enclosed by the curved bounding box of the edge of the first target subject; When the third main area meets the fourth preset threshold, the first interface is displayed; wherein, the first interface includes a first control corresponding to the virtual assistant, and the first control includes the selection control.
20. The method according to claim 17, characterized in that, The step of processing the original image using an image recognition model and an image restoration model in accordance with the control instructions to generate a first target image that does not have at least one second target subject includes: The image recognition model is used to process the original image in accordance with the control instructions to generate a second target image; The image restoration model is used to process the second target image and the original image to generate the first target image that does not have at least one second target subject.
21. The method according to claim 20, characterized in that, The step of using the image recognition model to process the original image in accordance with the control instructions to generate a second target image includes: The original image is processed using an image segmentation model in accordance with the control instructions to generate a second subject region corresponding to the second target subject. The second extraction model is used to process the second main body region to determine the second target image.
22. The method according to claim 21, characterized in that, The step of using an image segmentation model to process the original image in accordance with the control instructions to generate a second subject region corresponding to the second target subject includes: Extract the text features of the instruction control; Extract at least one image feature from the original image; the image feature includes a second target subject image feature; Based on the correlation between each of the image features and the second target subject image features, the second subject region in the original image corresponding to the instruction feature is obtained.
23. The method according to any one of claims 20-22, characterized in that, The step of activating the virtual assistant on the first interface to display the first control corresponding to the virtual assistant includes: Based on the original image, at least one target subject bounding box corresponding to the first target subject and at least one non-target subject bounding box corresponding to the second target subject are determined in the original image; When the number of target subject frames is greater than a third preset threshold, and the number of non-target subject frames is greater than the third preset threshold, the first control corresponding to the virtual assistant is displayed; wherein, the first control includes the selection control.
24. The method according to claim 23, characterized in that, The step of determining, based on the original image, at least one target bounding box corresponding to the first target subject and at least one non-target bounding box corresponding to the second target subject in the original image includes: The original image is processed using a subject recognition model to generate a third target image; the third target image includes the subject bounding box corresponding to the first target subject; The third target image is processed using an instance segmentation model to generate a fourth target image; the fourth target image includes at least one target subject bounding box corresponding to the first target subject, and at least one non-target subject bounding box corresponding to the second target subject.
25. An electronic device, characterized in that, include: A display screen, a memory, and one or more processors; the display screen, the memory, and the processors are coupled; wherein the memory stores computer program code, the computer program code including computer instructions, which, when executed by the processor, cause the electronic device to perform the image processing method as described in any one of claims 1-24.
26. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed electronically, cause the electronic device to perform the image processing method as described in any one of claims 1-24.
27. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the image processing method as described in any one of claims 1-24.