Image processing method and device
By classifying image content and integrating dedicated tools, combined with natural language interaction and multi-model collaborative scheduling, the problems of scattered image processing tools and poor scene adaptability are solved, achieving an efficient and convenient image processing experience.
Patent Information
- Application Number
- CN202511649931.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-01-30
AI Technical Summary
In existing technologies, image processing tools are scattered and have poor scene adaptability, resulting in inconvenience and low efficiency for users.
By classifying image content, the target image type is determined, a dedicated processing page is matched, and the target processing tool is invoked. Combined with natural language interaction and multi-model collaborative scheduling, scenario-based processing is achieved.
It achieves a precise classification, focused tools, convenient operation, and instant results in image processing, greatly improving processing efficiency and user satisfaction.
Smart Images

Figure CN121437992A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to the fields of natural language processing, image processing, computer vision and deep learning. Background Technology
[0002] With the widespread adoption of mobile smart terminals and the expansion of digital image application scenarios, user demand for end-to-end image services, from shooting to processing and analysis, continues to grow, with a particular emphasis on the convenience and intelligence of the processing process. Currently, mainstream cloud storage applications generally integrate AI (Artificial Intelligence) camera functionality to meet this demand. Among these, the AI camera, as a core interactive module, has become a crucial entry point connecting users with various image processing services. Its design revolves around improving user operational efficiency and aims to lower the barrier to entry for users through technological optimization.
[0003] In the field of intelligent image processing, existing technologies have formed a relatively mature technical system and application models. Among them, intelligent image recognition technology serves as the core support, enabling rapid identification of image content features after image capture, providing foundational data for subsequent services. Simultaneously, tool recommendation technology based on user demand data analysis is widely used. This technology can accurately match and distribute existing image tool capabilities to users' high-priority needs and core pain points in image processing, providing users with an efficient and tailored intelligent image processing experience. Summary of the Invention
[0004] This disclosure provides an image processing method, apparatus, device, storage medium, and program product.
[0005] In a first aspect, embodiments of this disclosure propose an image processing method, comprising: classifying a target image by image content to determine the target image type; determining a target processing page corresponding to the target image type, the target processing page including at least one target processing tool; and calling the target processing tool in the target processing page to process the target image and generate a processing result.
[0006] Secondly, embodiments of this disclosure provide an image processing apparatus, comprising: a classification module configured to classify the image content of a target image and determine the target image type; a determination module configured to determine a target processing page corresponding to the target image type, the target processing page including at least one target processing tool; and a processing module configured to call the target processing tool in the target processing page to process the target image and generate a processing result.
[0007] Thirdly, embodiments of this disclosure provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method as described in the first aspect.
[0008] Fourthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the methods described in the first aspect.
[0009] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0010] The key or essential features of the embodiments disclosed herein are not intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. The drawings are provided for a better understanding of the invention and are not intended to limit the scope of this disclosure. Wherein: Figure 1 This is a flowchart of an embodiment of the image processing method according to the present disclosure; Figure 2 This is a flowchart of yet another embodiment of the image processing method according to the present disclosure; Figure 3 This is a diagram of the closed-loop architecture for multi-scene image processing in AI cameras. Figure 4 This is a schematic diagram of the multi-scene image processing and interactive interface of an AI camera; Figure 5 This is a flowchart of AI camera image processing instruction scheduling and execution. Figure 6 A schematic diagram of a structural design of an image processing apparatus according to an embodiment of the present disclosure; Figure 7 A block diagram of an electronic device used to implement the image processing method of the embodiments of this disclosure. Detailed Implementation
[0012] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0013] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0014] Figure 1 A flowchart 100 is shown as an embodiment of an image processing method according to the present disclosure. The image processing method includes the following steps: Step 101: Classify the target image content to determine the target image type.
[0015] In this embodiment, the entity executing the image processing method can classify the target image content and determine the target image type.
[0016] The executing entity can be, for example, an application client that integrates AI camera functionality (such as a cloud storage application client), which can acquire target images in two core ways: one is images generated by the user in real time through the AI camera; the other is images already stored by the user from the local device. Images acquired in both ways are used as initial image processing materials to ensure coverage of both real-time shooting and existing image processing scenarios.
[0017] When classifying target images, the executing entity can employ an image content classification model. This model is equipped with dedicated image recognition feature operators, capable of accurately extracting multi-dimensional core features of the target image. Extraction dimensions include, but are not limited to: tables, watermarks, handwriting, QR codes, foreign language characters, the proportion of people in the subject, document text density, exam question type features, and the types of scene elements. After assigning weights and making comprehensive judgments on each dimension of features, the model outputs the target image type. The classification results strictly cover five major categories: portraits, documents, exam papers, scenes, and others, ensuring comprehensive coverage of mainstream daily image processing scenarios.
[0018] Considering the potential for bias in automatic model classification, the system also provides users with an interactive interface to manually switch classifications. Users can reselect the type that best suits their needs from five categories, ensuring classification accuracy while giving users full autonomy and preventing classification errors from affecting the subsequent processing experience.
[0019] Step 102: Determine the target processing page corresponding to the target image type. The target processing page includes at least one target processing tool.
[0020] In this embodiment, the executing entity can determine the target processing page corresponding to the target image type, and the target processing page includes at least one target processing tool.
[0021] The execution entity can automatically match and redirect to the corresponding dedicated processing page based on the target image type. Each processing page is meticulously designed based on the core user needs of the corresponding scenario, deeply integrating frequently used target processing tools in that scenario to ensure a high degree of tool-scenario compatibility and avoid functional redundancy.
[0022] If the target image type is a portrait, the corresponding target processing page integrates a full range of portrait beautification tools such as AI dress-up, AI background replacement, AI style effects, one-click beautification, hair styling tools, facial reshaping, skin beautification tools, body beautification tools, and teeth whitening.
[0023] If the target image type is a document, the corresponding target processing page integrates tools for removing handwriting, removing watermarks, image super-resolution, AI enhancement, text extraction, converting to Word, converting to PDF (Portable Document Format), converting to Excel, and document scanning, among other document digitization tools.
[0024] If the target image type is a test paper, the corresponding target processing page integrates specialized tools for learning scenarios such as test paper grading, problem solving, error correction, paper beautification, and handwriting removal.
[0025] If the target image type is a scene, the corresponding target processing page integrates a variety of AI filter tools such as AI brightening, one-click ultra-high definition, image quality enhancement, AI picture book style, AI comic style, healing anime style, and childlike picture book style.
[0026] If the target image type is other, the corresponding target processing page integrates comprehensive processing tools such as painting scanning, image enhancement, ID card scanning, and ID photo taking.
[0027] Meanwhile, an AI dialogue assistant module entry is fixed at the bottom of each type of target processing page. The entry style is uniform and easy to identify, providing users with a natural language interaction channel to meet the personalized and long-tail processing needs that cannot be covered by standardized tools.
[0028] Step 103: Use the target processing tool in the target processing page to process the target image and generate the processing result.
[0029] In this embodiment, the executing entity can call the target processing tool in the target processing page to process the target image and generate processing results.
[0030] Once a user enters the target processing page, the execution entity will automatically trigger basic processing operations for the corresponding scenario and display the effects in real time, following the principle of real-time processing design. This allows users to intuitively perceive the processing value. For example, portrait processing pages automatically perform basic beautification, document processing pages automatically perform sharpening and slight noise reduction, exam paper processing pages automatically perform paper noise reduction and contrast optimization, scene processing pages automatically perform color optimization and detail enhancement, and other processing pages automatically perform basic sharpening. All real-time processing effects can be restored with one click, ensuring the user's right to choose.
[0031] Users can initiate processing requests by clicking the target processing tool icon on the processing page. The tool then executes the corresponding atomic algorithm to precisely process the target image. For example, when calling the watermark removal tool, the algorithm intelligently removes the watermark from the image without damaging other content in the original image through steps such as watermark region recognition, edge detection, and background pixel repair. When calling the PDF conversion tool, the algorithm first extracts the text and layout information from the document image using OCR (Optical Character Recognition) technology, and then performs layout conversion according to the PDF standard format, ensuring that the text is copyable and the layout is consistent with the original. Figure 1 When the error correction tool is invoked, the algorithm automatically identifies the areas of incorrect questions in the test paper, extracts the questions and wrong answers, and allows users to add the correct answers to generate a structured error correction set.
[0032] Once processed, the generated results are displayed in image or document format in the central area of the processing page. Users can preview the effect using pinch-to-zoom, swipe left and right, and other gestures. If satisfied with the result, the user can trigger the export operation. For further optimization, users can continue to use other target processing tools or initiate personalized requests through the AI dialogue assistant module.
[0033] The system supports converting processed images into formats corresponding to the target image type: document images can be converted to Word, PDF, and Excel formats; portrait and landscape images can be converted to common image formats such as JPG (Joint Photographic Experts Group) and PNG (Portable Network Graphics); and ID photos can be saved in high-definition formats that meet printing standards. Users can choose to save the processed results to their local devices or share them directly to social media platforms and office software, achieving a complete solution for their image processing needs.
[0034] This disclosure provides an image processing method that solves the problems of scattered tools and poor scene adaptability in the prior art by deeply integrating scene-based classification with dedicated tools. It achieves a processing experience with accurate classification, focused tools, convenient operation, and real-time results, greatly improving image processing efficiency and user satisfaction.
[0035] Figure 2 A flow 200 of another embodiment of the image processing method according to the present disclosure is shown. The image processing method includes the following steps: Step 201: Classify the target image content to determine the target image type.
[0036] Step 202: Determine the target processing page corresponding to the target image type. The target processing page includes at least one target processing tool.
[0037] In this embodiment, the specific operations of steps 201-202 have been described. Figure 1 The steps 101-102 in the illustrated embodiments are described in detail and will not be repeated here.
[0038] Step 203: Obtain the user's image processing requirements.
[0039] In this embodiment, the executing entity can obtain the user's image processing requirements.
[0040] The execution entity can obtain users' image processing needs through two core methods, covering both standardized and personalized scenarios: First, standardized needs initiated by users clicking on tool icons on the target processing page, such as clicking the face-slimming tool to perform natural face-slimming on portraits, and clicking the handwriting removal tool to remove handwriting marks on exam papers; Second, natural language needs input by users through the AI dialogue assistant module at the bottom of the processing page. This method allows users to describe personalized needs in complete sentences, such as, "Make my skin more delicate, and create a hazy light spot effect in the background"; "Turn this landscape photo into a Chinese ink painting style, and highlight the mountains in the picture"; "Identify key clauses in a document and generate a summary," etc.
[0041] The AI dialogue assistant module offers flexible input methods, including but not limited to: clicking the recommended question above the input box (generated based on the current image type and high-frequency needs), manually entering text and clicking the send button, clicking the encyclopedia tag catch-all entry (automatically providing catch-all text and the latest result image on the current processing page), and clicking the recommended suggestion below the reply text (related needs generated based on previous interaction intents). The default prompt text for the input box is "Enter your request," which switches to "Pause to add new requirements" during the model's response to avoid repetitive input by the user.
[0042] Step 204: Input the image processing requirements and target image into the dialogue assistant model and output the processing results.
[0043] In this embodiment, the executing entity can input image processing requirements and target images into the dialogue assistant model and output the processing results.
[0044] The executing entity inputs the user's image processing requirements along with the latest target image (initial image or processed intermediate result image) into the dialogue assistant model. The model processes the image and outputs the results according to the following logic, ensuring accurate and efficient response: First, the dialogue assistant model performs semantic recognition and intent parsing on image processing requests. Through techniques such as keyword extraction, sentence structure analysis, and context association, it accurately determines the type of user intent. The intent can be classified into four main categories: image processing, image understanding, image-to-text, and other requests. Among them, the image processing intent is further subdivided into subcategories such as style conversion, element modification, and quality optimization.
[0045] If the user's intent is determined to be image processing, the model uses a tool matching engine to search for the existence of an atomic processing algorithm that satisfies this intent. The matching rules are based on the mapping relationship between natural language keywords and tool functions. If a matching atomic processing algorithm exists, it is directly invoked to process the target image and generate a processed image. For example, if the user's request is "brighten the image," the model finds the "AI brighten" atomic algorithm and invokes it to adjust the image brightness parameters according to the user's requested intensity, ensuring a natural effect. If no matching atomic processing algorithm exists, the model invokes an image creation model, using generative AI technology to fulfill the user's request and generate a processed image. For example, if the user's request is "generate an image of robots moving among green vegetation-covered buildings in a futuristic city," the model directly invokes a portrait creation model to generate a corresponding creative image based on the natural language description.
[0046] If the user's intent is determined to be image understanding or image-to-text, the model invokes a multimodal model to jointly process the user's intent and the target image, generating the intended text. For example, if the user's requirement is "describe the content of this image," the model analyzes the main elements, scene atmosphere, and key information in the image, and then outputs a structured text description; if the user's requirement is "identify key clauses in a document," the model extracts the text content through OCR, then identifies the core clauses through semantic analysis and outputs a summary; if the user's requirement is "provide outfit suggestions based on an image," the model identifies the clothing style and scene of the people in the image, generating multiple suitable outfit options.
[0047] If the user's intent is determined to be something else, the model will also call the multimodal model to generate the corresponding response text to meet the user's non-image processing needs. For example, if the user asks "What kind of plant is in this picture?", the model will use image recognition and encyclopedia data to output the plant's name and a brief description.
[0048] During processing, the model supports multi-round interactions, allowing users to continuously raise new requests based on the processing results. The model responds and updates the processing results in real time, ensuring seamless interaction. For example, if a user first requests "help me slim my face," and then requests "make my eyes bigger" after processing, the model automatically associates the previously processed image with the new processing instructions, eliminating the need for the user to repeatedly upload images.
[0049] In response to shutdown commands during processing, the model responds according to the following rules to ensure a good user experience: If a close command for the target processing page is detected after the image is generated, the target processing page is closed directly, the processed image is displayed in the current dialogue record of the dialog page, the result image is synchronously entered into the conversation record, and the most recent task is not displayed.
[0050] If a close command is detected on the target processing page before image generation, a secondary confirmation pop-up appears with the text "Image processing will be terminated after exiting, please confirm." If the user clicks "Confirm Exit," the target processing page closes, and the target image is displayed in the current chat history on the chat page. This task is not added to the "Recent Tasks" list. The front end performs a "Cancel Session" operation to prevent the result image of the asynchronous task from being inserted into the chat history later. The text in the processing is replaced with "You have stopped this task." If the user clicks "Cancel," the loading state continues.
[0051] If a new image processing request is obtained before image generation, the new request is input into the dialogue assistant model, and the new processed image is output and displayed in the current dialogue record on the dialogue page. After the original processed image is generated, it is displayed in the historical dialogue record on the dialogue page, ensuring the orderly processing of multiple requests and avoiding confusion of results.
[0052] Furthermore, the model employs an edge compression strategy, performing compression twice on the latest result image: once when the user enters the processing page and once when they open the AI dialogue assistant panel. Different compression parameters apply to different models / tools. For example, the multimodal model uses a compression resolution of 1024×1024, the portrait creation model uses 800×800, and the atomization algorithm uses 1280×1280, ensuring processing efficiency while avoiding excessive loss of image quality.
[0053] The output of the processing results follows these rules: If the image is being processed, the dialog page will insert the text "Image processing complete" and display an "Undo" button. After the user clicks it, the text will update to "Processing effect has been undone," and the processing page will resume displaying the image from the previous state. The undo operation is only valid for the most recent processing step.
[0054] Below the last processing result, a "Regenerate" button will be displayed. Regardless of whether the returned result is an image, text, or a shortcut, you can click it to re-initiate the request, and the model will re-execute the processing flow to generate a new result.
[0055] If an exception occurs during processing, such as loading timeout or token exceeding the limit, the model returns the fixed message "Processing failed, please retry" and provides a retry button. Users can click the button to re-initiate the processing request. If the token exceeds the limit, the current message will be used as the reply.
[0056] This disclosure provides an image processing method that, through natural language interaction and multi-model collaborative scheduling, not only meets users' standardized processing needs but also covers personalized and long-tail demands. At the same time, it ensures smooth operation through comprehensive interaction rules, greatly improving the flexibility and intelligence level of image processing.
[0057] Figure 3 The diagram illustrates a closed-loop architecture for multi-scene image processing in an AI camera. This architecture, centered on "multimodal input + scene-based classification + end-to-end processing + result output + record management," achieves a complete closed loop from image acquisition to consumption. This fundamentally changes the existing model where AI cameras merely serve as entry points for traffic diversion, resolving issues of fragmented functions and inconsistent user experiences. The specific architectural logic is as follows: Input Layer 301: Multimodal Input Ensures Recognition Stability. The architecture's input layer 301 supports multimodal image input, covering two core input scenarios: first, images captured in real-time by the user through an AI camera; and second, existing images imported by the user from their local device. After on-device preprocessing, the input images are transmitted to the classification layer 302 to ensure the stability and compatibility of subsequent processing. The preprocessing process does not alter the original image data, ensuring the integrity of the original image information.
[0058] Classification Layer 302: Dual Determination of Image Classification and Intent Understanding. Input images first enter Classification Layer 302, where multi-dimensional image features are extracted using a classification algorithm to automatically classify them into five categories. Classification accuracy is optimized through extensive sample training to ensure coverage of classification needs in everyday scenarios. Simultaneously, Classification Layer 302 combines user historical behavior data (such as past processing preferences and frequently used tools) with the input scenario (such as shooting time and device location) to preliminarily determine the user's potential processing intent using an intent understanding model (e.g., shooting documents on weekdays likely requires conversion to Word, while shooting portraits on weekends likely requires beautification). This provides a basis for subsequent tool recommendations and processing strategies, making the service more aligned with user expectations. Classification results support manual correction; users can adjust the image type through the classification switch to ensure classification accuracy.
[0059] Processing Layer 303: Scenario-Specific Processing with AI Dialogue as a Backup. The output of Classification Layer 302 triggers scenario-specific distribution in Processing Layer 303. Processing Layer 303 employs a dual-guarantee mechanism, with tool processing as the primary method and model processing as a backup. It supports both GUI (Graphical User Interface) and LUI (Language User Interface) interaction methods: For five major scenarios—portraits, documents, exam papers, scenes, and others—corresponding dedicated processing pages and toolsets are invoked. Each scenario processing follows the principle of real-time processing, automatically triggering basic optimization operations. Users can directly invoke tools for standardized processing via the GUI (clicking the tool icon). Tool operations support real-time preview, parameter adjustment, and one-click restoration to meet users' refined processing needs. If users have personalized needs or long-tail demands that tools cannot meet, they can initiate natural language requests (LUI method) through the AI dialogue assistant module at the bottom of each processing page. After the dialogue assistant model parses the request, it schedules either an atomic processing algorithm or an image creation model to execute the processing based on the existence of a matching atomic algorithm. The processing supports multi-round interactions, allowing users to continuously raise new requests based on the processing results. The model responds and updates the processing results in real time, ensuring seamless interaction. Processing layer 303 also has differentiated design capabilities, performing specific optimizations for user experience pain points in different scenarios: such as optimizing OCR recognition accuracy and format conversion consistency in document scenarios, optimizing the clarity of problem-solving steps and the structured organization of incorrect questions in test paper scenarios, and optimizing the naturalness of beautification and the accuracy of style conversion in portrait scenarios.
[0060] Output Layer 304: Multi-channel result export and session record management. The processing results (images or text) generated by Processing Layer 303 are distributed through Output Layer 304 to meet different user needs. For image results, users can save the processed image to their local device or share it directly to social media platforms and office software. The sharing process supports generating both original and compressed images to adapt to the transmission requirements of different platforms. For text results, long-press copying is supported, making it convenient for users to edit or paste for later use. Text such as outfit suggestions, document summaries, and problem-solving ideas can be directly copied to other applications. Simultaneously, the output layer 304 synchronously records session data, storing session records for the past year in reverse chronological order of session initiation time. Recorded content includes, but is not limited to, original images, user commands, processing results (images / text), processing time, and tools used. Users can view historical processing through the session record entry. Incomplete processing tasks are displayed in real-time in the records (in production / generated), with progress pushed in real-time. If progress synchronization fails, a fixed message "Image processing progress can be viewed in 'Recent Tasks'" and a jump to the recent tasks section are displayed. When the session record list exceeds one screen and the user scrolls to the bottom, an underlined message "Maximum of records displayed is for the past year" is displayed, reminding the user of the storage range.
[0061] This closed-loop architecture achieves full-process coverage from image acquisition, classification, processing to output and recording, solving the problems of functional fragmentation and poor scene connection in existing technologies, providing users with one-stop image processing services, and greatly improving processing efficiency and user experience.
[0062] Figure 4 The diagram illustrates the multi-scene image processing and interactive interface of an AI camera. This interface, centered around creative generation scenarios, constructs a complete interactive chain from "demand input - intelligent processing - result output." Details of each module are as follows: Creative Generation Scenario: The interface showcases the entire process of users refining creative images and interacting with natural language. First, an AI style selection panel is presented, providing access to functions such as "AI Playstyles" and "Portrait Beautification." Users can click on style thumbnails to preview the effects. Next, the user enters the natural language input stage. After typing "Help me brighten the image and add a sticker to the person's face," the system displays a "Image processing in progress" progress indicator and supports an "Pause and add new requirements" interaction, satisfying the user's need for independent control over the processing. After processing is complete, the user can issue the command "Help me describe the content of the image," and the system generates the text "This image is of a beautiful girl wearing a T-shirt and skirt, with a sunny and sweet smile," achieving a closed loop between image understanding and text generation needs.
[0063] The overall interface, through its scenario-based functional partitions, dual interactive entry points of natural language and tool icons, and processing progress and result annotations, solves the problem of one-stop operation experience in creative generation scenarios, and intuitively demonstrates the practical value of the technical solution in the integration of portrait creative processing and natural language interaction.
[0064] Figure 5 The flowchart illustrates the scheduling and execution of AI camera image processing commands. This process is designed around the entire chain of user command reception, parsing, scheduling, execution, and feedback. Standardized process control ensures the efficiency, accuracy, and consistency of command processing. The specific process is as follows: Step 501: The user issues a query: initiating the command process.
[0065] Users initiate image processing commands through two core channels to ensure that no requirements are missed in the collection of data: Standardized tool commands: Users initiate commands by clicking on the tool icons on the target processing page. Commands correspond one-to-one with tool functions. For example, clicking the face slimming tool corresponds to the natural face slimming processing command, clicking the handwriting removal tool corresponds to the command to remove handwriting marks from exam papers, and clicking the ID card background color change tool corresponds to the command to change the background color of ID card photos. These types of commands do not require additional parsing and are directly associated with the corresponding atomic algorithms.
[0066] Natural language commands: Users input through the AI dialogue assistant module, supporting complete sentence descriptions, such as "Change the background to a sunset at the beach and add a retro filter to the person", "Extract text from the document and generate an Excel spreadsheet", "Make this photo look like a Picasso painting". These commands need to be semantically parsed and matched with the resources.
[0067] Upon receiving the instruction, the system automatically associates it with the latest target image (initial image or processed intermediate result image) to form a "instruction + image" processing unit, ensuring the accuracy of the processed object and avoiding deviations in processing results due to image switching.
[0068] Step 502, Is the intent clear: Determine the clarity of the requirement.
[0069] The "instruction + image" processing unit is evaluated for intent clarity. If the user's intent is clear (e.g., the instruction contains clear tool pointers or effect descriptions), the process proceeds to step 504 to determine the intent; if the user's intent is ambiguous (e.g., only inputting "edit" or "optimize" the image), the process proceeds to step 503 to guide the user to supplement information.
[0070] Step 503: Guide users to supplement information: clarify vague needs.
[0071] The system initiates follow-up questions through the AI dialogue assistant module, such as "Do you want to beautify, style-change, or improve the quality of the image?", guiding the user to provide specific information. After the user provides the information, the system returns to step 501, where the user issues a query, and the process is restructured into a "command + image" processing unit to avoid invalid processing.
[0072] Step 504, determine intent: classify requirement types.
[0073] Classify and judge clear user intentions to determine their respective types of requests (image understanding / image-to-text, other requests, image processing requests, etc.) to provide a basis for subsequent tool / model scheduling.
[0074] Step 505, Image Understanding / Image-to-Text Request: Non-image processing requirements.
[0075] If the user's request is image understanding (such as "describe the content of this image") or image-to-text (such as "generate a story based on the image"), proceed to step 506 to invoke the multimodal model response process.
[0076] Step 506: Call the multimodal model to respond: non-image-related requirement processing.
[0077] The multimodal model is invoked to jointly process the user intent and the target image, generating intent copy (such as text description, story content, analysis report, etc.). After processing, step 507 is directly displayed to show the returned copy.
[0078] Step 507: Directly display the returned text: Text-related results output.
[0079] The copywriting results generated by the multimodal model are streamed to the dialog page. Recommended suggestions (related requirements generated based on the current copywriting) are displayed below the content. Clicking on the recommended suggestion allows users to directly initiate a new command. The copywriting supports long-press copying for easy reuse by users.
[0080] Step 508, Other requirements: General non-image processing requirements.
[0081] If the user's request is a general request that is not clearly categorized (such as "What is the name of the plant in this picture?"), proceed to step 509 to invoke the multimodal model response process.
[0082] Step 509: Call the multimodal model to reply: General request processing.
[0083] The multimodal model is invoked to parse the user's intent and generate corresponding response text (such as plant names and descriptions, general knowledge answers, etc.). After processing, the process proceeds to step 510 to directly display the return text.
[0084] Step 510: Directly display the returned text: general text type result output.
[0085] The generic copy-type results generated by the multimodal model are streamed to the dialog page, and the display format is the same as the direct display of the returned copy in step 507. It supports operations such as copying and recommended sug interaction.
[0086] Step 511, Image Processing Requirements - Existing AI Camera Tools: Standardized Scene Processing.
[0087] If the user's request is image processing and there is a matching atomic processing algorithm (such as watermark removal, PDF conversion, hair seam repair, etc.), proceed to step 512 to call the existing tool processing flow.
[0088] Step 512: Call existing tools for processing: atomic algorithm scheduling.
[0089] The corresponding atomic processing algorithm is directly scheduled to process the target image. The scheduling process employs an edge compression strategy to ensure a balance between processing efficiency and image quality. During algorithm execution, a progress bar provides real-time feedback on the processing progress, allowing the user to intuitively understand the processing status. After processing is complete, step 513, the tool's real-time processing flow, is executed.
[0090] Step 513, whether the tool processes in real time: verify the processing effect.
[0091] Verify the tool's processing results to determine if they meet the user's requirements. If they do, proceed to step 514 to directly display the result graph and insert the result history; if they do not, proceed to step 515 to directly display the jump instruction and insert the jump instruction history.
[0092] Step 514: Directly display the results of the historical data insertion: output the image-type results.
[0093] The processed image results are displayed at the top of the dialog page or in the center of the processing page, supporting zooming in and out with two fingers to view details. The dialog page includes the text "Image processing complete" and displays an undo button. When the user clicks the button, the text updates to "Processing effect has been undone," and the processing page resumes displaying the image from the previous state. Below the last processing result, a regenerate button is displayed. Clicking this button allows the user to resend the current command and initiate a processing request, generating a new processing result. Simultaneously, this result is written to the conversation history module and inserted into the history.
[0094] Step 515: Directly display the history of jump instructions and insert instructions: function entry guide.
[0095] If the processing result is a function entry requirement (such as recommending similar tools), insert the text "Based on the image, we recommend the following tools for you" into the dialog box, displaying a shortcut entry style. After the user clicks, they will enter a separate tool processing page. After the processing page is saved, they will enter the result page according to the saving logic of each processing page. Insert the function entry command into the history, but do not insert the result image or document.
[0096] Step 516: After exiting, record the abnormal status without inserting the result image or document.
[0097] If an exception occurs during processing that causes the process to be interrupted (such as the user actively closing the processing page and the task not being completed), the system records the status (such as processing interruption). No result image or document is inserted into the history record to ensure the information is complete when tracing back later.
[0098] Step 517, Image Processing Request - No Existing Tools: Model-Based Handling.
[0099] If the user requests image processing but there is no matching atomic processing algorithm, proceed to step 518 to call the image creation model response process.
[0100] Step 518, call the image creation model to reply: Generative AI processing.
[0101] The image creation model is scheduled. After receiving instructions and the compressed image, the model performs generative processing to generate the processed image. After processing, step 519 is taken to directly display the result image and insert the historical record result flow.
[0102] Step 519: Directly display the result of the historical data insertion: output the result of the generative image class.
[0103] The image-type results generated by the image creation model are displayed at the top of the dialog page or in the center of the processing page. The display and interaction logic is consistent with the result image history insertion result in step 514. At the same time, the result is synchronously saved to the conversation record module and inserted into the history record.
[0104] Step 520: The user sends another query: multi-turn interaction support.
[0105] Users can initiate a new query based on the current processing result or session record. The process returns to step 501, where the user issues a query, supporting multiple rounds of continuous interaction to achieve a deep fulfillment of needs.
[0106] This flowchart, through the standardized end-to-end design of steps 501-520, ensures that every step of the user command process, from receiving feedback to completion, has clear processing rules. This solves the problems of chaotic command processing and unsmooth interaction in existing technologies, and significantly improves the reliability of command processing and user experience.
[0107] Further reference Figure 6 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of an image processing apparatus, which is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0108] like Figure 6 As shown, the image processing apparatus 600 of this embodiment may include: a classification module 601, a determination module 602, and a processing module 603. The classification module 601 is configured to classify the target image content and determine the target image type; the determination module 602 is configured to determine the target processing page corresponding to the target image type, the target processing page including at least one target processing tool; and the processing module 603 is configured to call the target processing tool in the target processing page to process the target image and generate a processing result.
[0109] In this embodiment, the specific processing of the classification module 601, the determination module 602, and the processing module 603 in the model training device 600, and the resulting technical effects, can be found in reference to [reference needed]. Figure 1 The relevant descriptions of steps 101-103 in the corresponding embodiments will not be repeated here.
[0110] In some optional implementations of this embodiment, the classification module 601 is further configured to: input the target image into the image content classification model, use the image recognition feature operator of the image classification model to identify the target image features, and perform image content classification based on the target image features to output the target image type.
[0111] In some optional implementations of this embodiment, the processing module 603 is further configured to: acquire the user's image processing requirements; process the target image using the target processing tool corresponding to the image processing requirements, and generate processing results.
[0112] In some optional implementations of this embodiment, the processing module 603 is further configured to: input the image processing requirements and the target image into the dialogue assistant model, and output the processing results.
[0113] In some optional implementations of this embodiment, the processing module 603 is further configured to: use a dialogue assistant model to perform intent parsing on the image processing request and determine the user intent; in response to determining that the user intent is image processing, process the target image based on the user intent to generate a processed image; in response to determining that the user intent is image understanding or graph-to-text, call a multimodal model to process the user intent and the target image to generate intent text.
[0114] In some optional implementations of this embodiment, the processing module 603 is further configured to: in response to determining that an atomic processing algorithm that satisfies the user's intent exists, call the atomic processing algorithm that satisfies the user's intent to process the target image and generate a processed image; in response to determining that no atomic processing algorithm that satisfies the user's intent exists, call the image creation model to process the target image and generate a processed image.
[0115] In some optional implementations of this embodiment, the processing module 603 is further configured to: in response to detecting a close instruction of the target processing page after the processed image is generated, close the target processing page and display the processed image in the current dialogue record of the dialogue page; in response to detecting a close instruction of the target processing page before the processed image is generated, close the target processing page and display the target image in the current dialogue record of the dialogue page.
[0116] In some optional implementations of this embodiment, the processing module 603 is further configured to: in response to obtaining a new image processing requirement before the image is generated, input the new image processing requirement into the dialogue assistant model, output a new processed image, display the new processed image in the current dialogue record of the dialogue page, and display the processed image in the historical dialogue record of the dialogue page after the image is generated.
[0117] In some optional implementations of this embodiment, the image processing apparatus 600 further includes an output module configured to convert the processed image into a format corresponding to the target image type and save it locally and / or publish it to the target social platform.
[0118] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0119] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0120] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0121] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 702 or a computer program loaded into random access memory (RAM) 703 from storage unit 708. The RAM 703 may also store various programs and data required for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.
[0122] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0123] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as image processing methods. For example, in some embodiments, the image processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the image processing method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform image processing methods by any other suitable means (e.g., by means of firmware).
[0124] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0125] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0126] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0127] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0128] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0129] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.
[0130] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution provided in this disclosure can be achieved, and this is not limited herein.
[0131] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An image processing method, comprising: performing image content classification on a target image to determine a target image type; determining a target processing page corresponding to the target image type, the target processing page comprising at least one target processing tool; calling the target processing tool in the target processing page to process the target image to generate a processing result.
2. The method of claim 1, wherein, The target image type is determined by performing image content classification on the target image, comprising: inputting the target image into an image content classification model, identifying target image features using a feature recognition operator of the image classification model, and performing image content classification based on the target image features to output the target image type.
3. The method of claim 1, wherein, The target image is processed by calling the target processing tool in the target processing page to generate a processing result, comprising: obtaining an image processing requirement of a user; processing the target image using a target processing tool corresponding to the image processing requirement to generate the processing result.
4. The method of claim 3, wherein, The target image is processed by calling the target processing tool in the target processing page to generate a processing result, comprising: inputting the image processing requirement and the target image into a dialogue assistant model to output the processing result.
5. The method of claim 4, wherein, The target image is processed by calling the target processing tool in the target processing page to generate a processing result, comprising: performing intent analysis on the image processing requirement using the dialogue assistant model to determine a user intent; in response to determining that the user intent is image processing, processing the target image based on the user intent to generate a processed image; in response to determining that the user intent is image understanding or image generation, calling a multi-modal model to process the user intent and the target image to generate an intent script.
6. The method of claim 5, wherein, The target image is processed based on the user intent to generate a processed image, comprising: in response to determining that there is an atomic processing algorithm that meets the user intent, calling the atomic processing algorithm that meets the user intent to process the target image to generate the processed image; in response to determining that there is no atomic processing algorithm that meets the user intent, calling an image creation model to process the target image to generate the processed image.
7. The method of claim 5, wherein, The target image is processed based on the user intent to generate a processed image, comprising: in response to detecting a closing instruction of the target processing page after the processed image is generated, closing the target processing page and displaying the processed image in the current conversation record of the dialogue page; in response to detecting a closing instruction of the target processing page before the processed image is generated, closing the target processing page and displaying the target image in the current conversation record of the dialogue page.
8. The method of claim 5, wherein, The target image is processed based on the user intent to generate a processed image, comprising: In response to obtaining a new image processing requirement before the processing image is generated, the new image processing requirement is input to the dialogue assistant model, a new processing image is output, the new processing image is displayed in the current dialogue record of the dialogue page, and after the processing image is generated, the processing image is displayed in the historical dialogue record of the dialogue page.
9. The method of any one of claims 5-8, wherein, The method further includes: Converting the processing image into a format corresponding to the target image type and saving to the local and / or publishing to the target social platform.
10. An image processing apparatus, comprising: a classification module configured to perform image content classification on a target image and determine a target image type; a determination module configured to determine a target processing page corresponding to the target image type, the target processing page comprising at least one target processing tool; a processing module configured to invoke a target processing tool in the target processing page to process the target image and generate a processing result.
11. An electronic device, comprising: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-9.
12. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method of any one of claims 1-9.
13. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-9.