Precise image editing and generating method and device, equipment and storage medium
By importing selected images on vector canvas and using human-computer interaction and diffusion models for image editing, the problem of difficulty in realizing accurate editing of complex images in the prior art is solved, and efficient and flexible intelligent image editing is achieved.
Patent Information
- Application Number
- CN202510267197.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
Existing text-based image editing methods are difficult to achieve accurate editing of images containing moving objects, specific shapes or complex scene distributions, reducing the efficiency of using artificial intelligence image drawing.
By importing selected images on vector canvas, selecting the image to be edited using human-computer interaction, and fusion processing with the initial image of the target position through the viewing angle difference method, and combining the diffusion model for content compensation or blank compensation, the precise editing and movement of the image is achieved.
It realizes the function of editing or moving local images in the overall image, improves the efficiency of intelligent drawing, and provides high-precision, flexibility and user-friendly image editing experience.
Smart Images

Figure CN120198524A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, specifically to the field of computer digital painting, and particularly to an image precise editing and generation method, device, equipment, and storage medium. Background Art
[0002] With the rapid development of diffusion model technology, image generation technology has achieved remarkable results in multiple fields. As a creative form of expression, children's painting has also benefited from this technological progress. However, although existing diffusion models can generate diverse and high-fidelity images, precise editing of children's paintings still faces some special challenges.
[0003] Currently, in applications related to children's painting, text-based image editing methods usually can only achieve global modifications and are difficult to meet the editing requirements involving moving objects, specific shapes, or complex scene distributions, reducing the usage efficiency of artificial intelligence image drawing. Summary of the Invention
[0004] Aiming at the technical problem that when the existing human vision test device conducts vision detection on the elderly, it cannot display the colors that the elderly can perceive, a method, device, equipment, and medium for interface color display are provided.
[0005] According to the first aspect, an image precise editing and generation method is provided, and the method includes:
[0006] Import a selected picture based on a created vector canvas to generate an initial image;
[0007] Select the image to be edited in the initial image in a human-computer interaction manner, move the image to be edited to the target position and perform integration processing, including fusing the image to be edited with the initial image at the target position through the perspective difference method, and compensating the content or blank of the initial position area of the image to be edited using a diffusion model;
[0008] Output or save the integrated initial image.
[0009] According to the second aspect, an image precise editing and generation device is provided, including:
[0010] An image generation unit for importing a selected picture based on a created vector canvas to generate an initial image;
[0011] An image fusion and compensation unit for selecting the image to be edited in the initial image in a human-computer interaction manner, moving the image to be edited to the target position and performing integration processing, including fusing the image to be edited with the initial image at the target position through the perspective difference method, and compensating the content or blank of the initial position area of the image to be edited using a diffusion model;
[0012] An image storage unit for integrating and outputting or saving the processed initial image.
[0013] According to a third aspect, there is provided an electronic device, including: one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method of any one of the embodiments in the image precise editing and generation method.
[0014] According to a fourth aspect, there is provided a computer-readable storage medium having a computer program stored thereon, which when executed by a processor, implements the method of any one of the embodiments in the image precise editing and generation method.
[0015] According to the solution of the embodiment of the present application, an initial image is generated by importing a selected picture based on a created vector canvas. In the initial image, the image to be edited is selected in a human-computer interaction manner. The image to be edited is moved to a target position and integrated and processed, including that the image to be edited is fused with the initial image at the target position by the perspective difference method, and the initial position area of the image to be edited is compensated for content or blank by using a diffusion model, so as to realize the function of editing or moving a local image in the overall image, improve the efficiency of intelligent drawing, and combine the powerful function of the diffusion model to achieve a high-precision, flexible and user-friendly image editing experience. Traditional image editing tools often face limitations in local editing and are difficult to perform precise operations in the global context. The proposed technical solution effectively overcomes these limitations. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects and advantages of the present application will become more apparent:
[0017] Figure 1 is an exemplary system architecture diagram to which some embodiments of the present application can be applied;
[0018] Figure 2 is a flowchart of an embodiment of the image precise editing and generation method according to the present application;
[0019] Figure 3 is a schematic diagram before and after specifically fusing pictures according to the image precise editing and generation method of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The following describes exemplary embodiments of the present application in conjunction with the accompanying drawings. Various details of the embodiments of the present application are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0021] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0022] Figure 1 An exemplary system architecture 100 is shown, which can apply embodiments of the method or device for accurately editing and generating images of the present application.
[0023] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0024] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the terminal devices 101, 102, 103, such as video applications, live broadcast applications, instant messaging tools, email clients, social platform software, etc.
[0025] The terminal devices 101, 102, 103 here can be hardware or software. When the terminal devices 101, 102, 103 are hardware, they can be electronic devices, including but not limited to smartphones, tablets, laptop portable computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they can be installed in the above-listed electronic devices. It can be implemented as multiple software or software modules (such as multiple software or software modules for providing distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.
[0026] The server 105 can be a server that provides various services, such as a background server that provides support for the terminal devices 101, 102, 103. The background server can collect image data, process it, and feedback the processing results to the terminal devices.
[0027] It should be noted that the method provided by the embodiments of the present application can be executed by the server 105 or the terminal devices 101, 102, 103. Correspondingly, the image precise editing and generating device can be set in the server 105 or the terminal devices 101, 102, 103.
[0028] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in
[0029] Continuing to refer to Figure 2 , a flowchart of an embodiment of the image precise editing and generating method according to the present application is shown. The image precise editing and generating method includes the following steps:
[0030] S201: Import a selected picture based on the created vector canvas to generate an initial image;
[0031] S202: Select the image to be edited in the initial image in a human-computer interaction manner. The image to be edited is moved to the target position and integrated and processed, including fusing the image to be edited with the initial image at the target position by the perspective difference method, and performing content compensation or blank compensation on the initial position area of the image to be edited using a diffusion model;
[0032] S203: Output or save the integrated and processed initial image.
[0033] In the traditional method: It is very difficult to meet the local editing requirements in the existing image processing methods because the traditional methods often rely on regional masks to modify specific parts, but it is very difficult to maintain the harmony and beauty of the overall picture. However, the method of the present invention can perform local editing on specific objects in a painting work and at the same time retain the natural layout of the overall scene. Children's paintings are often full of whimsy. Therefore, when processing their creations, local editing can make operations such as moving objects and adding details simple and straightforward. This flexibility not only improves the creative freedom of children but also encourages them to express their ideas boldly.
[0034] In the aspect of basic image generation, accurately conveying the spatial positions of various elements in children's ideal paintings is a challenge. Many times, when children describe their creativity, it may be difficult for them to express the specific relationships between objects through language. Generating children's paintings through serializable layouts helps them better achieve the imagined picture layout, thereby enhancing their creative experience. The direct manipulation interface is of great significance in the interaction between children and the image editing system. Traditional natural language interfaces face certain difficulties in the precise editing tasks of children's paintings because children often have difficulty avoiding ambiguity when expressing their ideas. However, by combining natural language instructions with direct manipulation, a more intuitive editing experience can be provided for children, enabling them to easily adjust, modify, and recreate specific objects during the creative process.
[0035] In this embodiment, the execution entity (such as Figure 1 the server or terminal device shown) on which the image precise editing generation method runs generates an initial image, and then performs integration processing on the selected image to be edited. The integration processing includes, first, the fusion after the movement of the image, and second, the supplementation or filling of the blank area formed at the initial position after the movement of the image to be edited.
[0036] For the method provided in the above embodiment of the present application, importing the selected picture based on the created vector canvas in S201 to generate the initial image includes the following steps:
[0037] The selected images in the image set are imported into the vector canvas to generate an initial image. The image set pre-exists in the database. The image set includes pictures of different types or attributes. The database updates the image set in real-time or periodically through a port, or, downloaded images are imported into the vector canvas to generate an initial image. The downloaded images are pictures downloaded from the Internet or local area network through a download port. It should be noted that: A Vector Canvas generally refers to creating an editable canvas based on vector graphics, where the elements can be scaled arbitrarily without losing clarity. Vector graphics are composed of points, lines, paths, and shapes, which are different from pixel maps (dot-matrix-based images) and are suitable for illustration, logo design, and other graphic design tasks. For example, using AI to generate a vector canvas, diffusion models and other generative models can generate vector graphics according to text descriptions. These graphics can be adjusted through automated tools or manual design. The graphics generated by AI can be exported in vector formats such as SVG. Another example is using a web application to generate a vector canvas. Through a simple online tool, vector graphics can be directly drawn, edited, and exported in a browser. For example, Figma is a popular browser-based design tool that supports the design and collaboration of vector graphics; Vectr: A free online vector graphics design tool that supports creating and editing vector canvases and can support layering and coordinate systems, enabling each pixel operation of the user to correspond precisely. Users can freely drag and adjust the position of images on the canvas. The entire operation interface is flexible and intuitive. The images selected by the user can automatically match and adapt to the size of the canvas, ensuring their integrity in the component hierarchy and avoiding image distortion or disproportion. At the same time, expand the original source of the pictures and obtain basic pictures through the network to avoid limitations on the types and attributes of pictures.
[0038] Furthermore, the method of the present invention can play an auxiliary role in the treatment of autistic children. Autistic children are usually more sensitive to visual information. Editable pictures can help autistic children understand and express their needs, emotions, and thoughts. By showing social scenarios through pictures, it helps them learn appropriate social behaviors and responses, and can also help autistic children better understand abstract concepts and chronological order, such as daily activity arrangements, and promote emotion management. By showing different emotions through pictures, it helps them recognize and understand their own and others' emotions. Specifically,
[0039] 1. Apply the method of the present invention in an emotion recognition training system
[0040] Based on the precise image editing technology of the present invention, a dynamic expression database is constructed to generate more than 2,000 micro-expression features (such as eye expressions and mouth corner arcs). 3D virtual characters with controllable intensity gradients (20%-100%) are generated through AI to guide children to recognize emotional changes. The existing system is used to capture children's facial expression feedback in real time, dynamically adjust the expression complexity, and form a personalized learning curve to facilitate doctors' analysis of children's treatment effects.
[0041] 2. Apply the method of the present invention to a social scenario simulation system
[0042] Use image generation technology to create an interactive virtual scene, supporting custom light (50 - 6500K color temperature) and spatial layout parameters. Generate a social scene with 8 - 12 people, and achieve editable interaction of objects through semantic segmentation technology. For example, adjust the body language parameters of virtual characters (such as gesture angles ±15°) to train children to understand social distances.
[0043] 3. Strengthen the feedback mechanism
[0044] Develop a real-time feedback engine in combination with the Generative Adversarial Network (GAN). When a child completes the interaction, a reward animation (with a precision of 512×512px) is generated within 0.5 seconds. The system analyzes the child's attention heat map (sampling rate 120Hz) through the Transformer architecture and automatically optimizes the subsequent training content.
[0045] The method provided in the above embodiments of the present application includes three parts in S202, where
[0046] The first part: The specific implementation manner of selecting the image to be edited in the initial image in a human-computer interaction manner. For example, receive the user's voice input or mouse click on the image to be edited in the initial image in a human-computer interaction manner to generate a selection box. The selection box can cover or include the image to be edited. Preferably, the selection box is configured with a magnifying icon. Clicking on the magnifying icon can magnify the selection box to ensure that the selection box includes the selected image to be edited. Generally, the selection box is a quadrilateral box, which is simple to generate and ensures the rapidity of selection, that is, the box selection of the image. The methods include:
[0047] Rectangular Selection is the most common selection method. By clicking and dragging the mouse, a rectangular area is created to select a part of the image. This method is simple and intuitive and is applicable to most scenarios. Or, Circular / Elliptical Selection is similar to Rectangular Selection, but the selected area is circular or elliptical. This is usually used when a circular or symmetric shape needs to be selected, such as selecting circular or elliptical objects, or creating an irregular selection area for artistic effects. Or, the Lasso Tool is similar to Polygonal Selection, but it is more free. The user can freely draw the selection area by hand to flexibly select any shaped area. Or, the Magic Wand Tool is a tool that automatically selects adjacent areas based on color similarity. When the user clicks on a point in the image, the Magic Wand Tool will automatically select the area with a color similar to that point to form an "intelligent" selection area. It can also be the Quick Selection Tool or Path Selection Tool in the prior art or other software or tools.
[0048] Furthermore, in order to simplify the computational workload of subsequent image editing calculations and reduce the complexity of the calculations, matte extraction processing is required. For example, the selection box calls an external multi-modal model through the API port for OCR recognition pre-training to perform matte extraction processing on the image to be edited, so that the shape of the selection box can match the outer contour shape of the image to be edited. For example, the initial image includes a coral in the middle of the canvas and a fish at the bottom of the canvas. Now, taking the "fish" as the selected image to be edited, an arc-shaped box that fits the shape of the "fish" needs to be generated. Through OCR recognition pre-training, the computer can recognize the fish and simulate its outer contour. Only the selected "fish" is taken as the image to be edited, and matte extraction processing is performed to reduce the computational workload of subsequent image editing and grayscale processing and reduce the computational cost of the hardware.
[0049] The second part is the filling or supplementing of the blank area. After the image to be edited is dragged or moved, a blank area is formed in the initial position area of the image to be edited. The initial position area of the image to be edited uses a diffusion model for content compensation or blank compensation. Specifically,
[0050] Obtain the feature parameters of all images within a specified range around the image to be edited, and use an extended model to perform an image diffusion algorithm calculation to fill the blank area with images and colors. The feature parameters include color, hue, and image depth information. The purpose is that through intelligent reasoning and context understanding, the image diffusion algorithm can generate natural and rich content for the blank area, providing a higher-quality filling effect than traditional image editing methods. Specifically, the diffusion algorithm can analyze the color information of the surrounding area and perform appropriate color filling, not just simple color copying, but can intelligently generate a gradient that is consistent with the color of the surrounding area, ensuring that the filled area matches the hue of the original image. For example, in the reconstruction of the sky, ocean, or light and shadow effects, it can ensure natural color transitions and avoid abrupt boundaries.
[0051] The third part: The image to be edited is fused with the initial image at the target position through the perspective difference method, including: determining the target position. Specifically, when the user stops the movement or dragging of the image to be edited in a human-computer interaction manner, the final position of the image to be edited is used as the target position, and the fusion is performed at the target position.
[0052] Obtain the picture depth, light and shadow relationship, and color information of the original image at the target position, etc.
[0053] The picture depth, color information, and light and shadow relationship are used to simulate the image based on the human perspective difference.
[0054] Convert the simulated image into a transition image with color depth ranging from white to black. The transition image is used to present the three-dimensional sense of the image. Specifically, based on the binocular perspective difference of a person, the area where the angle formed by the distance to the target position and the binoculars is less than the threshold is used as the far transition image, and the area where the angle formed by the distance to the target position and the binoculars is greater than or equal to the threshold is used as the near transition image. The threshold is selected according to factors such as the age of the population or other factors.
[0055] Perform gradient gray processing on the near transition image to form a depth image. Specifically, within the range of 0° - 180°, for every 1° decrease in the binocular perspective difference, the gray value is increased by 1. That is, 0° represents a completely black image, and 180° represents a completely white image. The depth image can present a three-dimensional sense and avoid excessive drift of the three-dimensional image. For example, if the situation of the picture floating occurs, it is likely to cause visual dizziness.
[0056] Adjust the light and shadow relationship of the image to be edited at the target position to be consistent with the light and shadow relationship of the original image at the target position, and adaptively match the gray value of the depth image with the picture depth of the original image at the target position. For example, the gray change does not exceed 1 / 180.
[0057] Summary of the technical solutions of the above multiple specific embodiments: The core idea is to use a diffusion model to intelligently process the user's editing operations during image editing, receive multi-modal instructions from the user through a user interface, process these instructions by serializing the multi-modal instructions into text form, and then use layout-based image generation. The multi-modal model follows the input image and multi-modal instructions and generates a transformed image that solves the instructions. The method of the present invention does not directly transform the image using the diffusion model, but manipulates a part of the structure of the image in the form of a spatial layout, representing the object specified by the bounding box and text description, converts the multi-modal instructions into text form. Suppose the user wants to move a specific object in a combination, a complex scenario. The user can pass a bounding box, which can be represented in text form, such as "{x:150,y:400,width:100,height:100}", and the destination is represented as "{x:144,y:132}". By inputting the instructions into the generated transformed layout and then providing it to the layout-to-image generation system, such as GLIGEN (GLIGEN (Grounded-Language-to-Image Generation) is a new type of image generation method. It adds support for localization inputs on the basis of existing pre-trained text-to-image diffusion models. Specifically, GLIGEN injects different localization conditions (such as bounding boxes, key points, etc.) of the input into the new trainable layer through a gating mechanism, so as to achieve control over open-world image generation. The advantage of this method is that it retains a large amount of conceptual knowledge of the pre-trained model while improving the controllability and accuracy of the generated images. GLIGEN's zero-shot performance on the COCO and LVIS datasets is significantly better than the existing supervised layout-to-image baselines. In addition, GLIGEN supports multiple input methods, including text entity + box, image entity + box, image style + text + box, and text entity + key point.
[0058] User Interface In addition to developing an LLM-based system for handling multimodal user commands, we aimed to develop a simple user interface that does not require familiarity with complex graphic design software. Rather than a complex interface with multiple layers of drop-down menus like WIMP, we only have five tools in the toolbar, an interactive canvas, and a command input form. The tools supported by the interface are: selection operations for selecting objects to draw, a bounding box tool, a star tool for specifying 2D points on the canvas, and a reload tool for regenerating the latest layout. Despite its simplicity, users can perform a wide array of image manipulations. Instead of having pre-specified tools for performing operations such as moving objects, adding objects, or changing the appearance of an object, we allow users to specify these transformations using flexible language commands. Users can easily specify geometric objects (such as bounding boxes) using direct manipulation, which can be used as "parameters" for these natural language editing commands. As the user draws each object, it appears in the description text box with a reference symbol, just like the words in the instruction text box. For layout-based image generation, we utilized GLI-GEN, which is built on top of stable diffusion and has the following steps:
[0059] 1. Create a Canvas
[0060] Goal: Create a new canvas, which is the basis for the entire operation.
[0061] Content: The canvas is filled with coordinates to ensure that every pixel corresponds to an exact coordinate. Because the canvas is vector, it can be resized dynamically, and no matter how the user stretches the canvas, the coordinate system can expand accordingly.
[0062] 2. Import the image
[0063] Goal: Drag the image selected by the user into the canvas.
[0064] Content: The user can select an image and drag it into the newly created canvas. The system will automatically fit the image's size to fit the canvas's boundaries while maintaining the integrity of the image.
[0065] 3. Select the editing scope
[0066] Goal: To perform partial editing on an imported image.
[0067] Content: The user uses the mouse to select a specific range on the image, which corresponds to the coordinate system on the canvas. After the user selects the range, the range is clearly positioned on the canvas, which is convenient for subsequent processing.
[0068] 4. Drag and drop editing
[0069] Goal: Move the selected image area to a new location.
[0070] Content: The user drags the selected range from the original position (point A) to the target position (point B). During this process, the system needs to apply the capabilities of the context model to analyze and process the surrounding content.
[0071] 5. Compensation and Fusion of Vacant Areas
[0072] Goal: Ensure that during the dragging process, the vacant part is naturally filled and seamlessly fused with the original content.
[0073] Content:
[0074] When the selected area moves from point A to point B, a vacancy is formed at point A. At this time, the system will utilize the context learning ability to analyze the surrounding content and spread it to fill this vacancy.
[0075] Meanwhile, when the selected area lands at point B, point B and its surrounding areas will be fused with the dragged image. The system will analyze the content at point B and diffusively fuse the dragged image with the existing parts around it to ensure a natural and harmonious overall effect.
[0076] 6. Completing the Editing
[0077] Goal: Complete the entire image editing process and ensure that the final result maintains the integrity and natural beauty of the image.
[0078] Content: Through the above steps, the local editing of the image is achieved, the vacant part is compensated and fused, and finally the user's dragging operation is completed, realizing precise and fitting image editing.
[0079] In this process, the core points lie in two aspects: Firstly, establish an accurate and adjustable coordinate system canvas; Secondly, achieve the compensation and fusion of content through the context learning ability to meet the user's needs during the image editing process.
[0080] The main beneficial effects achieved by this process are as follows:
[0081] Firstly, with the support of the diffusion model and LLM, the specific needs of users can be intelligently responded to during the editing process, improving the accuracy of editing. At the same time, combined with an intuitive interface design, users can easily perform image editing operations, greatly reducing the usage threshold. Through this technological innovation, the application of the diffusion model enables the system to generate natural transition effects in a dynamic environment (such as adjusting the image layout), avoiding obvious splicing marks caused by traditional editing methods. In the application of this technology, users can freely adjust the position and content of the image, and the system can automatically analyze the context to ensure the consistency of the overall visual effect and stimulate the creative potential of users.
[0082] Application Examples
[0083] Example 1: Moving an Object in an Image
[0084] User Operations:
[0085] The user imports an image containing a fish and a coral.
[0086] The user uses the mouse to select the position of the fish and drags it to the upper left corner of the canvas.
[0087] The user enters the natural language instruction: "Move the fish to the upper left corner."
[0088] System Processing:
[0089] 1. The system records the coordinates of the selected area (x: 150, y: 400, width: 100, height: 100).
[0090] 2. The system combines the drag operation and the natural language instruction to generate a multimodal instruction:
[0091] {
[0092] "operation": "move",
[0093] "object": "fish",
[0094] "source": {"x": 150, "y": 400, "width": 100, "height": 100},
[0095] "destination": {"x": 50, "y": 50}
[0096] }
[0097] 3. The system uses a diffusion model to generate compensatory content to fill the gap in the original position of the fish.
[0098] 4. The system moves the fish to the target position and seamlessly integrates it with the surrounding content.
[0099] Final Effect: The fish is moved to the upper left corner of the canvas, the gap in the original position is naturally filled, and the overall image remains harmonious.
[0100] Example 2: Replacing an Object in an Image
[0101] User Operations:
[0102] The user imports an image containing a cat and a sofa.
[0103] The user uses the mouse to select the position of the cat and enters the natural language instruction: "Replace the cat with a dog."
[0104] System processing:
[0105] 1. The system records the coordinates of the boxed area (x: 200, y: 300, width: 120, height: 120).
[0106] 2. The system combines the box selection operation with natural language instructions to generate multimodal instructions:
[0107]
[0108] 3. The system uses a diffusion model to generate an image of a dog and replaces it at the position of the cat.
[0109] 4. The system seamlessly integrates the dog's image with the surrounding content, as shown in Figure 3 shown;.
[0110] Final effect: The cat is replaced by a dog, and the replaced image remains natural and harmonious.
[0111] Specific function implementation
[0112] 1. Dynamic canvas construction
[0113]
[0114]
[0115] 2. Image import and adaptation
[0116]
[0117] 3. Precise area selection and dragging
[0118]
[0119]
[0120] 4. Content compensation and integration
[0121]
[0122]
[0123] 5. def process_instruction(instruction, canvas):
[0124]
[0125] Application example: Resize an object in an image
[0126] Scenario description:
[0127] The user imports an image containing a tree and grassland and hopes to double the size of the tree. The user selects the position of the tree by drawing a box with the mouse and enters the natural language instruction: "Double the size of the tree."
[0128] System processing flow:
[0129] 1. Image import and adaptation:
[0130] After the user imports the image, the system automatically adjusts the size and proportion of the image to ensure that it remains clear when displayed on the canvas.
[0131] The image is converted into a vector format for subsequent editing operations.
[0132] 2. Precise area selection:
[0133] The user uses the mouse to select the position of the tree, and the system records the coordinates of the selected area (x: 300, y: 200, width: 150, height: 200).
[0134] The system obtains the selected area through the following function:
[0135]
[0136] 3. Generate multimodal instructions:
[0137] The system combines the selection operation and the natural language instruction to generate multimodal instructions:
[0138]
[0139] 4. Resize the object:
[0140] The system uses a diffusion model to generate an image of the enlarged tree and replaces it in the original position.
[0141] The system resizes the object through the following function:
[0142]
[0143] 5. Content compensation and fusion:
[0144] The system seamlessly fuses the enlarged tree with the surrounding content to ensure a natural and harmonious visual effect of the overall image.
[0145] The system performs content compensation and fusion through the following function:
[0146]
[0147]
[0148] Final effect:
[0149] The tree is magnified two times, and the magnified tree naturally blends with the grassland and the surrounding environment, and the overall image remains harmonious.
[0150] Beneficial effects: This technical solution has multiple beneficial effects, making it have significant advantages in the field of image processing:
[0151] 1. High-precision editing: With the support of machine learning algorithms, the specific needs of users can be intelligently responded to during the editing process, improving the accuracy of editing.
[0152] 2. User-friendliness: Combined with an intuitive interface design, users can easily perform image editing operations, greatly reducing the usage threshold, and are suitable for users at different levels, especially in fields such as children's painting or creative design.
[0153] 3. Intelligent content generation: The application of the diffusion model enables the system to generate natural transition effects in a dynamic environment (such as adjusting the image layout), avoiding obvious splicing traces caused by traditional editing methods.
[0154] 4. Enhanced creative flexibility: Users can freely adjust the position and content of the image, and the system can automatically analyze the context to ensure the coordination of the overall visual effect, stimulating the creative potential of users.
[0155] Overall: It overcomes the limitations of traditional image editing tools, not only improving the accuracy and flexibility of editing, but also greatly improving the user experience, providing a more user-friendly creative platform for users.
[0156] In the second aspect, referring to an image precise editing and generating device, the device includes:
[0157] An image generation unit, configured to import a selected picture based on a created vector canvas to generate an initial image;
[0158] An image fusion and compensation unit, configured to select an image to be edited in the initial image in a human-computer interaction manner, move the image to be edited to a target position and perform integration processing, including fusing the image to be edited with the initial image at the target position by the perspective difference method, and performing content compensation or blank compensation on the initial position area of the image to be edited by using the diffusion model;
[0159] An image storage unit, configured to output or save the integrated initial image.
[0160] In some optional implementation manners of this embodiment, at least the following two functions are provided in the above device, including:
[0161] After the image to be edited is dragged or moved, a blank area is formed in the initial position area of the image to be edited. The initial position area of the image to be edited is compensated for content or blank using a diffusion model. Specifically,
[0162] Obtain the feature parameters of all images within a specified range around the image to be edited, perform an image diffusion algorithm calculation using an extended model, and fill the blank area with images and colors. The feature parameters include color, hue, and image depth information.
[0163] The image to be edited is fused with the initial image at the target position through the perspective difference method, including:
[0164] Obtain the screen depth, light and shadow relationship, and color information of the original image at the target position;
[0165] The screen depth, color information, and light and shadow relationship are used to simulate the image from the perspective difference of a person;
[0166] Convert the simulated image into a transitional image with a color depth ranging from white to black. The transitional image is used to present the three-dimensional sense of the image. Among them, based on the binocular perspective difference of a person, those with an angle formed by the distance to the target position and the binoculars less than the threshold are used as far transitional images, and those with an angle formed by the distance to the target position and the binoculars greater than or equal to the threshold are used as near transitional images;
[0167] Perform gradient gray processing on the near transitional image to form a depth image. Among them, within the range of 0° - 180°, for every 1° decrease in the binocular perspective difference, the gray value increases by 1.
[0168] According to an embodiment of the present application, the present application also provides an electronic device and a readable storage medium.
[0169] Block diagram of an electronic device for an image precise editing generation method according to an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0170] The electronic device includes: one or more processors, a memory, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. The various components are interconnected using different buses and can be mounted on a common motherboard or otherwise mounted as required. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory for displaying graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, if necessary, multiple processors and / or multiple buses can be used in conjunction with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, with each device providing part of the necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system).
[0171] The memory is the non-transitory computer-readable storage medium provided by this application. Among them, the memory stores instructions executable by at least one processor to enable at least one image precise editing generation method. The non-transitory computer-readable storage medium of this application stores computer instructions, and these computer instructions are used to cause a computer to execute the image precise editing generation method provided by this application.
[0172] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the interface color display method in the embodiments of this application. The processor executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory, that is, implements the image precise editing generation method in the above method embodiments.
[0173] The memory can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the image precise editing generation method, etc.
[0174] The electronic device may further include: an input device and an output device. The processor, the memory, the input device, and the output device can be connected via a bus or other means.
[0175] The input device can receive input digital or character information, as well as generate key signal inputs related to user settings and function control of an electronic device for the test method of the Timed Up and Go Test, such as input devices like touchscreens, keypads, mice, trackpads, touchpads, pointing sticks, one or more mouse buttons, trackballs, joysticks, sensors that can capture human motion information and / or physiological information, etc. The output device can include display devices, auxiliary lighting devices (e.g., LEDs), and tactile feedback devices (e.g., vibration motors), etc. The display device can include, but is not limited to, liquid crystal displays (LCDs), light-emitting diode (LED) displays, and plasma displays. In some embodiments, the display device can be a touchscreen, a head-mounted display (HMD).
[0176] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuit systems, application specific ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0177] These computational programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disks, optical disks, memories, programmable logic devices (PLDs)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0178] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0179] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0180] A computer system can include a client and a server. The client and the server are generally far apart from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other.
[0181] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, which contains one or more executable instructions for implementing a specified logical function.
[0182] The units involved in the embodiments described in the present application can be implemented in software or in hardware. The described units can also be provided in a processor, for example, it can be described as: a processor includes a plurality of computing units. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, a parameter calculation unit can also be described as "a calculation unit for image parameters".
[0183] As another aspect, the present application also provides a computer-readable medium, which may be included in the device described in the above embodiments; or may exist separately without being assembled into the device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the device, the device is configured to: select an image to be edited in the initial image in a human-computer interaction manner, move the image to be edited to a target position and perform integration processing, including fusing the image to be edited with the initial image at the target position by using the perspective difference method, and compensating the content or blank of the initial position area of the image to be edited by using a diffusion model.
[0184] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present application.
Claims
1. A method for accurately editing and generating an image, the method comprising: Import the selected picture based on the created vector canvas to generate the initial image; An image to be edited is selected from the initial image in a human-computer interactive manner, and the image to be edited is moved to a target position and integrated, including fusing the image to be edited with the initial image at the target position by a perspective difference method, and performing content compensation or blank compensation on the initial position area of the image to be edited by using a diffusion model; The processed initial image is integrated and output or saved.
2. The image accurate editing and generating method according to claim 1, wherein: Import the selected image based on the created vector canvas to generate the initial image, including: An image selected from an image set is imported into the vector canvas to generate an initial image. The image set is pre-stored in a database. The image set includes pictures of different types or attributes. The database updates the image set in real time or regularly through a port. Alternatively, a downloaded image is imported into the vector canvas to generate an initial image. The downloaded image is a picture downloaded from the Internet or a local area network through a download port.
3. The image accurate editing and generating method according to claim 1, wherein: Selecting an image to be edited from the initial image in a human-computer interactive manner includes: The user selects the image to be edited in the initial image by voice input or by mouse click in a human-computer interaction manner, and generates a filter box, wherein the filter box can cover the image to be edited.
4. The image accurate editing and generating method according to claim 3, wherein: Generates a filter box, including: The filter box is configured with a magnifying badge, and the filter box can be enlarged by clicking the magnifying badge; and / or the filter box calls an external multimodal model through an API port to perform OCR recognition pre-training, and performs cutout processing on the image to be edited, so that the shape of the filter box can match the outer contour shape of the image to be edited.
5. The image accurate editing and generating method according to any one of claims 1 to 4, wherein: After the image to be edited is dragged or moved, the initial position area of the image to be edited forms a blank area, and the initial position area of the image to be edited uses a diffusion model to perform content compensation or blank compensation, including: The characteristic parameters of all images within the specified range outside the image to be edited are obtained, and the image diffusion algorithm is calculated using the extended model to fill the blank area with images and colors. The characteristic parameters include color, hue and image depth information.
6. The image precise editing and generating method according to any one of claims 1 to 4, wherein the image to be edited is fused with the initial image of the target position by using a perspective difference method, comprising: Determine the target location, wherein When the user stops moving or dragging the image to be edited in a human-computer interaction manner, the final position of the image to be edited is used as the target position, and fusion is performed at the target position.
7. The image accurate editing and generating method according to claim 6, wherein: Performing fusion at the target location includes: Obtaining picture depth, light and shadow relationship, and color information of the original image at the target position; The image depth, color information and light-shadow relationship are simulated by using the human visual angle difference; The simulated image is converted into a transition image with a color depth from white to black, wherein the transition image is used to present the stereoscopic sense of the image, wherein, based on the difference in binocular visual angle of a person, a far transition image is formed when the angle between the target position and the binocular is less than a threshold, and a near transition image is formed when the angle between the target position and the binocular is greater than or equal to the threshold; The near-transition image is subjected to gradient grayscale processing to form a depth image, wherein within the range of 0°-180°, the grayscale value increases by 1 for every 1° decrease in binocular visual angle difference; The light and shadow relationship of the image to be edited at the target position is adjusted to be consistent with the light and shadow relationship of the original image at the target position, and the grayscale value of the depth image is adaptively matched with the picture depth of the original image at the target position.
8. An image accurate editing and generating device, the device comprising: An image generation unit, used for generating an initial image by importing a selected picture based on the created vector canvas; An image fusion compensation unit, used for selecting an image to be edited from the initial image in a human-computer interactive manner, moving the image to be edited to a target position and integrating the image, including fusing the image to be edited with the initial image at the target position by a perspective difference method, and performing content compensation or blank compensation on the initial position area of the image to be edited by using a diffusion model; The image storage unit is used to integrate the processed initial image for output or storage.
9. A device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Intelligent image processing method and device, computer equipment and storage medium
CN121095394A
High-efficiency multi-round image editing method for high-speed-low-speed tool path proxy
CN121236229A