Free training image editing method based on self-supervised learning
Through the self-supervised learning image editing method, the image completion and reconstruction task training model is used to solve the problem of unstable label data dependence and generation quality, and efficient and accurate image editing is achieved, which is suitable for applications such as social media and e-commerce platforms.
Patent Information
- Application Number
- CN202510523899.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-08
AI Technical Summary
The dependence of existing image editing technologies on large amounts of labeled data leads to high labor costs and unstable image quality, especially in complex editing scenarios, which is difficult to meet high-quality needs.
The self-supervised learning method is adopted to train the image editing model through image completion and reconstruction of self-supervised tasks, and combine data normalization and enhancement technology to automatically extract features from unlabeled images and edit them according to user instructions.
Significantly improve image editing efficiency and accuracy, lower technical thresholds, realize personalized editing, and is suitable for scenarios such as social media and e-commerce platforms, and improve work efficiency.
Smart Images

Figure CN120451335A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image editing, and in particular to a free training image editing method based on self-supervised learning. Background Art
[0002] The image editing technology in the prior art has the following defects:
[0003] (1) Dependence on large amounts of labeled data:
[0004] Existing image editing technologies typically require large amounts of labeled data for training and optimization. Data labeling is not only time-consuming and labor-intensive, but also involves high labor costs. Especially for scenarios requiring diverse and high-quality image editing, obtaining labeled data is often a bottleneck, limiting the widespread adoption and application of this technology.
[0005] (2) The quality of generated images is unstable:
[0006] While some existing technologies utilize generative adversarial networks (GANs) for image generation, the quality of the generated images can be unstable in some cases. Especially when the editing content is complex, the generated images may appear distorted, blurry, or unnatural, failing to meet high-quality editing requirements.
[0007] Therefore, it is necessary to provide a free training image editing method based on self-supervised learning to significantly improve the efficiency and accuracy of image editing while lowering the technical threshold for users. Summary of the Invention
[0008] The purpose of the present invention is to provide a free training image editing method based on self-supervised learning to significantly improve the efficiency and accuracy of image editing while lowering the technical threshold for users.
[0009] In order to solve the problems existing in the prior art, the present invention provides a free training image editing method based on self-supervised learning, comprising the following steps:
[0010] S1: Collect data;
[0011] S2: Self-supervised training of image editing models;
[0012] S21: Self-supervised training of image data to form a self-supervised learning model in the following way:
[0013] Self-supervised training sets up a self-supervised image completion task, which requires the image editing model to predict and generate the content of the missing area given a partial image;
[0014] The self-supervised training sets up an image reconstruction self-supervised task, wherein the image reconstruction self-supervised task requires the image editing model to perform multiple transformations on the input image, and then trains the image editing model to restore the original image from the transformed image;
[0015] S22: By understanding and processing the editing instructions, the self-supervised learning model is able to perform editing operations on the image according to the instructions, so that the user's image editing instructions are aligned with the image features generated by the self-supervised learning model.
[0016] Optionally, in the self-supervised learning based free training image editing method, data is collected from public datasets, user-generated content, e-commerce product images and synthetic images.
[0017] Optionally, in the self-supervised learning-based free training image editing method, the changes in the image reconstruction self-supervised task include cutting, rotation and / or blurring.
[0018] Optionally, in the self-supervised learning-based free training image editing method, after completing the image reconstruction self-supervised task, the self-supervised training of the image data further includes the following steps:
[0019] Perform normalization and data augmentation operations on the image.
[0020] Optionally, in the self-supervised learning-based free training image editing method, the image normalization operation is to perform standardization processing on the image to adjust the pixel values to a uniform range;
[0021] Data augmentation is the process of enhancing images by random rotation, cropping, scaling, and / or color transformation.
[0022] Optionally, in the free training image editing method based on self-supervised learning, in S22, the image is edited according to the instruction as follows:
[0023] Modification of the content of the image;
[0024] Make adjustments to the style, structure, and details of your images.
[0025] Compared with the prior art, the present invention has the following advantages:
[0026] (1) The free training image editing method based on self-supervised learning provided by the present invention can significantly improve the efficiency and accuracy of image editing while lowering the technical threshold of users.
[0027] (2) Through self-supervised learning, the system automatically extracts effective features from unlabeled images without requiring a large amount of labeled data, significantly reducing data labeling costs. Users can easily implement personalized image editing through a simple and intuitive interface without having to master complex image processing skills.
[0028] (3) Whether adjusting image color and style or modifying details, the present invention provides rich and flexible editing capabilities while ensuring the overall harmony of the image and the fidelity of its details. Through intelligent command parsing and image feature alignment, the system can accurately execute the user's editing needs, enhancing the personalization and flexibility of image editing.
[0029] (4) The editing process of the present invention is efficient and time-saving, making it particularly suitable for application scenarios such as social media, advertising design, and e-commerce platforms, significantly improving work efficiency. The continuous optimization of self-supervised learning enables the system to continuously improve editing results and has broad adaptability to meet the needs of different fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 A flowchart of a method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0031] The following is a more detailed description of the specific embodiments of the present invention with reference to schematic diagrams. The advantages and features of the present invention will become more apparent from the following description. It should be noted that the drawings are greatly simplified and not to exact scale, and are only used for the purpose of conveniently and clearly illustrating the embodiments of the present invention.
[0032] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.
[0033] The image editing of the existing technology has the following defects: (1) Dependence on a large amount of labeled data: Existing image editing technology usually requires a large amount of labeled data for training and optimization. Data labeling is not only time-consuming and labor-intensive, but also involves high labor costs. Especially for scenarios that require diverse and high-quality image editing, the acquisition of labeled data is often a bottleneck, limiting the popularization and application of the technology. (2) The quality of generated images is unstable: Although some existing technologies use generative adversarial networks (GANs) for image generation, in some cases, the quality of the generated images is unstable. Especially when the editing content is more complex, the generated images may be distorted, blurred or unnatural, and cannot meet the high-quality editing requirements. In order to solve the problems existing in the existing technology, the present invention provides a free training image editing method based on self-supervised learning, such as Figure 1 As shown, the following steps are included:
[0034] S1: Collect data; preferably, data can be collected from public datasets, user-generated content, e-commerce product images, and synthetic images.
[0035] Specifically, public datasets include ImageNet, COCO, ADE20K and other public datasets. These datasets contain a large number of natural images, covering various scenes and object categories, and are suitable for the training of self-supervised learning models.
[0036] User-generated content includes user-generated images collected through social media, online platforms, or mobile devices, such as personal photos and product display images. These images have strong personalized characteristics and can provide more diverse training samples for the model.
[0037] E-commerce product images include: Product images provided by merchants on e-commerce platforms provide important data for self-supervised learning in the e-commerce field. By collecting and processing product images (such as clothing, electronics, and home furnishings), the model can learn how to edit and optimize images based on product characteristics such as category, size, and color.
[0038] Synthetic images, including those generated by computers or using 3D rendering software, can provide models with extremely rich scenarios. These synthetic images can usually be controlled and precisely designed, making them ideal for training specific scenarios, such as objects, scenes, and characters in virtual worlds.
[0039] S2: Self-supervised training of image editing models;
[0040] S21: In the first stage, self-supervised training of image data is performed to form a self-supervised learning model in the following way:
[0041] Self-supervised training involves the self-supervised task of image completion, which requires the image editing model to predict and generate the content of the missing area given a partial image. This task helps the image editing model learn the local structure and context of the image, and further understand how elements in the image are combined to form a complete scene. For example, when inputting a missing face image, the image editing model can predict the missing facial features (such as eyes, nose, etc.).
[0042] Self-supervised training sets up an image reconstruction self-supervised task, which requires the image editing model to perform multiple transformations on the input image (including transformations such as cropping, rotation and / or blurring), and then trains the image editing model to restore the original image from the transformed image; this not only helps the image editing model learn the details and global structure of the image, but also enables the image editing model to capture subtle differences in the image, such as edges, textures, etc.
[0043] The core idea of self-supervised training is to guide the model to learn the underlying features of an image by designing appropriate tasks in the absence of labeled data. The goal of this stage is to enable the image editing model to automatically learn useful image representations from a large number of unlabeled images. These representations will provide a foundation for subsequent image editing tasks. To guide the model to learn meaningful features, the present invention employs two self-supervised tasks: image completion and image reconstruction. Each task forces the image editing model to learn certain structural information or content features of the image, enabling it to generate reasonable output in subsequent image editing.
[0044] To ensure that the image editing model can process various types of image data, the self-supervised training of image data also includes the following steps after completing the image reconstruction self-supervised task:
[0045] Perform image normalization and data augmentation. Image normalization standardizes images, adjusting pixel values to a uniform range (e.g., [0, 1] or [-1, 1]) to ensure consistency in training data. Data augmentation enhances images through random rotation, cropping, scaling, and / or color transformation to increase dataset diversity and improve model robustness.
[0046] S22: The second stage involves understanding and processing the editing instructions, enabling the self-supervised learning model to perform editing operations on the image according to the instructions, so that the user's image editing instructions are aligned with the image features generated by the self-supervised learning model. The image editing operations performed according to the instructions are as follows: (1) modifying the image content; (2) adjusting the image style, structure, and details to ensure that the generated result meets the user's expectations.
[0047] The goal of the second stage is to align the user's image editing instructions with the image features generated by the self-supervised learning model to achieve accurate image editing.
[0048] In summary, the present invention has the following advantages compared with the prior art:
[0049] (1) The free training image editing method based on self-supervised learning provided by the present invention can significantly improve the efficiency and accuracy of image editing while lowering the technical threshold of users.
[0050] (2) Through self-supervised learning, the system automatically extracts effective features from unlabeled images without requiring a large amount of labeled data, significantly reducing data labeling costs. Users can easily implement personalized image editing through a simple and intuitive interface without having to master complex image processing skills.
[0051] (3) Whether adjusting image color and style or modifying details, the present invention provides rich and flexible editing capabilities while ensuring the overall harmony of the image and the fidelity of its details. Through intelligent command parsing and image feature alignment, the system can accurately execute the user's editing needs, enhancing the personalization and flexibility of image editing.
[0052] (4) The editing process of the present invention is efficient and time-saving, making it particularly suitable for application scenarios such as social media, advertising design, and e-commerce platforms, significantly improving work efficiency. The continuous optimization of self-supervised learning enables the system to continuously improve editing results and has broad adaptability to meet the needs of different fields.
[0053] The above description is merely a preferred embodiment of the present invention and does not limit the present invention in any way. Any person skilled in the art who, without departing from the scope of the present invention, makes any equivalent substitution, modification, or other changes to the technical solution and technical content disclosed in the present invention shall be deemed to be within the scope of the present invention and still fall within the scope of protection of the present invention.
Claims
1. A free training image editing method based on self-supervised learning, characterized in that The following steps are involved: S1: Collect data; S2: Self-supervised training of image editing models; S21: Self-supervised training of image data to form a self-supervised learning model in the following way: Self-supervised training sets up a self-supervised image completion task, which requires the image editing model to predict and generate the content of the missing area given a partial image; The self-supervised training sets up an image reconstruction self-supervised task, wherein the image reconstruction self-supervised task requires the image editing model to perform multiple transformations on the input image, and then trains the image editing model to restore the original image from the transformed image; S22: By understanding and processing the editing instructions, the self-supervised learning model is able to perform editing operations on the image according to the instructions, so that the user's image editing instructions are aligned with the image features generated by the self-supervised learning model.
2. The self-supervised learning-based free training image editing method according to claim 1, characterized in that: Data is collected from public datasets, user-generated content, e-commerce product images, and synthetic images.
3. The free training image editing method based on self-supervised learning according to claim 1, characterized in that: Image reconstruction in self-supervised tasks involves variations such as cropping, rotation, and / or blurring.
4. The self-supervised learning-based free training image editing method according to claim 1, characterized in that: After completing the image reconstruction self-supervision task, the self-supervised training of image data also includes the following steps: Perform normalization and data augmentation operations on the image.
5. The free training image editing method based on self-supervised learning according to claim 4, characterized in that: Image normalization is to standardize the image and adjust the pixel values to a uniform range; Data augmentation is the process of enhancing images by random rotation, cropping, scaling, and / or color transformation.
6. The free training image editing method based on self-supervised learning according to claim 1, characterized in that: In S22, the image is edited according to the instruction as follows: Modification of the content of the image; Make adjustments to the style, structure, and details of your images.