An artificial intelligence-based image editing system, method, device, and medium
The AI-based image editing system enables automated segmentation, generation, and ambient color matching of image elements, solving the problems of cumbersome operation and insufficient security in existing technologies. It improves the efficiency and quality of image editing and meets the diverse needs of enterprises.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QINGDAO NANLUODAO TECHNOLOGY DEVELOPMENT CO LTD
- Filing Date
- 2026-01-15
- Publication Date
- 2026-06-02
Smart Images

Figure CN122134856A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an image editing system, method, device and medium based on artificial intelligence. Background Technology
[0002] With the large-scale application of Artificial Intelligence Generated Content (AIGC) technology, the digital content creation industry is transforming from "human-led" to "intelligence-driven." Image scene element replacement and atmosphere optimization have become core requirements for improving creative efficiency and enriching content expression in industries such as advertising, film and television, gaming, and e-commerce. Enterprises and creators are placing higher demands on image editing in terms of efficiency, personalized effects, scene adaptability, and data security.
[0003] Current image scene element replacement technologies can be broadly categorized into three types: manual editing, automated replacement based on traditional computer vision, and AIGC-assisted replacement. However, existing technologies cannot automatically recommend suitable replacement elements based on scene semantics. Users must manually select appropriate element resources, resulting in a cumbersome and inefficient process that fails to meet the demands of large-scale, rapidly iterating, and highly secure enterprise-level business needs. Summary of the Invention
[0004] To address the aforementioned issues, this application proposes an artificial intelligence-based image editing system, comprising: a semantically enhanced segmentation module for preprocessing the input image and segmenting the preprocessed image to obtain target element contour masks and subject protection masks; a platform-based AIGC generation module for constructing composite instructions based on pre-set multi-dimensional constraints, performing model adaptation based on the composite instructions, and generating a verification iteration mechanism to determine image elements; a multi-dimensional atmosphere collaborative adaptation module for fusing the generated image elements with the original image through atmosphere depth perception and adaptive adjustment; and a full-link automation and batch processing module for performing end-to-end automated processing and batch optimization based on a pre-set workflow engine.
[0005] In one example, the semantically enhanced segmentation module includes: an image preprocessing unit, which removes noise from the image using a pre-set noise reduction algorithm and unifies the resolution and color space of the image; divides the image and calculates the local variance and average gradient of the divided image to determine smooth regions and edge regions based on the local variance and average gradient; applies Gaussian filtering and median filtering with corresponding parameters to each region to preserve edge information; a fusion segmentation model unit, which performs deep fusion and pixel segmentation using a pre-set segmentation architecture; and an optimization unit, which eliminates edge jaggedness and outputs a mask through morphological operations and Bézier curve smoothing.
[0006] In one example, the platform-based AIGC generation module includes: a multi-dimensional constraint instruction construction unit, which parses multi-dimensional constraints according to a pre-set instruction parsing engine, the multi-dimensional constraints including scene feature constraints, user requirement constraints, and segmentation mask constraints, and generates composite instructions based on the multi-dimensional constraints; a model adaptation unit, which guides scene feature vectors, limits mask boundaries, and transforms requirement parameters through a pre-set AIGC model; and a generation verification iteration unit, which determines the matching degree between image elements and multi-dimensional constraints through a similarity algorithm, and automatically iterates based on the matching degree.
[0007] In one example, the multi-dimensional atmosphere collaborative adaptation module includes: an atmosphere depth perception unit, which extracts features from the image, including color, lighting, texture, and semantic atmosphere features, the semantic atmosphere features including scene type, emotional tone, and semantic association, and generates an atmosphere vector based on the features; an adaptive adjustment unit, which performs color adaptation, lighting adaptation, texture adaptation, and semantic adaptation on the image based on the atmosphere vector; and a global optimization unit, which performs color balance, sharpening, and contrast adjustment on the image to ensure that the image elements are completely integrated with the original image.
[0008] In one example, the end-to-end automation and batch processing module includes: an end-to-end automation unit, which performs an end-to-end automated process through a pre-set workflow engine, the automated process including preprocessing, segmentation, generation, replacement, adaptation, and output; a batch optimization unit, which introduces a platform-based parallel computing framework to perform batch processing of the images; a batch adaptation unit, which is used to balance batch unified parameters and individual personalized instructions; and a data security unit, which is used to encrypt the storage and transmission of original images and generated data to ensure data security.
[0009] In one example, it also includes a feedback-driven iteration module for collecting user feedback information to perform iterative optimization based on the feedback information.
[0010] In one example, the feedback-driven iteration module includes: a feedback collection unit, which collects user feedback information on the processing results, including scores, text modification suggestions, and labeled areas; a model optimization unit, which incorporates the feedback data into a training set and iteratively optimizes the segmentation model, AIGC-generated constraint coefficients, and adaptation algorithm thresholds based on the training set; and a template expansion unit, which automatically generates industry templates so that users can directly call the industry templates for batch processing.
[0011] On the other hand, this application also proposes an AI-based image editing method, applied to an AI-based image editing system as described in any of the examples above. The method includes: preprocessing the input image and performing image segmentation on the preprocessed image to obtain a target element contour mask and a subject protection mask; constructing a composite instruction based on pre-set multi-dimensional constraints, performing model adaptation based on the composite instruction, and generating a verification iteration mechanism to determine image elements, wherein the composite instruction includes semantic text, feature vectors, and mask boundaries; extracting features from the preprocessed image, determining an atmosphere vector based on the extracted features, and adaptively adjusting the image based on the atmosphere vector to fuse the generated image elements with the original image; performing end-to-end automated processing and batch optimization based on a pre-set workflow engine; and collecting user feedback information for iterative optimization based on the feedback information.
[0012] On the other hand, this application also proposes an artificial intelligence-based image editing device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the artificial intelligence-based image editing device to perform: the method described in the example above.
[0013] On the other hand, this application also proposes a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured as described in the example above.
[0014] This application's AI-based image editing system features a semantically enhanced segmentation module. This module unifies image resolution and color space through an image preprocessing unit, accurately dividing smooth and edge regions while preserving edge information. A fusion segmentation model unit achieves deep fusion and pixel segmentation, while an optimization unit eliminates jagged edges, accurately obtaining the outline of target elements and the subject's protective mask. A platform-based AIGC generation module constructs multi-dimensional constraint instructions. Through a model adaptation unit, scene feature vectors are guided, mask boundaries are defined, and required parameters are transformed. A verification iteration unit is then generated to ensure a high degree of matching between image elements and constraints. A multi-dimensional atmosphere collaborative adaptation module first deeply perceives the atmosphere to generate vectors, then adaptively adjusts colors, lighting, etc., and finally performs global optimization to perfectly blend elements with the original image. A full-link automation and batch processing module implements an end-to-end automated process, introducing a parallel computing framework to support batch processing, balancing unified parameters and personalized instructions, and ensuring data security. Furthermore, a feedback-driven iteration module collects user feedback, incorporates the data into the training set to optimize model parameters, and can automatically generate industry templates for users to use. The entire system forms a complete and efficient closed loop, from segmentation, generation, and adaptation to batch processing and feedback optimization. It can meet diverse image editing needs, improve processing efficiency and quality, and bring users a convenient, accurate, and safe image editing experience.
[0015] This application contains three core innovations. First, a cross-device adaptive editing architecture. Its innovative design supports five types of devices, including PCs, mobile devices, and workstations, through a device type recognition module. It automatically adjusts algorithm complexity, using a lightweight model branch for mobile devices and a full-precision model for workstations, achieving a single architecture adaptable to multiple scenarios. In terms of quantitative results, the inference latency for mobile devices is no more than 300 milliseconds, a 50% reduction compared to the technology disclosed in CN120303690A; the processing accuracy for workstations is no less than 98.5%, and the difference in editing effects across different devices is no more than 2%, solving the pain points of existing technologies that focus on single-device optimization and fragmented cross-device experiences. Second, a multimodal instruction precision parsing algorithm. This algorithm innovatively integrates the BERT-base model for text semantic parsing with the Wav2Vec2.0 model for speech instruction recognition, supporting mixed text and speech instructions. Through an attention mechanism, it focuses on core editing needs, achieving an instruction parsing accuracy of no less than 95%. In terms of quantitative results, compared to the single text commands disclosed in CN113516148B, the parsing efficiency has been improved by 40%, and the recognition accuracy of complex commands, such as replacing the background with an autumn forest and enhancing warm tones while preserving the edge details of the person, has been improved by 35%, lowering the user's operational threshold. Thirdly, the real-time feedback closed-loop optimization mechanism innovatively converts user feedback, such as ratings and labeled regions, into model optimization parameters in real time. The dynamic adjustment range of the segmentation model edge weights is ±0.05, and the iteration step size of the AIGC-generated constraint coefficients is 0.03, achieving an end-to-end closed loop of editing-feedback-optimization. The quantitative results show that after 1000 user feedback iterations, the satisfaction rate of the editing effect increased from the initial 82% to 94%, and the rate of repeated editing decreased by 60%, solving the shortcomings of existing technologies that have fixed effects and cannot be personalized for optimization. Attached Figure Description
[0016] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart illustrating an image editing method based on artificial intelligence, as described in an embodiment of this application. Figure 2 This is a schematic diagram of an artificial intelligence-based image editing device according to an embodiment of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0018] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0019] Current image scene element replacement technologies can be broadly categorized into three types, each with significant technical limitations. First, traditional manual editing methods rely on professional tools like Photoshop and After Effects, using manual image cutout, element stitching, and parameter fine-tuning to replace scene elements. However, this approach suffers from several core problems: the process is extremely cumbersome, requiring 1 to 3 hours to process a single complex scene image, making it unsuitable for batch processing; it demands high levels of professional skill and aesthetic judgment from operators, resulting in inconsistent output across different operators; this approach only achieves "physical replacement" of elements, failing to accurately match the original image's lighting, color scheme, and texture, leading to issues like "element floating" and "atmosphere fragmentation" after replacement, resulting in poor visual harmony; furthermore, the high labor costs make it difficult to support large-scale content production for enterprises. Second, automated replacement solutions based on traditional computer vision typically utilize image segmentation algorithms, such as Mask R-CNN and U-Net, to extract the outline of the target element, call a pre-defined element library to complete the replacement operation, and optimize the adaptation effect through simple color adjustments. For example, some image editing tools use this technology for their "one-click background replacement" function. However, this type of solution has obvious technical bottlenecks: the segmentation accuracy is limited by the complexity of the scene, and the segmentation effect is poor for transparent objects, hair edges, and low-contrast elements, easily resulting in jagged edges, missing areas, and unnatural replacement boundaries; element replacement relies on a preset resource library and lacks dynamic generation capabilities, making it difficult to meet users' needs for replacing personalized and scarce scene elements; atmosphere adaptation only stays at the level of basic color adjustment, without involving the collaborative optimization of light and shadow direction, texture, and scene semantics, and there is still a noticeable disconnect between the replaced element and the original image; it is difficult to support batch processing while also taking into account personalized parameter customization, and adaptability and efficiency cannot be guaranteed at the same time; and there is a lack of data security protection mechanisms, making data leakage easy during batch processing. Thirdly, existing AIGC-assisted replacement solutions generate replacement elements by calling general AIGC models, such as StableDiffusion and Midjourney, and then combine them with traditional image editing techniques to complete the replacement operation.While such solutions address some of the needs for "personalized element generation," they still suffer from key drawbacks: poor controllability of generated results; general AIGC models struggle to accurately respond to constraints such as "element size adaptation, fixed position, and style consistency with the original image," often requiring manual adjustments and resulting in inefficiency; lack of deep scene atmosphere perception and collaborative adaptation mechanisms; resulting in color, lighting, and texture of generated elements that are out of sync with the original image, necessitating manual fine-tuning to achieve harmony; fragmented element generation and replacement processes, failing to achieve full-chain automation from "requirement input - element generation - intelligent replacement - atmosphere adaptation," still requiring manual intervention to connect each stage; reliance on external general AIGC model interfaces, lacking deep integration with proprietary platforms; data security and generation efficiency limited by external networks and interfaces, and inability to be customized for specific business scenarios; and unclear copyright ownership of generated results, posing potential legal risks.
[0020] To address the aforementioned problems, this application provides an artificial intelligence-based image editing system, comprising: The semantically enhanced segmentation module is used to preprocess the input image and perform image segmentation on the preprocessed image to obtain the target element contour mask and the subject protection mask; The platform-based AIGC generation module is used to construct composite instructions based on pre-set multi-dimensional constraints, perform model adaptation based on the composite instructions, and generate a verification iteration mechanism to determine image elements. A multi-dimensional atmosphere collaborative adaptation module is used to fuse the image generation elements with the original image through atmosphere depth perception and adaptive adjustment; The end-to-end automation and batch processing module is used to perform end-to-end automated processing and batch optimization based on a pre-configured workflow engine. The feedback-driven iteration module is used to collect user feedback information and perform iterative optimization based on the feedback information.
[0021] In one embodiment, the semantically enhanced segmentation module includes an image preprocessing unit, a fusion segmentation model unit, and an optimization unit.
[0022] The image preprocessing unit receives the original image and the required instructions. After identifying the types of elements to be replaced or protected, it uses a self-developed adaptive noise reduction algorithm that integrates Gaussian filtering and median filtering techniques. First, it removes noise from the image, converts it to the RGB color space, and standardizes the resolution to 1080P or 4K to adapt to different needs, while preserving EXIF lighting parameters. Next, it performs region analysis and parameter calculations, dividing the input image into several small blocks, such as 32x32 pixel blocks. For each block, it calculates the local variance and average gradient of its pixel values. A small local variance indicates a smooth region, such as a solid-color wall in a warehouse; a large average gradient indicates the region contains edges or textures, such as a trench coat collar. Adaptive filtering is then applied. For smooth regions (low variance, low gradient), a Gaussian filter with larger parameters (e.g., a 5x5 kernel with σ of 1.5) and a 5x5 window median filter are used to strongly eliminate uniform noise. For textured or edge regions (high variance, high gradient), a Gaussian filter with smaller parameters (e.g., a 3x3 kernel with σ of 0.5) and a 3x3 window median filter are used to slightly reduce noise while preserving edge information to the maximum extent. Finally, intelligent fusion is performed, weighting the two filtering results for each region. The weights are dynamically adjusted according to the texture complexity of the region. In smooth regions, the Gaussian filter result has a higher weight (e.g., set to 0.7) to ensure smoothness; in textured edge regions, the median filter result has a higher weight (e.g., set to 0.7) to better preserve edges and eliminate salt-and-pepper noise. Taking the processing of a trench coat product image as an example, the algorithm performs strong noise reduction on smooth areas such as the warehouse background, while performing weak noise reduction on textured edge areas such as the fur collar and the edge of the transparent packaging bag, thus preserving clear boundaries for subsequent segmentation operations.
[0023] In one embodiment, the fusion segmentation model unit adopts a "ViT-L global semantic extraction + improved Mask R-CNN" architecture. This architecture is a feature-level deep fusion architecture, not a simple concatenation. ViT-L, as the "semantic brain," provides global contextual guidance for the "local perception" of Mask R-CNN.
[0024] In ViT-L's global semantic extraction, the entire image is input, and a Transformer encoder is used to obtain a feature map rich in global semantic information. This feature map encodes scene-level semantics, such as "indoor warehouse"; subject-level semantics, such as "trench coat subject"; and relation-level semantics, such as "trench coat wrapped in a transparent bag". In the feature fusion stage, the global semantic feature map output by ViT-L is upsampled or transformed and then fused layer by layer with the multi-scale local feature maps extracted by the CNN backbone network of Mask R-CNN, such as ResNet-FPN. This allows Mask R-CNN to perceive the global context while extracting local features.
[0025] The improved Mask R-CNN's RPN (Region Proposal Network) proposes regions on the feature map after feature fusion. By incorporating global semantics, the RPN can more accurately propose regions related to the target semantics, such as "background," while suppressing irrelevant regions. Its masking branch introduces an edge attention mechanism, further improving the segmentation accuracy for fur collars and transparent edges.
[0026] For example, ViT-L first understands that the image is "a trench coat in a warehouse". When Mask R-CNN processes a local region, global semantics guides it: if the region is a shelf, it should be classified as "background"; if it is the outline of a trench coat, even if the edges are blurred, it should be classified as "foreground" based on global cognition.
[0027] To adapt to the computing power of the self-developed platform, the ViT-L model extracts global semantic features, such as "outdoor rainy day - transparent raincoat - backpack as the main subject," enhancing the understanding of complex scenes. The extracted global semantic features need to cover three levels: scene level, i.e., the overall environment of the image, such as "indoor warehouse"; subject level, i.e., the category and attributes of the main object, such as "windbreaker" and "has a fur collar"; and relationship level, i.e., the spatial and logical relationships between objects, such as "transparent bag wrapping the windbreaker" and "windbreaker is in the center." The extraction process is as follows: the input image is divided into fixed-size image blocks, such as 16x16 pixels, and each image block is linearly projected into a feature vector. A learnable classification label is added before the vector sequence, and positional encoding is added to all vectors to preserve spatial information. The sequence is input into ViT-L's Transformer encoder. Through a multi-layer self-attention mechanism, the features of each image block interact with the features of all other image blocks, and the classification label ultimately converges the global information of the entire image. At output, the output vector corresponding to the classification label is taken as the global semantic feature vector of the image; this vector is the dense numerical representation of the decoded global semantics.
[0028] The RPN network has been optimized, and five new small-sized, high-resolution anchor frames have been added to improve the detection accuracy of small elements and elements with blurred edges. The five anchor frames are defined by pixel area and aspect ratio: 8x8 pixels, 1:1 aspect ratio, for detecting extremely small targets such as small buttons and ornaments on a trench coat; 16x16 pixels, 1:1 aspect ratio, for detecting slightly larger buttons, zipper pulls, and small labels on packaging bags; 24x24 pixels, 1:1 aspect ratio, for detecting medium to small parts such as cuffs and pocket flaps on a trench coat; 16x16 pixels, 1:2 aspect ratio, for detecting thin seams and narrow decorative strips; and 16x16 pixels, 2:1 aspect ratio, for detecting horizontal small targets such as brand logos and belt loops.
[0029] An edge attention mechanism is introduced to focus on edge pixel features, optimize mask generation, and achieve pixel-level segmentation. The edge attention branch starts from a shared feature map and calculates an edge attention map through several convolutional layers. Each pixel in this map has a value between 0 and 1, representing the probability that the location belongs to the edge of an object. The edge attention map is multiplied by the feature map of the main mask branch. The pixel feature values in edge regions are amplified, while those in non-edge regions are relatively suppressed. This essentially gives the main branch a hint, prompting it to focus on the boundaries. The modulated feature map is then convolved and upsampled to generate the final segmentation mask. In an example, when segmenting the fur collar of a trench coat, this mechanism generates an attention map highlighting the fur collar's edge. The main branch is thus guided to finely process the boundaries of each individual hair, resulting in a mask with a natural, textured feel rather than a jagged, hard edge.
[0030] In one embodiment, the optimization unit combines morphological dilation and erosion operations with a Bézier curve smoothing algorithm to eliminate jagged edges and output a target element contour mask and a main body protection mask, achieving a segmentation accuracy of over 99.5%. In the morphological operations, the opening operation employs an erosion-then-dilation approach. First, a small kernel, such as a 3x3 kernel, is used for erosion to remove isolated noise points and small burrs at the edges. Then, dilation restores the approximate size of the main body. This process smooths the contour and eliminates small particles. The closing operation, on the other hand, dilates first and then erodes, which can be used to fill small holes inside the mask caused by prediction errors. During the Bézier curve smoothing process, a sequence of contour points is first extracted from the morphologically processed mask. Then, key control points are selected based on the curvature changes of the contour points. Next, a third-order Bézier curve is fitted between adjacent key points to generate a smooth curve path. Finally, the smoothed curve is converted back into a pixel-level binary mask. Taking the fur collar of a trench coat as an example, its initial mask edges are jagged. After removing obvious noise through morphological operations, Bézier curve fitting is used to transform the jagged lines into smooth arcs, ultimately resulting in a visually natural fur outline.
[0031] In one embodiment, the platform-based AIGC generation module includes a multi-dimensional constraint instruction construction unit, a model adaptation unit, and a generation verification iteration unit.
[0032] The multi-dimensional constraint instruction building unit includes a self-developed instruction parsing engine. This engine is a multimodal information fusion unit responsible for translating different types of constraints into unified control signals that the AIGC model can understand. Its core is based on feature space alignment and cross-attention control technology. During implementation, the scene feature constraints are first parsed. It receives the atmosphere vector F from the atmosphere perception module, which contains information such as color and lighting. The encoder within the engine converts this into a scene style condition vector C_scene, used to control style consistency during generation. For example, the generated lighting texture must match the original image's trench coat. Next, user requirement constraints are parsed. An NLP model is used to parse user instructions, such as "outdoor maple forest, realistic, warm color tone," converting them into a standardized text embedding vector C_text and generation parameters P_user. For example, "warm color tone" is quantized into a color temperature parameter. Then, segmentation mask constraints are parsed. The target mask image obtained from the segmentation module is encoded into a spatial feature map M_mask through a small CNN. This map is used to strictly limit the generation area of the AIGC model, achieving pixel-level localization. Finally, the scene feature constraints are integrated, namely the color, light and shadow, and texture vectors of the atmosphere perception module, user requirement constraints, namely size, position, style, and details, and segmentation mask constraints, namely the outline of the target element, to generate a composite instruction of "semantic text + feature vector + mask boundary".
[0033] The model adaptation unit is based on a self-developed AIGC model with an optimized Diffusion architecture, and adds a platform-based constraint layer. This constraint layer has three key functions: first, it uses scene feature vectors to guide the generated content to maintain stylistic consistency; second, it strictly limits the generation range through mask boundaries to ensure accurate and controllable content generation; and third, it transforms user requirement parameters into specific attribute constraints, such as converting requirements like "retro style" and "low saturation" into attribute conditions that the model can recognize, thereby improving the quality and relevance of the generated content.
[0034] The generation and verification iteration unit incorporates a similarity algorithm that verifies the matching degree between the generated elements and the given constraints. If the matching degree is lower than 96%, the system will automatically start the iterative generation process, performing a maximum of 3 iterations to ensure that the final generated result meets the requirements without manual adjustment.
[0035] In one embodiment, the multi-dimensional atmosphere collaborative adaptation module includes an atmosphere depth perception unit, an adaptive adjustment unit, and a global optimization unit.
[0036] The atmosphere depth perception unit includes a self-developed feature extraction engine, a modular visual feature quantification system that integrates traditional algorithms with lightweight models. This engine transforms the subjective perception of "atmosphere" into objective and computable feature vectors. For example, it first extracts color features by converting the image to the HSV color space and calculating global and regional hue histograms, mean saturation, mean brightness, and other statistical quantities to form a color feature vector. For instance, it analyzes that the original image's dominant hue angle is 25°, presenting a warm yellow tone. Next, it extracts light and shadow features, starting with the brightness channel. By analyzing the distribution of highlights and shadows, it estimates the direction of the light source, such as the vector [-0.8, -0.2] indicating light coming from the upper left. It also estimates the light intensity, such as a scalar of 0.75. Finally, it extracts texture features using local binary mode or gray-level co-occurrence matrix algorithms to calculate texture indices such as image roughness and contrast, forming a texture feature vector. Finally, semantic atmosphere features are extracted. Using pre-trained scene recognition and object detection models, scene categories (e.g., "indoor warehouse") and subject categories (e.g., "trench coat") are identified. Then, through sentiment analysis or semantic embedding models, a semantic atmosphere feature vector is generated. The engine ultimately extracts four types of features: color features (covering RGB distribution, hue angle, saturation, and brightness statistics); lighting features (including direction, intensity, and shadow parameters); texture features (including roughness, smoothness, and material type); and semantic atmosphere features (involving scene type, emotional tone, and semantic associations), generating a unified atmosphere vector.
[0037] In terms of color adaptation, the adaptive adjustment unit employs a self-developed color transfer algorithm. The core principle of this algorithm lies in extracting color feature information from the original image, including key elements such as color distribution, hue, saturation, and brightness. Then, based on this feature information, the colors of the generated elements are reshaped and adjusted. By establishing a color mapping relationship between the original image and the generated elements, the generated elements are made as close to the original image in color as possible, achieving a color deviation control within 3% and realizing harmonious color adaptation. Specific steps include image preprocessing, color feature extraction, establishing a color mapping relationship, color transfer and adjustment, processing, and output. For lighting and shadow adaptation, shadows, highlights, and reflections are precisely added to the generated elements based on the lighting and shadow parameters of the original image. For texture adaptation, texture mapping and detail enhancement techniques are used to unify the texture style of the generated elements with the original image. Semantic adaptation optimizes the attributes of the generated elements based on scene semantics. For example, in an autumn scene, a "fallen leaf" texture is added to the generated elements, making the generated content more in line with the scene atmosphere.
[0038] The global optimization unit performs adaptive color balancing to make the color distribution of generated elements more harmonious and natural; it also performs sharpening to improve image clarity and adjusts contrast to enhance the image's depth. Through this series of operations, it ensures that the generated elements can be completely integrated with the original image.
[0039] In one embodiment, the end-to-end automation and batch processing module includes an end-to-end automation unit, a batch optimization unit, a batch adaptation unit, and a data security unit.
[0040] The fully automated unit relies on a self-developed workflow engine to build an end-to-end automated process from "preprocessing" to "segmentation", then to "generation", "replacement" and "adaptation", and finally to "output", with a single image processing time of no more than 2 minutes.
[0041] The batch optimization unit, by introducing a platform-based parallel computing framework, has powerful batch processing capabilities, supporting the simultaneous processing of 1,000 images per batch, with a total processing time of no more than 20 minutes for 100 images.
[0042] The batch adaptation unit has a flexible processing mode, which can both support efficient processing of batch images by uniformly applying the same parameters, and can also receive personalized instructions for individual images to generate customized output. While improving processing efficiency, it fully meets diverse customization needs.
[0043] The data security unit enables the entire processing flow to operate in a closed-loop manner, eliminating reliance on external interfaces and ensuring the independence and security of data processing. In the data storage and transmission stages, both raw images and generated data are protected using AES-256 encryption technology, strictly complying with relevant regulatory requirements.
[0044] In one embodiment, the feedback-driven iteration module includes a feedback acquisition unit, a model optimization unit, and a template expansion unit.
[0045] The feedback collection unit collects user ratings for the processing results, ranging from 1 to 10 points. It can also collect submitted text-based modification suggestions and mark the areas that need to be modified. The platform will automatically associate this feedback information with the corresponding processing parameters.
[0046] The model optimization unit periodically integrates the collected feedback data into the training set and uses this data to iteratively optimize the edge weights of the segmentation model, the constraint coefficients generated by AIGC, and the threshold of the adaptation algorithm in order to continuously improve the performance of the model and algorithm.
[0047] The template extension unit has the ability to automatically generate industry-specific templates, covering multiple fields such as e-commerce, film and television, and games. Users can directly call these templates to batch process images, which greatly improves processing efficiency.
[0048] In one embodiment, regarding cross-device adaptive algorithms, the device type identification module automatically classifies devices by analyzing device hardware parameters such as the number of CPU cores, GPU memory, and screen resolution. For mobile devices, such as smartphones, a series of optimization measures are implemented. In terms of model lightweighting, the segmentation model uses MobileNet instead of ResNet, reducing the number of parameters by 70%; in terms of algorithm simplification, the Gaussian filter kernel size is fixed at 3×3, and the number of Bézier curve control points is reduced to 3; in terms of data compression, intermediate feature maps use FP16 precision storage, reducing memory usage by 50%. For workstation devices, a full-precision mode is enabled. The model fully loads the complete architecture of ViT-L and the improved MaskR-CNN, achieving a segmentation accuracy of no less than 98.5%; during batch processing, 8 threads are used for parallel computation, processing 100 images takes no more than 15 minutes.
[0049] The multimodal instruction parsing process is as follows: For text instructions, semantic vectors are extracted using the BERT-base model with a dimension of 768, and the keyword recognition accuracy is no less than 96%; for speech instructions, the Wav2Vec2.0 model converts speech into text, with a recognition accuracy of no less than 93%, and supports both Mandarin Chinese and English; for mixed instruction fusion, attention weight allocation is adopted, with a text weight of 0.6 and a speech weight of 0.4, to generate a unified instruction vector, and the mixed instruction parsing accuracy is no less than 95%.
[0050] The real-time feedback closed-loop execution steps are as follows: First, feedback is collected. Users provide feedback through ratings, such as 1-10 points; textual opinions, such as "the background color is too dark"; and rectangular boxes annotating the target area. Feedback data is uploaded to the server in real time. Next, parameter mapping is performed. When the rating is below 6 points, model parameter adjustments are triggered, and the weights of the corresponding modules are optimized. For example, annotating the edge area will increase the edge weights of the segmentation model by 0.05. Then, iterative training is performed. Small-batch training is automatically performed every morning at midnight, with a batch size of 32 and a learning rate of 1e-5. The updated model takes effect the next day. Finally, the effect is verified. After each iteration, 100 editing tasks are randomly selected for verification. If the effect improvement is not less than 5%, the parameters are retained; otherwise, the previous version is rolled back.
[0051] The data interaction between modules is as follows: The semantically enhanced segmentation module outputs the target element contour mask to the platform-based AIGC generation module. The mask has the same resolution as the original image, as well as the edge feature vector with a dimension of 256. Data transmission is compressed using the protobuf format, and the transmission latency does not exceed 50ms. The multi-dimensional atmosphere collaborative adaptation module outputs atmosphere adaptation parameters, including color deviation values, light and shadow adjustment coefficients, and user editing logs, to the feedback-driven iteration module. After associating with user feedback, these parameters are stored in the training dataset. The end-to-end automation and batch processing module sends device type parameters, batch processing quantity, and security encryption keys to each functional module. Asynchronous communication is achieved using the MQTT protocol to ensure system stability during batch processing.
[0052] like Figure 1 As shown in the figure, an image editing method based on artificial intelligence provided in this application includes: S101. Preprocess the input image and perform image segmentation on the preprocessed image to obtain the target element contour mask and the subject protection mask.
[0053] S102. Construct composite instructions based on pre-set multi-dimensional constraints, perform model adaptation based on the composite instructions, and generate a verification iteration mechanism to determine image elements. The composite instructions include semantic text, feature vectors, and mask boundaries.
[0054] S103. Extract features from the preprocessed image, determine the atmosphere vector based on the extracted features, and adaptively adjust the image based on the atmosphere vector to fuse the generated image elements with the original image.
[0055] S104. Perform end-to-end automated processing and batch optimization based on the pre-configured workflow engine.
[0056] S105. Collect user feedback information to perform iterative optimization based on the feedback information.
[0057] In one embodiment, the platform initializes by launching a self-developed AIGC creation platform, loading a pre-trained semantically enhanced segmentation model, a self-developed controllable AIGC model, an atmosphere feature extraction engine, an outdoor scene label library, and an AES-256 encryption module. This strong segmentation model integrates ViT-L and an improved Mask R-CNN. Next, parameter configuration is performed, setting the resolution to 1920×1080, the segmentation confidence threshold to 0.97, the AIGC generation iteration to 50 steps, the atmosphere adaptation deviation to no more than 3%, the batch thread count to 32, and the generation matching threshold to 96%. The encryption key is automatically generated and associated with the user's account. Users upload 200 images of trench coats containing transparent packaging bags and fur collars. The uploaded data is transmitted via AES-256 encryption. Simultaneously, users input requirements, requesting the background to be replaced with an outdoor maple forest scene, with a realistic style, warm colors, highlighting the trench coat, and batch processing. Ten images are specified as being of a cloudy maple forest. The system parses the unified and personalized constraints. Subsequently, batch preprocessing is performed, including adaptive noise reduction, RGB color space conversion, resolution unification, and preservation of lighting parameters. In the semantic segmentation stage, ViT-L extracts the global semantics of "warehouse background - trench coat body - transparent packaging bag". RPN generates 18 anchor boxes to locate the warehouse background, and the edge attention mechanism optimizes the mask, outputting the background outline mask and the trench coat protection mask, achieving a batch segmentation accuracy of 99.3%. During scene atmosphere depth perception, atmosphere features are extracted in batches, such as an RGB distribution of R=30%, G=35%, B=35%, light and shadow direction diagonally upward, medium-coarse texture, and semantics of "product display - minimalist", generating a unified atmosphere vector F; and generating a low-saturation warm color vector F1 for 10 specified images, with a hue angle of 15°-30° and a 20% reduction in brightness. In the platform-controlled AIGC generation process, constraint instructions are constructed: 95% of the images use F vectors paired with the "realistic maple forest + warm color tone" instruction, and 5% use F1 vectors paired with the "overcast maple forest + low saturation" instruction, all bound to a warehouse background mask. The self-developed AIGC model is guided by the constraint layer for generation. After verification, three images were iterated once to achieve a 96.8% matching degree, and the rest met the standard on the first try. An improved Poisson fusion algorithm is used for replacement fusion, fitting elements and optimizing the edge transition of the transparent packaging bag. During multi-dimensional atmosphere collaborative adaptation, batch color adjustments are performed, such as adjusting the G channel ratio of the maple forest to 32% and increasing the R channel of the warm color tone image to 38%; adding an oblique upper projection with an intensity of 70% of the original image; adjusting the texture roughness of maple leaves to unify the texture; adding "fallen leaves" details to enhance the semantics of the autumn scene; and simultaneously performing batch color balancing and sharpening to ensure the trench coat subject stands out. The entire process takes 38 minutes, with an average of 11.4 seconds per image, and a list of results is generated for users to preview. The user provided feedback on two overcast images, suggesting that the low saturation be increased by 10%. The platform adjusted the F1 vector to re-adapt the output and collected the feedback data, then encrypted and exported JPG / PNG files. The entire data process was closed-loop with no leakage.At the end of the month, model parameters were iterated based on feedback data to improve the accuracy of adaptation to cloudy scenes. Regarding key technical details, the segmentation model was trained using the COCO dataset and 300,000 e-commerce product images. ViT-L froze the parameters of the first 16 layers, and the anchor boxes were optimized to 18 types. The self-developed AIGC model adopted a Diffusion architecture and a cross-attention constraint layer, balancing efficiency and performance through 50 iterations. The computing power configuration consisted of a GPU cluster containing six RTX4090s, with each GPU consuming no more than 12GB of resources, supporting single-image processing on PCs with RTX3070 or higher. Security was ensured through storage encryption and an SSL transmission channel, accessible only to authorized users, complying with data security regulations.
[0058] In one embodiment, this application employs a "ViT-L global semantic extraction + improved Mask R-CNN" architecture. ViT-L provides global context guidance, while the improved Mask R-CNN's RPN integrates global semantics to more accurately propose regions. The mask branch introduces an edge attention mechanism to address the accuracy issues in segmenting transparent objects and hair edges. A multi-dimensional constraint instruction construction unit translates different constraints into a unified signal, using scene feature vectors, mask boundaries, and user-required parameters to guide generation, improving content quality and relevance. This application transforms subjective atmosphere into objective feature vectors, with an adaptive adjustment unit adapting from multiple dimensions including color, lighting, texture, and semantics, and a global optimization unit ensuring complete integration of generated elements with the original image.
[0059] like Figure 2 As shown in the illustration, this application also provides an image editing device based on artificial intelligence, including: At least one processor; and, A memory that is communicatively connected to at least one processor; wherein, The memory stores instructions that can be executed by at least one processor to enable an artificial intelligence-based image editing device to perform the method described in the above embodiments.
[0060] This application also provides a non-volatile computer storage medium storing computer-executable instructions, which are configured as described in the above embodiments.
[0061] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0062] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0063] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0064] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware components.
[0065] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and medium embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the description of the method embodiments.
[0066] The devices and media provided in this application are one-to-one with the methods. Therefore, the devices and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0067] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0068] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0069] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0070] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0071] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0072] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0073] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0074] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. An image editing system based on artificial intelligence, characterized in that, include: The semantically enhanced segmentation module is used to preprocess the input image and perform image segmentation on the preprocessed image to obtain the target element contour mask and the subject protection mask; The platform-based AIGC generation module is used to construct composite instructions based on pre-set multi-dimensional constraints, perform model adaptation based on the composite instructions, and generate a verification iteration mechanism to determine image elements. A multi-dimensional atmosphere collaborative adaptation module is used to fuse the image generation elements with the original image through atmosphere depth perception and adaptive adjustment; The end-to-end automation and batch processing module is used to perform end-to-end automated processing and batch optimization according to a pre-set workflow engine. The end-to-end automation and batch processing module includes a device type identification unit, which is used to parse device hardware parameters and automatically adjust algorithm complexity to adapt to multiple types of devices, including PCs, mobile devices, and workstations. The feedback-driven iteration module is used to collect user feedback information and perform iterative optimization based on the feedback information.
2. The system according to claim 1, characterized in that, The semantically enhanced segmentation module includes: The image preprocessing unit uses a pre-set noise reduction algorithm to remove noise from the image and unifies the resolution and color space of the image; it divides the image and calculates the local variance and average gradient of the divided image to determine smooth regions and edge regions based on the local variance and average gradient; it applies Gaussian filtering and median filtering with corresponding parameters to each region to preserve edge information. The fusion segmentation model unit uses a pre-set segmentation architecture to perform deep fusion and pixel segmentation; The optimization unit eliminates edge jaggedness and outputs a mask through morphological operations and Bézier curve smoothing.
3. The system according to claim 1, characterized in that, The platform-based AIGC generation module includes: The multi-dimensional constraint instruction construction unit parses multi-dimensional constraints according to a pre-set instruction parsing engine. The multi-dimensional constraints include scene feature constraints, user requirement constraints, and segmentation mask constraints. It generates composite instructions based on the multi-dimensional constraints. The composite instructions are used by the parsing engine and support mixed text and speech instructions. The model adaptation unit guides scene feature vectors, limits mask boundaries, and transforms requirement parameters through a pre-set AIGC model. A verification iteration unit is generated, and the matching degree between image elements and multi-dimensional constraints is determined by a similarity algorithm. The unit then automatically iterates based on the matching degree.
4. The system according to claim 1, characterized in that, The multi-dimensional atmosphere collaborative adaptation module includes: The atmosphere depth perception unit extracts features from the image, including color, light and shadow, texture, and semantic atmosphere features. The semantic atmosphere features include scene type, emotional tone, and semantic association. An atmosphere vector is generated based on the features. An adaptive adjustment unit performs color adaptation, lighting adaptation, texture adaptation, and semantic adaptation on the image based on the atmosphere vector. The global optimization unit performs color balance, sharpening, and contrast adjustments on the image to ensure that image elements are fully integrated with the original image.
5. The system according to claim 1, characterized in that, The end-to-end automation and batch processing module includes: The fully automated unit performs end-to-end automated processes through a pre-set workflow engine. The automated processes include preprocessing, segmentation, generation, replacement, adaptation, and output. The batch optimization unit introduces a platform-based parallel computing framework to perform batch processing on the images; Batch adaptation unit is used to balance batch uniform parameters and individual personalized instructions; The data security unit is used to encrypt the storage and transmission of original images and generated data to ensure data security.
6. The system according to claim 1, characterized in that, Also includes: The feedback-driven iteration module is used to collect user feedback information and perform iterative optimization based on the feedback information.
7. The system according to claim 6, characterized in that, The feedback-driven iteration module includes: The feedback collection unit collects user feedback on the processing results, including ratings, text modification suggestions, and marked areas. The model optimization unit incorporates feedback data into the training set and iteratively optimizes the segmentation model, AIGC-generated constraint coefficients, and adaptation algorithm thresholds based on the training set. The template extension unit automatically generates industry templates, allowing users to directly call these templates for batch processing.
8. An image editing method based on artificial intelligence, characterized in that, The method, applied to an AI-based image editing system as described in any one of claims 1-7, comprises: The input image is preprocessed, and the preprocessed image is segmented to obtain the target element contour mask and the subject protection mask. A composite instruction is constructed based on pre-set multi-dimensional constraints. The model is adapted based on the composite instruction, and a verification iteration mechanism is generated to determine image elements. The composite instruction includes semantic text, feature vector, and mask boundary. Feature extraction is performed on the preprocessed image, an atmosphere vector is determined based on the extracted features, and the image is adaptively adjusted based on the atmosphere vector to fuse the generated image elements with the original image. End-to-end automated processing and batch optimization are performed based on a pre-configured workflow engine, automatically adapting to device types and adjusting algorithm complexity; Collect user feedback information to perform iterative optimization based on the feedback information.
9. An image editing device based on artificial intelligence, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the artificial intelligence-based image editing device to perform the method as described in claim 8.
10. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are configured as described in claim 8.